Back in 2011, Harvard Professor Raj Chetty and two esteemed colleagues (John Friedman at Brown University and Jonah Rockoff at Columbia University) published a dazzling study of teachers, asserting that the best teachers are those whose students get improved scores. Those students have a higher income ($250,000 over their lifetimes), and enjoy a multitude of benefits, all because of that one teacher who induced them to have higher scores. President Obama cited Chetty’s research in his 2012 State of the Union address to show how important it was to find the “best” teachers and fire the “worst’ teachers.
Chetty’s work supported the Obama-Duncan Race to the Top plan to encourage evaluating teachers by the test scores of their students. States that evaluated teachers by their students’ scores were eligible to apply for a share of RTTT funding. Those who did not were not eligible.
Most states, eager for a share of the $5 billion prize, agreed to adopt what was called “value-added modeling” or “value-added measurement.” (VAM)
I posted dozens of times about the flaws of VAM, first of all, because the American Statistical Association said that teachers account for only 1-14% of score changes; most changes were attributable to home life and school system issues. Secondly, because the VAM concept is very unstable and is highly affected by student demographics. Only teachers of reading and math in grades 4-9 could even be assessed by the annual tests, which are mandated only in grades 3-8. Schools started attributing scores to teachers not in those subjects and not in those grades, tied to the work of other teachers in the school.
The Los Angeles Times engaged researchers to calculate VAM scores for the district’s teachers, and the newspaper published them alongside the names of teachers. It was humiliating for teachers, but Arne Duncan thought this disclosure was wonderful.
I happened to be in Los Angeles on the day that a fifth grade teacher committed suicide after he received a poor VAM score. No one knows if that was the reason for his suicide, but it may have been. By all accounts, he was a good teacher in a tough school.
The New York Post did the same for New York City teachers. It listed names and scores. The teacher identified by the newspaper as the city’s “worst” teacher was hounded by Post reporters seeking an interview. It turned out that she taught classes of new immigrants, who moved in and then out of her class as they learned enough English to join regular classes. Her students in September were not the same students by June. The VAM scores for her were meaningless, as they were for other teachers. Teachers of the gifted saw few if any gains because their students were at the top year after year. Expert math teacher Gary Rubinstein noticed that some teachers had high scores in one subject, but not in the other. Should half the teacher get a bonus while the other half was fired?
Freddie deBoer, an independent writer who earned a doctorate in English and education assessment, re-evaluated Raj Chetty’s famous study.
It’s a long review, and I won’t post it all. Please open the link and read it.
He begins:
For a long time I’ve been getting some version of the comment, “What about Chetty!” in response to my perspective on education, as in Raj Chetty, the economist who for the past decade has made a lot of waves asserting that our education problems are straightforwardly the product of bad teachers and that replacing them will have implausibly large economic effects. I tend to try and work from a broader perspective than “this is why I think this guy is wrong,” but I get this request so often, here you go. This is why I think Raj Chetty is wrong.
Few empirical claims in modern education policy have traveled farther than Chetty et al’s findings on teacher “value-added.” In his famous 2014 American Economic Review papers, he and his coauthors reported that students assigned to better (excuse me, higher value-added) teachers were more likely to attend college, earn higher salaries, save for retirement, and avoid teen pregnancy, and that replacing a teacher in the bottom five percent of the distribution with an average teacher would raise the present value of a single classroom’s lifetime earnings by roughly $250,000. Chetty’s research findings in this domain had been floating around for awhile at the time of publication, and President Obama cited the figure in his 2012 State of the Union address, and the judge who decided Vergara v. California leaned on it to strike down California’s teacher tenure laws. Take that, teachers! The findings are arresting, the dataset is impressive – 2.5 million children, linked to IRS tax records! – and the policy implications are clean: identify and remove bad (pardon me, low “value-added”) teachers, watch outcomes improve. It’s exactlythe kind of story our neoliberal policy establishment is desperate to tell, and was clearly catnip to the Obama administration, which was doggedly attached to a simplistic vision of delivery through better education, where the gutting of the uneducated labor market was ameliorated by turning every last child in the United States into a genius, scaling up the Stanford-to-Google pipeline until every American could pass through it.
Unfortunately, the Chetty story is ultimately another neoliberal just-so story, that is to say, a fable, a legend, a myth. The closer you look at what the “value-added” construct actually measures, how stable those measurements are, and how the Chetty results have fared under replication, the more reason there is to doubt both the magnitude of the claimed effects and, more fundamentally, whether “teacher quality” as the literature operationalizes it is a coherent, measurable attribute at all. (Spoiler: it is not.) Let us count the problems.
The construct itself puts the thumb on the scale. The first problem is conceptual. In the Chetty et al. studies, a teacher’s “value-added” is the residual variation in a student’s standardized test scores that remains after controlling for prior achievement and some demographic covariates. It’s not a measure of pedagogical skill, content knowledge, classroom climate, the cultivation of curiosity, or any other property normally meant by “good teaching.” It’s a statistical residual on a narrow set of assessments, usually math and reading tests in grades three through eight. That residual is then defined as quality. I want to be clear about this: any portion of variability in student outcomes that Chetty et al cannot or will not identify otherwise is assumed to be a product of teacher inputs. Since Chetty’s whole project is to argue that educational outcomes are the result of teacher quality, this is what we used to call begging the question – that is, he’s assuming the point he wants to prove, asserting the desired conclusion as a premise, by acting as though any uncaptured variation is necessary evidence of teaching quality. And it gets worse in the telling. When advocates and journalists and politicians summarize his work, the construct expands silently from “the part of test-score gains Chetty cannot otherwise explain” to “good teachers,” and the slippage is rarely flagged. But that’s the whole game, you guys.
Open the link and enjoy!











