Bruce Baker is one of the stellar scholars in the field of economics of education. He is currently at the University of Miami. He is a Professor and Chair of the Department of Teaching and Learning in Miami’s School of Education and Human Development.
In this post, he takes issue with the economic prescriptions of Margaret Roza. Her work has been embraced by the privatizers.
Baker begins:
My critiques of Marguerite Roza, the Edunomics Lab, and Education Resource Strategies — with receipts. Every quotation is verbatim from the linked post or article.
24 blog entries, Aug 2009 – Aug 2026 (and counting)
21 taking on Roza, from CRPE to Gates to Georgetown
4 taking on ERS or Karen Hawley Miles
3 on Edunomics graphs, so far
The short version
I’ve spent an embarrassing share of my career, seventeen years and counting, cleaning up after Marguerite Roza. It has come in two waves. The first ran from 2009 to 2014, when she was at the Center on Reinventing Public Education (CRPE) and then advising the Gates Foundation, and her work was all over the federal policy conversation. The second started in 2025, when the Edunomics Lab she now runs at Georgetown began handing state policymakers graphs so bad I made a video about them. Education Resource Strategies (ERS) comes up less often, but it keeps turning up in the same places. Its founder, Karen Hawley Miles, co-wrote the Houston/Cincinnati weighted student funding (WSF) “success story” with Roza. Stephen Frank presented ERS slides alongside Roza’s fabricated graph at the 2011 Regents symposium. And ERS graded Baltimore’s Fair Student Funding against Baltimore’s own formula.
It comes down to the same four problems, over and over:
Blame the districts. The claim goes like this: states have fixed between-district inequity, so the problem left is within districts, and weighted student funding fixes that. The evidence for that sweeping national claim turns out to be “one or a handful of deeply flawed analyses,” mostly Roza’s Texas Weighted Student Index study. That study checks schools against the district’s own spending priorities, not against what kids actually need.
Productivity without the arithmetic.Stretching the School Dollar, Curing Baumol’s Disease, and the USDOE productivity page that showcased them offer spending cuts relabeled as “cost savings,” without a single cost-effectiveness analysis. Hank Levin laid out how to do one back in 1983. It isn’t a secret.
Evidence that was simply made up. Roza’s 2011 “productivity curve” had no data, no definitions and no connection to anything real. Researchers in the room said the claims were “simply made up.” ERS’s contribution at the same event: all teacher pay above the starting salary is waste.
Money-doesn’t-matter graphics. Edunomics’ long-term trend graphs start the clock in 2013, the one year that all but guarantees spending up and scores down, then skip the cost adjustment and stretch the axes. Their scatterplots throw every school onto one chart with no cost adjustment and call the resulting cloud a finding.
Yes, my tone has changed. In 2011 I called this work “methodologically flimsy” and “hack research.” By 2025 I was calling it “intentionally deceitful,” and I stand by that. Once you’ve been told, repeatedly and in public, exactly why a graph misleads, and you keep putting it in front of legislators anyway, “sloppy” stops being the right word. As I put it on Bluesky: “for anyone still using this kind of garbage, you’ve been on notice for years.”
The peer-reviewed version is politer, but it says the same thing. I’ve never said within-district inequity isn’t real. It is. What I have said, with data (EPAA, 2009), is that the WSF showcase districts were no more responsive to student need than districts without WSF. WSF was sold on advocacy research that “identifies the politically motivated solution then seeks to prove that it works.” And you can’t fix a district its state has starved by rearranging what’s left inside it. My 2013 NEPC review gave ERS’s Baltimore analysis the same treatment: grading a formula against itself isn’t an equity analysis. It’s a tautology.
Please open the link and read the rest of Bsker’s analysis. This post analyzes what happens when economic analysis becomes untethered from the needs of students.
The Center for Research on Education Outcomes (CREDO) has positioned itself as the think tank that writes trustworthy evaluations of charter schools. In a sense, this is strange because CREDO’s leadership and its funding are not neutral on the subject of charter schools; they are decidedly pro-charter.
Freddie DeBoer argues on his blog that CREDO has consistently reported that charter schools find no gains for charter schools as compared to public schools. But it’s 2023 report was widely reported as the year of the Big Charter Breakthrough. He believes they are honorable in reporting their findings, but they pulled a fast one on the media and public by inventing a new metric that cannot be compared to other measures: “days of learning.”
DeBoer shows how the CREDO metric magnifies charter school results, which are actually small enough to be counted as “no effect.”
DeBoer begins:
In 2023 the Center for Research on Education Outcomes, known as CREDO, released its third national study of charter schools. The research covered 1.8 million students across twenty-nine states plus Washington D.C. and New York City. CREDO’s study purported to offer the most definitive evidence yet about the relative school quality of charter schools as compared to traditional public schools. Certainly our media, political establishment, and the nonprofit sector have acted as though what CREDO has to say is definitive; the 2023 study is, I believe, the most-cited piece of charter research in existence. Every pro-charter politician reaches for it, your Jonathan Chait neoliberal “reform” types treat its findings with near-mythic reverence, people who otherwise have no exposure to educational research wave at it as the final word on the subject. When the average person talks loosely as though there’s some sort of generic charter advantage, they’re usually referencing CREDO research, even if they don’t know they are.
This is all rather surprising, given that CREDO has not produced robust evidence in favor of charter schools, not in the 2023 study or its precursors. As we shall see, even if we assume that CREDO has solved the problem of the fundamentally non-random samples inherent to the task of comparing charter to public – which we should not, because they have not – the evidence supplied by CREDO in the favor of charter schools is so marginal and so noisy as to be meaningless from a policy perspective. And yet the message that the media took and shared from that report – the good news!, the vindication!, the headlines in Education Week and the New York Post!, the press releases from the National Alliance for Public Charter Schools! – was that the typical charter student was the beneficiary of a meaningful learning boost compared to their traditional public school counterpart. Why did this message spread so widely and resonate so loudly, if I’m right that even CREDO’s own numbers show very modest differences? Because of a remarkably savvy bit of marketing in how CREDO’s numbers were presented. The report did not show its findings in terms of effect size differences between charter and public, which would be the field’s standard way of sharing results. Rather, the outcome that was shared again and again was that a typical charter school student in the dataset had gained sixteen additional “days of learning” per year in reading and six in math. For a few months there you couldn’t read educational media without running into those numbers: sixteen additional days of learning in reading per year, six in math. It became a bit of holy writ in the endless quest to make bad students good with ACCOUNTABILITY and market reforms.
Now, me, I immediately gravitate to that last bit: what on Earth is a “day of learning,” and since when is that a statistical measure? “Days of Learning”: A Jury-Rigged Metric Invented for the Purpose of Inflating the Effect If you share my confusion over that metric, don’t blame yourself. “Days of learning” is not merely an unusual measure of student performance; it’s a boutique metric that CREDO invented themselves for their charter studies. It’s like YouTube changing how views are counted so the numbers look bigger.
This should raise a number of red flags, most importantly when it comes to comparison with other research. We have standardized metrics in social science research for a reason. One core reason among them is mutual interpretability, the ability to compare outcomes involving different interventions that are measured with different instruments and in different units. That’s why effect size exists, and various expressions of effect size have served as the dominant measure of interest in all manner of human research for decades precisely because effect size enables ease of comparison. If you want to compare the influence of the same intervention when one study uses grades as the DV and another used test scores, effect size enables you to do that. If you want to compare two different types of intervention, effect size enables you to do that too. There are some caveats, as there always are, but the basic, elegant logic of effect size makes the measure remarkably flexible and portable: effect size tells us how a given intervention moves average outcomes along the distribution of outcomes. If you want to know more there are plenty of simple explainers of effect size out there; here’s mine. There are some fields where effect size is used less often (economics, for example) but the interpretability and flexibility of effect size metrics like Cohen’s d are powerful advantages. Not surprisingly, then, it’s the dominant reported measure in education research.
So if effect size is both useful and common, and is in fact the dominant measure of interest in the relevant field, why would CREDO come up with its own boutique measure? Why invite the potential for confusion with such consequential research? Occam’s razor suggests a simple answer: because CREDO is a pro-charter organization that’s intent on convincing stakeholders that charters have meaningfully superior outcomes when compared to traditional public schools… and their own research demonstrates the opposite. Their own numbers, in fact, suggest that the returns on charter schools are meager even if we ignore the selection problems that continue to plague charter advocacy.
This is a much longer post about the inadequacy of research on charter schools. It’s well-written and thoughtful. Please open the link and finish reading.
Back in 2011, Harvard Professor Raj Chetty and two esteemed colleagues (John Friedman at Brown University and Jonah Rockoff at Columbia University) published a dazzling study of teachers, asserting that the best teachers are those whose students get improved scores. Those students have a higher income ($250,000 over their lifetimes), and enjoy a multitude of benefits, all because of that one teacher who induced them to have higher scores. President Obama cited Chetty’s research in his 2012 State of the Union address to show how important it was to find the “best” teachers and fire the “worst’ teachers.
Chetty’s work supported the Obama-Duncan Race to the Top plan to encourage evaluating teachers by the test scores of their students. States that evaluated teachers by their students’ scores were eligible to apply for a share of RTTT funding. Those who did not were not eligible.
Most states, eager for a share of the $5 billion prize, agreed to adopt what was called “value-added modeling” or “value-added measurement.” (VAM)
I posted dozens of times about the flaws of VAM, first of all, because the American Statistical Association said that teachers account for only 1-14% of score changes; most changes were attributable to home life and school system issues. Secondly, because the VAM concept is very unstable and is highly affected by student demographics. Only teachers of reading and math in grades 4-9 could even be assessed by the annual tests, which are mandated only in grades 3-8. Schools started attributing scores to teachers not in those subjects and not in those grades, tied to the work of other teachers in the school.
The Los Angeles Times engaged researchers to calculate VAM scores for the district’s teachers, and the newspaper published them alongside the names of teachers. It was humiliating for teachers, but Arne Duncan thought this disclosure was wonderful.
I happened to be in Los Angeles on the day that a fifth grade teacher committed suicide after he received a poor VAM score. No one knows if that was the reason for his suicide, but it may have been. By all accounts, he was a good teacher in a tough school.
The New York Post did the same for New York City teachers. It listed names and scores. The teacher identified by the newspaper as the city’s “worst” teacher was hounded by Post reporters seeking an interview. It turned out that she taught classes of new immigrants, who moved in and then out of her class as they learned enough English to join regular classes. Her students in September were not the same students by June. The VAM scores for her were meaningless, as they were for other teachers. Teachers of the gifted saw few if any gains because their students were at the top year after year. Expert math teacher Gary Rubinstein noticed that some teachers had high scores in one subject, but not in the other. Should half the teacher get a bonus while the other half was fired?
It’s a long review, and I won’t post it all. Please open the link and read it.
He begins:
For a long time I’ve been getting some version of the comment, “What about Chetty!” in response to my perspective on education, as in Raj Chetty, the economist who for the past decade has made a lot of waves asserting that our education problems are straightforwardly the product of bad teachers and that replacing them will have implausibly large economic effects. I tend to try and work from a broader perspective than “this is why I think this guy is wrong,” but I get this request so often, here you go. This is why I think Raj Chetty is wrong.
Few empirical claims in modern education policy have traveled farther than Chetty et al’s findings on teacher “value-added.” In his famous 2014 American Economic Reviewpapers, he and his coauthors reported that students assigned to better (excuse me, higher value-added) teachers were more likely to attend college, earn higher salaries, save for retirement, and avoid teen pregnancy, and that replacing a teacher in the bottom five percent of the distribution with an average teacher would raise the present value of a single classroom’s lifetime earnings by roughly $250,000. Chetty’s research findings in this domain had been floating around for awhile at the time of publication, and President Obama cited the figure in his 2012 State of the Union address, and the judge who decided Vergara v. California leaned on it to strike down California’s teacher tenure laws. Take that, teachers! The findings are arresting, the dataset is impressive – 2.5 million children, linked to IRS tax records! – and the policy implications are clean: identify and remove bad (pardon me, low “value-added”) teachers, watch outcomes improve. It’s exactlythe kind of story our neoliberal policy establishment is desperate to tell, and was clearly catnip to the Obama administration, which was doggedly attached to a simplistic vision of delivery through better education, where the gutting of the uneducated labor market was ameliorated by turning every last child in the United States into a genius, scaling up the Stanford-to-Google pipeline until every American could pass through it.
Unfortunately, the Chetty story is ultimately another neoliberal just-so story, that is to say, a fable, a legend, a myth. The closer you look at what the “value-added” construct actually measures, how stable those measurements are, and how the Chetty results have fared under replication, the more reason there is to doubt both the magnitude of the claimed effects and, more fundamentally, whether “teacher quality” as the literature operationalizes it is a coherent, measurable attribute at all. (Spoiler: it is not.) Let us count the problems.
The construct itself puts the thumb on the scale. The first problem is conceptual. In the Chetty et al. studies, a teacher’s “value-added” is the residual variation in a student’s standardized test scores that remains after controlling for prior achievement and some demographic covariates. It’s not a measure of pedagogical skill, content knowledge, classroom climate, the cultivation of curiosity, or any other property normally meant by “good teaching.” It’s a statistical residual on a narrow set of assessments, usually math and reading tests in grades three through eight. That residual is then defined as quality. I want to be clear about this: any portion of variability in student outcomes that Chetty et al cannot or will not identify otherwise is assumed to be a product of teacher inputs. Since Chetty’s whole project is to argue that educational outcomes are the result of teacher quality, this is what we used to call begging the question – that is, he’s assuming the point he wants to prove, asserting the desired conclusion as a premise, by acting as though any uncaptured variation is necessary evidence of teaching quality. And it gets worse in the telling. When advocates and journalists and politicians summarize his work, the construct expands silently from “the part of test-score gains Chetty cannot otherwise explain” to “good teachers,” and the slippage is rarely flagged. But that’s the whole game, you guys.
Despite appearances, the most powerful person in the Trump administration is not Donald Trump: it’s Russell Vought, Director of Office of Management and Budget. He is the brains of this administration. Vought was at the Heritage Foundation and was one of the writers of project 2025. He controls the budget and makes the decisions about which government programs should live or die. Trump has impulses, whims, and passing fancies; Vought is methodical and determined to impose his rightwing views on the entire federal government. Every federal grant, Vought believes, should align with Trump’s anti-woke, anti-DEI agenda.
The White House is seeking to exert more control over billions of dollars in annual government grants, aiming to restrict a vast swath of funding — in health, housing, science and transportation — so that it primarily serves the purposes and organizations politically aligned with President Trump.
While the administration says that its primary goal is to safeguard taxpayer money, its proposal amounts to a major escalation in its attempt to reimagine the nation’s spending, even as Congress and the courts continue to rebuke the president for abusing such powers.
Mr. Trump’s ambitions were made clear in a roughly 400-page blueprint that was released to little fanfare on Friday. If finalized, it would require all federal grants to be approved by the president’s political appointees, who must ensure that the money would “demonstrably advance the president’s policy priorities.”
For the agencies that issue those awards and the nonprofit groups, local governments, universities and other entities that receive the money, the Trump administration would also impose a set of highly prescriptive and political criteria.
The government could not issue grants to projects or groups that “deny the biological reality of sex or the sex binary in humans,” for example. Nor could it seek to fund initiatives that “promote anti-American values,” contribute to illegal immigration, advance diversity, equity and inclusion or assist in voter registration.
The rules would further limit the ability of grant recipients to engage in some “issue advocacy.” Those that are funded would be scrutinized for their compliance with “religious liberty laws” and their “memberships and affiliations” with outside groups. And they could face the outright termination of their grants if the Trump administration someday determines that their actions are not in the “public interest.”
The restrictions echo the string of executive orders that Mr. Trump signed shortly after returning to office, many of which have been challenged or blocked in court. This time, however, the White House has pursued its restrictions by proposing a regulation, which is expected to become final after the government solicits public comment. The result could be applied far more broadly, and perhaps in ways that are harder to fight legally or undo later, according to budget experts.
The consequences could fall hardest on health and science, a field in which Mr. Trump has pursued some of the steepest cuts in his second term.
In exchange for federal assistance, researchers would face limits on the subjects that they can explore, the foreign labs with which they may collaborate and even the conferences at which they can appear. Dr. Georges C. Benjamin, the chief executive of the American Public Health Association, a professional organization and advocacy group, said the policy could “devastate innovation, science and research” in the United States.
Paul L. Thomas of Furman University has been a persistent critic of the narrative about the “Mississippi Miracle.” The story gained great traction when New York Times‘ columnist Nicholas Kristof took it national on September 1, 2023, in an article titled: “America Has a Reading Problem. Mississippi Has a Solution.” The “miracle” supposedly was accomplished without doing anything to improve the lives of children and their families, without even raising teachers’ salaries. The “science of reading” did the trick; that, plus holding back third graders who didn’t pass the final reading test.
Many articles have been written since then recycling the claim that the “science of reading” was largely responsible for the impressive growth in Mississippi’s fourth grade reading scores on NAEP (the National Assessment of Educational Progress), which is administered every two years. If only states forced teachers to teach the “science of reading,” there would be no failure in reading (except, of course, for the students who were retained in third grade and not participants in the fourth grade testing.)
The “Mississippi Miracle” allegedly occurred within the context of a “Southern Surge,” where low-spending, non-union states like Alabama and Louisiana also participated in a miraculous increase in reading scores. These professors complexified that claim recently.
“No story has caught the imagination of education reformers this decade quite like the ‘Mississippi miracle,’” Rachel Canter asserts in The Atlantic, adding:
Other states are now trying to emulate what Mississippi did. Those efforts largely revolve around adopting what’s known as the “science of reading”— a set of principles and teaching techniques, including phonics, that are grounded in decades of empirical research.
Canter, the Director of Education Policy at the Progressive Policy Institute, released as well a report on Mississippi reading and education reform, noting:
I personally spent 17 years helping state leaders run that race. As the head of Mississippi First, a nonprofit I founded in 2008, I played a hand in, and sometimes led, many of the state’s key education policy conversations with the legislature while also working with the Mississippi Department of Education to implement the reform agenda. This is my insider’s view of what policymakers, philanthropists, and pundits should know about what really happened.
Both Canter’s article and her report are lessons themselves in how education reform in the US works, specifically during this cycle driven by the “science of reading” and “science of learning.”
Notably, Canter mentions “empirical research,” yet neither a magazine article nor a think tank report meet the standards of “scientific” championed by “science of” reformers—experimental/quasi-experimental research published in peer-reviewed journals [1].
Also, Canter’s article introduces on a larger scale one of the many multiverses of the “science of reading” existing currently.
The article and report express what Mississippi officials have been arguing for a while: Mississippi reform is not a miracle; it is many years of hard and complex work.
Canter, in fact, seems to double-down on Mississippi reform is effective due to high-stakes accountability (the core of education reform since Reagan, reform that has never worked but perpetuated a permanent cycle of crisis and reform in the US).
I will return to Canter’s argument about Mississippi’s reform success, but I think the criticism of overly simplistic stories about the Mississippi “miracle” are valid and many are beginning to acknowledge that news articles and podcasts have driven reductive and misguided reading reform, policy, and classroom practice [2].
In short, a lesson we should learn, finally, is to reject “miracle” narratives in education.
Lessons Ignored (And Questions Unanswered)
The problem with Canter’s article and report (beyond that they lack experimental rigor) is that her claims are just as misleading and often just as incomplete as the media stories being sold.
One lesson ignored in the Mississippi story is that it suffers from “the moment” syndrome. I have been asking since the start of the “miracle” narrative: Why haven’t we looked at the historical increase in grade 4 NAEP reading scores, including an ignored spike well before the 2019 christening of “miracle”?:
A bigger lesson, however, is taking greater care when deciding if reforms work as well as what causes that success. Related, as well, is assuring that the data used to decide success or failure represents learning.
Here the Mississippi story is much different that the media “miracle” or Cantor’s argument that high-stakes accountability has worked in the state.
Several questions must be answered.
If Mississippi’s reform has worked, why does the state have the same wealth and race gaps as in 1998?
If Mississippi’s reform has worked, why does the state continue to retain about 9000 K-3 students per year?
And most significantly, if Mississippi reform has worked, do the test score increases in grade 4 represent greater student learning?
There is little scientific evidence on this important question, but the evidence is suggesting a principle by Gerald Bracey: “Rising test scores do not necessarily mean rising achievement.”
When grade 8 data are compared to grade 4, those analyses seem accurate since states behind Mississippi in grade 4 catch and pass by grade 8 (include the subgroup of Black students):
The irony here is that in 2019 when Hanford declared Mississippi reading reform a “miracle,” many uncritically jumped on that bandwagon.
The Atlantic article is receiving the same uncritical and effusive response—although it is no more credible.
Canter offers just a different compelling but ultimately misleading story.
As of 2026, there simply is no empirical evidence Mississippi’s reading reform has worked.
There remains no “science” in the multiverse of “science of reading” stories.
[1] One frustrating aspect of the “science of reading” movement has been the demand for “science” while advocates tend to use anecdotes, cherry pick evidence, and ignore research counter to their stories. Note the expectations, often ignored, for “scientific” by The Reading League:
If you want to help with the costs of keeping my public work open access and free, please DONATE.
Subscribe to Paul Thomas
Launched 4 months ago
P.L. Thomas, Professor of Education (Furman University, Greenville SC), is the poetry editor for English Journal. NCTE named Thomas the 2013 George Orwell Award winner. Follow his work @plthomasEdD.
So-called reformers continue to pursue a fantasy: they believe that changing the governance of public schools will lead to improvement in the education of children.
And so they advocate for mayoral control, state takeovers, charter schools, vouchers.
They choose to ignore the overwhelming consensus among education researchers that the home lives of children has a far greater impact on children’s school performance than the governing structure.
If you’re worried about corporations taking over public schools, this next sentence will not allay your worry: The state of Indiana just turned over much of the responsibilities for the city of Indianapolis’ schools–public and charter–to something called the Indianapolis Public Education Corporation (IPEC).
“This new organization will be charged with building and transportation management for both charter and traditional public schools,” reports Governing. “It will also be charged with creating a single set of evaluation criteria for both types of schools.”
Admittedly, it’s a nonprofit corporation and its nine board members are appointed by the mayor with statutes to determine the corporate board’s membership: three come from the Indianapolis Public Schools [IPS] board of commissioners–which still exists, three from the charter school industry, and three with administrative and financial expertise.
In reality, however, four board members are from the charter industry. Its board chairman is David Harris, who founded the Mind Trust-Indianapolis, the driving force pushing the charterization of the district. Harris is the President and CEO of Christal House International that operates the Christal House Academy charter chain. According to the organization’s 990 tax form, in 2025, Harris received $554,148 in compensation.
But IPEC inserts a layer of control and bureaucracy beyond–or, better put, around–Indianapolis’ elected school board. At least as troubling as that is the fact IPEC was given the authority to levy property taxes that it can use to fund–with public money–charter schools. This puts charter schools on equal footing–and funding–with public schools–a dangerous precedent that is certain to be attempted elsewhere.
One of IPEC’s first orders of business is likely to be placing on the November ballot an operating referendum since IPS’ expires at the end of this year. While Indianapolis Public Schools expects a $40 million deficit for the year, that might have been addressed if IPS’s attempt to place an operating referendum on the 2023 ballot hadn’t been derailed by the charter school industry and the Greater Indianapolis Chamber of Commerce.
For public school advocates, the implications of the new, non-elected board are clear–and disturbing.
“What is occurring in Indianapolis is part of a growing movement to destroy the neighborhood school governed by the community and replace it with a corporate vision of schooling that sees the marketplace and competition as the primary drivers of quality,” Carol Burris, executive director of the Network for Public Education, tells In the Public Interest. “We are now more than thirty years into the charter school experiment, and we have yet to see the miracle.”
Paul Thomas is a professor at Furman University. He has taken a leading role in refuting claims for the “science of reading.”
There are many successful ways to teach reading. some children arrive at school knowing how to read, because a parent read with them every day. Phonics is important. The joy of reading is important. Comprehension is important. Legislatures should not mandate one way to teach.
If you pay attention to the non-stop moral panicking around reading fanned by mainstream media, you may have seen this click-bait headline: Did New York blow $10 million on reading instruction that doesn’t work?
The article repeats tired and misleading (often false) stories about the failures of balanced literacy, NAEP reading scores, the “success” of Mississippi and Louisiana, the promise of structured literacy, and the National Reading Panel as well as the one research study on phonics that is linked.
Let’s consider these:
*No scientific studies identify a reading crisis in the US causally linked to balanced literacy, reading programs, lack of phonics instruction, or inadequacy of teacher education. (Aydarova, 2025; Reinking, Hruby, & Risko, 2023).
*The media hyper-focuses only on grade 4 NAEP reading scores, but notice how the story changes once we consider grade 8—states behind MS in grade 4 catch and pass MS by grade 8 (primarily because MS inflates their grade 4 scores by excessive grade retention, like FL):
*Structured literacy (scripted curriculum) is whitewashing the reading curriculum and restricting teacher autonomy and professionalism. (Khan, et al., 2022; Parsons, et al., 2025; Rigell, et al., 2022).
*The linked research suggests it replicates findings of the NRP; however, the NRP did not prove systematic phonics outperformed whole language or was a silver bullet. As Diane Stephens explains about the findings on phonics: “Minimal value in kindergarten; no conclusion about phonics beyond grade 1 for ‘normally developing readers’; systematic phonics instruction in grades 2-6 with struggling readers has a weak impact on reading text and spelling; systematic phonics instruction has a positive effect in grade 1 on reading (pronouncing) real and nonsense words but not comprehension; at-risk students benefit from whole language instruction, Reading Recovery, and direct instruction.” Further, while the article quotes from the research report, it doesn’t include this much more tentative hedge: “These findings suggest that SL approaches may yield larger positive effects on student learning compared to BL approaches.” At best, structured literacy is no better or worse than whole language or balanced literacy, but to be clear, there is no “settled science” that is works.
But the bigger problem is not that mainstream media continues to repeat misinformation, but that it fails to offer the full story.
Note that the “literacy experts” quoted in the article are supporting structured literacy programs (scripted curriculum), and some of those experts are co-authors of those programs.
Further, these experts are promoting a different teacher training program than the one being attacked in the article, and many states are spending 10s of millions of dollars on that program—LETRS. (My home state of SC a few years ago allocated $11 million for one year, for example.)
What’s missing in this story?
There are two high-quality studies that were released in 2025 on the effectiveness of LETRS, but so far, there have not been click-bait scare headlines about those findings:
*A review (Rowe & Thrailkill, 2025) of reading policy in North Carolina concludes:
Despite LETRS’ claim that it helps educators “distinguish between the research base for best practices and other competing ideas not supported by scientific evidence” (Lexia Learning, 2022, p. 4), we noticed a pattern of misinterpretation, selective inclusion, and omission of literacy research. LETRS is a prime example of a common problem with the deployment of research for educational policy and instructional decision-making, in that multiple claims are not substantiated by a close reading of the original research cited (cf. Hodge et al., 2020).
*And Gearin, et al. (2025) found:
[Abstract] We investigated whether Language Essentials for Teachers of Reading and Spelling: 3rd Edition (LETRS) was related to student reading ability by comparing the average third-grade reading achievement of schools that used LETRS to that of schools that used exclusively other professional development experiences in the context of Colorado’s Read Act. Guided by What Works Clearinghouse Standards, we conducted a quasi-experiment with propensity score matching and an active comparison group. We supplemented our primary intent-to-treat analysis with three sensitivity analyses designed to demonstrate the robustness of our claims. Effect estimates for completing a LETRS volume on educator knowledge ranged from 0.82 to 0.94. Students’ third-grade reading achievement did not statistically differ for schools that adopted LETRS compared with other professional development experiences in any model, suggesting that LETRS was comparable to the other programs at improving third-grade reading achievement at the school level.
The “science of reading” movement is a political, ideological, and market-based attack on teachers and public education, and the only people profiting off yet another moral panic are the media, political leaders, education reformers, entrepreneurs, and of course, the education market place.
The full story is never covered, because the real story about reading simply isn’t that profitable.
When an education policy is tried and failed, then tried again and continues to fail, that policy may justly beee called “zombie policy.” It survives despite experience..
Tom Ultican, retried teacher of physics and advanced mathematics in California, here describes such a policy. It is called “grade retention,” but is more commonly known as flunking a student because he or she is not “ready” to be promoted with peers. The short-term effect may seem successful: test scores. But the long-term effect on students’ success is typically negative.
Ultican writes:
Twenty-six American states have a mandatory third-grade retention policy for students who do not pass the state’s reading exam and Maryland is set to implement that policy in 2027. According to researchers, this is bad thinking based on intuition not science. Writing for Education Trust, Brittney Davis declared, “The research is clear that grade retention is not effective over time, and it is related to many negative academic, social, and emotional outcomes for students — especially students of color who have been retained.”
Economist Jiee Zhong won her PhD from Texas A&M in 2024 and is now an assistant professor of economics at the University of Miami. Last year, she just finished a very impressive study on the effects of grade retention for Texas third graders. Texas abandoned mandatory third-grade retention in 2009.
Zhong studied outcomes of third-graders from 2002-03, 2003-04 and 2004-05 school years who took the Texas reading exam that carried retention consequences. This large data set allowed her to use a fuzzy regression discontinuity design to extract many results. By 2024, the students studied were all young adults over 26 years of age. She was able to evaluate their education, social and economic outcomes using powerful math techniques.
Zhong concluded:
“I find that third-grade retention significantly reduces annual earnings at age 26 by $3,477 (19%). While temporarily improving test scores, retention increases absenteeism, violent behavior, and juvenile crime, and reduces the likelihood of high school graduation.”
For one outcome, she investigated a group of students who barely passed or barely failed the reading test. She learned that the barely failing students earn $1,682 (11.3%) less at age 23 than the barely passing students. Zhong noted that 64.2% of barely passing students graduated from high school while just 55.1% of the barely failing students graduated. She observed that both of these results were statistically significant at a 5% level.
Zhong also noticed a racial disparity. She reports, “White students experience a sharp 43.8 percentage point decline in high school graduation probability, higher than the reductions for Black (17.6 percentage points) and Hispanic students (0.6 percentage points).”
These results from 2025 add more weight to similar results that previous researchers have reported.
The Retention Illusion
In January 2025, Duke University in Chapel Hill, North Carolina published a linked series of three policy briefs concerning grade retention by Claire Xia and Elizabeth Glennie, Ph.D. The Duke researchers stated, “The majority of published studies and decades of research indicate that there is usually little to be gained, and much harm that may be done through retaining students in grade.”
They also mention the grade retention illusion is held by many community members, administrators and teachers who believe grade retention is helpful and needed. The Duke researchers stated, “The findings that retention is ineffective or even harmful in the long run seem counterintuitive.” This belief is so strong that on the 31st Annual Phi Delta Kappa/Gallop Poll, 72% of the public favor stricter promotion standards even if significantly more students would be held back. Other studies show the public being strongly opposed to social promotion believing low-achieving students will continue to fall farther behind.
The study found that students who had been consistently taught by teachers using “the science of reading” were gaining basic literacy skills, but were limited in their comprehension of what they read. They could read the words, but they couldn’t step back and explain what they had read.
Let’s back up for a few minutes and see this new study in historical perspective. The “science of reading” was based on the recommendations of the National Reading Panel. That panel was established by Congress in 1997 to determine the best, most effective ways to teach reading. Most of its 14 members were academics. In 2000, the panel released its report, callled Teaching Children to Read: An Evidence-Based Assessment. It recommended that effective reading instruction should include:
Phonics: Explicit, systematic instruction in letter-sound correspondences.
Fluency: Guided oral reading to encourage automaticity.
Vocabulary: Direct and indirect instruction of word meanings.
Comprehension: Teaching specific strategies for understanding text.
When George W. Bush became President in 2001, his education agenda featured the findings of the National Reading Panel. Dr. Reid Lyon, the organizer of the panel, became President Bush’s advisor. Bush’s No Child Left Behind legislation included $6 billion for reading instruction, based on the recommendations of the National Reading Panel, as well as an independent evaluation of its results.
Independent evaluators reviewed the progress of students in the districts that implemented the panel’s recommendations.
In 2008, they published their conclusions:
Reading First had a statistically significant positive impact on multiple practices promoted by the program, including the amount of instructional time spent on the five essential components of reading instruction (phonemic awareness, phonics, vocabulary, fluency, and comprehension) and professional development in scientifically based reading instruction.
Reading First did not produce a statistically significant impact on student reading comprehension test scores in grades one, two, or three.
Reading First had a statistically significant positive impact on first graders’ decoding skills in Spring 2007.
After the $30 million study, involving four major research organizations, reported that “the science of reading” improved decoding skills but not comprehension, enthusiasm for the NRP report waned.
But the NRP report found a second life less than a decade after it seemed to have faded.
Emily Hanford, a journalist who worked for American Public Media, began researching early literacy in 2016. Her 2022 podcast Sold a Story maintained that the source of poor literacy skills could be traced to the work of Marie Clay and Lucy Calkins, both of whom were advocates of balanced literacy, which did not incorporate the findings of the NRP.
Hanford became an advocate for “the science of reading” and the revival of phonics.
Many states enacted legislation mandating “the science of reading” and banning “three-cuing” and other elements of Calkins’ program.
“The science of reading” is unquestionably the dominant mode of teaching reading today.
Harkay wrote:
Four school districts in major urban areas using the science of reading found while students are grasping basic literacy skills, limitations toward deeper comprehension still exist, according to a new study.
The “Robust Reading Comprehension” report, conducted by nonprofit research organization SRI, examined literacy instruction in districts in Texas, Maryland, North Carolina and Virginia that have been using materials rooted in the popular phonics-based literacy approach for at least five years.
Through numerous classroom observations, teacher surveys and interviews with district officials in Aldine Independent School District, Baltimore City Public Schools, Guilford County Schools and Richmond Public Schools, researchers found a majority of reading lessons lacked “depth” – meaning foundational skills were mainly limited to working on single words rather than reading them in sentences.
Comprehension lessons in later elementary grades also mainly focused on completing a task, such as identifying a main character, rather than using a text for discussion and understanding its purpose.
“You’re not able to really think about the unpacking of a complicated sentence. You’re not thinking about really intentional vocabulary instruction or the building of kids’ word knowledge over time,” said Dan Reynolds, one of the lead authors of the report. “Ultimately, how should we be framing kids to read? Are we teaching our K-4 kids that reading is just tasks? Are we teaching them that they just need to label stuff and fill out graphic organizers?”
In recent years, nearly every state has passed science of reading laws, including many that have limited the type of programming and instructional materials a school can use – a move that has drawn some criticism that it’s too restrictive and that the instruction faces its own limitations.
The report defined surface literacy skills as a student’s ability to complete tasks and understand texts based on their literal meeting while robust instruction would further push a child to understand, evaluate and synthesize what they had read for its significance.
The study said its “comprehension observations alone are more rigorous than nearly all studies conducted in the last 50 years.” It’s not expected to be representative of reading instruction across the country, Reynolds said, but “we have four big districts in four different states, and we saw this pattern happening in all four of them with three different curricula.”
The study also found that teachers struggled with implementing comprehension-focused learning materials and said many times the curriculum was too dense, required substantial planning or may not have been developmentally appropriate. Professional development opportunities for these educators were also limited.
Researchers reported less than a quarter of observed comprehension lessons were engaging in robust learning. More than two-thirds of the lessons focused on “surface-level” comprehension.
“It seems that these curriculums are designed to build knowledge and they don’t develop meaning, and so then why read about the Civil War or about insects?” said Katrina Woodworth, director at SRI’s Center for Education Research & Improvement. “The point is to both teach reading and to build students’ knowledge base so that they have more scaffolding for future learning of both content and meaning.”
The SRI researchers also found that many review tools that measure comprehension don’t make a distinction between surface-level and robust instruction and skills. So, while educators are tasked with meeting a baseline standard, like having a child compare and contrast a text, it may be “unintentionally encouraging teachers to focus on surface-level goals,” the report said.
Without distinction, it weakens instruction for students and can later manifest as a skills disadvantage, Reynolds said.
“Districts had done so much to get the kids all the way there [with literacy], but it was losing voltage in the end,” Reynolds said. “If we can actually shift the way that districts are thinking about improving their comprehension instruction, they can take that all the way home and deliver really high quality comprehension instruction because so many pieces are already in place.”
Reynolds and one of his fellow co-authors, Sara Rutherford-Quach, said they saw glimpses of “magic” in the classroom when students understood a passage in wide-ranging contexts, which is the type of instruction they’re hoping to see districts incorporate more of in early grades.
“The kids were way more engaged,” Rutherford-Quach said. “Surface-level is important and necessary in some cases, … but it really is fundamentally different when you start talking about meaning and making it matter to the kids, and you see that they’re invested in it.”
Reynolds added that it’s unlikely robust comprehension could make up 100% of lessons in the classroom, but “we are thinking that if we can shift that needle from 24% robust lessons up to 50 or 60, then that would be a real catalyst for comprehension growth.”
The report recommended district leaders create “a shared vision for robust comprehension and define what it means for students, teachers, schools and the district,” and align how to best measure the extent of learning. It also called for better professional learning structures that could help model and rehearse robust comprehension work.
Previous reporting from The 74 found the percentage of recent high school graduates who lack “robust” comprehension skills is the highest it’s ever been, according to 2023 data. The sooner districts can engrain literacy skills that go beyond just explicit tasks, the easier it will be as they continue through the K-12 system, Reynolds said.
“I see the distinction between surface level and robust comprehension as critical to comprehension in fifth grade, but I also see it in the kids when they’re in 12th grade. Surface level comprehension and robust comprehension is the difference between a two on the AP exam and a three,” he said.
One evaluation in 2008. Another evaluation in 2026. Same conclusions. What have we learned?
The ultimate expert, Jeanne Chall, had it right. A former kindergarten teacher who became a renowned Harvard professor, she was commissioned by the Carnegie Corporation to review the research on reading. In her 1967 book, Learning to Read: The Great Debate, she concluded that the best approach was: both. Start early with phonics, she said, then transition to excellent children’s literature. If we continued to swing from extreme to extreme–from phonics to whole word, from whole word to phonics–she predicted, we would forever be trapped in that pendulum.
John Thompson, historian and retired teacher, reports on the latest education news from Oklahoma: the Chamber of Commerce is intent on reviving the failed test-and-punish agenda of the Bush-Obama years, plus the so-called “Mississippi Miracle,”which is credited with amazing results in reading.
John writes:
Once again, attacks on under-funded Oklahoma public schools are examples of the threats the nation’s schools face. Yes, we’ve gotten rid of State Superintendent Ryan Walters, but I’m more worried about today’s “accountability-driven” mandates, such as those pushed by the Chamber of Commerce.
On the other hand, our public schools have a history of receiving support from holistic, bottom-up efforts by a variety of excellent social work agencies, nonprofits, volunteers, and innovative educators.
These partners remind me of 1990’s, when student performance was growing. The head of the Oklahoma City Public School System curriculum department dropped into my History classroom, saying that she had been watching me teach, and I might like to try something new. She suggested that I start the year with the 20th century to get my kids hooked on history. Then, around Thanksgiving, we would return to the beginning of the subject, and reteach the 20th century.
It was a brilliant approach, supported by cognitive science. And it showed inner city students respect by nurturing meaningful, challenging instruction. The result was that my kids worked from bell-to-bell, from day one to their graduation day, learning how to learn.
I doubt that would be allowed today, when “everyone” is pressured to be on the “same page,” often requiring the same type of data-driven instruction.
Then, as the No Child Left Behind Act of 2001 approached, our principal gave us aligned and paced curriculum guidelines; They are now pervasive. She said that she knew we wouldn’t use it, but rather than throwing it away, we should keep it handy in case a top administrator visited the class.
Before NCLB, we had the autonomy to adjust our lessons in order to promote in-depth learning. For instance, when my students came to class carrying Ralph Ellison’s The Invisible Man, which they were reading in their English class, I would quickly change my schedule. Our History class would learn about Ellison’s childhood in Oklahoma City, and how his famous “Battle Royal” scene was inspired by a cruel joke that was played on him when he applied for a job.
And, around that time, the bipartisan MAPS for Kids succeeded in saving the OKCPS from a financial collapse by raising taxes. MAPS for Kids listened to educators, parents, top national education and cognitive science researchers, and students; it called for the meaningful instruction which treated high-challenge kids with the same respect and opportunities that are bestowed on students in the exurbs.
I was in the room when MAPS and OKCPS leaders agreed that educators should receive a clear message – their job is to teach the Standards of Instruction, not to standardized tests.
I was then in the room when top district administrators were supposed to reveal the agreement to a committee of principals. The committee chair started with summaries of ridiculous policies that had been imposed over the years. Principals replied with absurd, but hilarious stories, about the tumultuous effects of non-educators’ political demands.
But, the administrator then said that we would have to dramatically expand standardized testing.
When I pushed back, a highly respected administrator put her hands on my shoulders, and said, “John, I’ve always said you don’t make a hog heavier by weighing it. But this is politics. We have no choice.”
When NCLB and subsequent corporate school reforms were implemented, the supposed goal was using top-down, accountability mandates to rapidly transform schools serving our poorest children of color. But in my experience, those were the students who were most damaged by output-driven reforms that forced teachers to be “on the same page” when teaching the same lessons.
Reformers also brought frequent benchmark testing into schools. Lacking explicit stakes, benchmarks could have created a culture of testing for diagnostic, not accountability, purposes. In my experience, however, the test prep culture, combined with more frequent tests, further undermined the teacher autonomy required for holistic instruction.
Today, the campaign for the “Science of Reading,” now known as the “Mississippi Miracle,” is driven by “extensive use of formative and benchmark assessments to track student progress and inform instructional differentiation.” The American Federation of Teachers’ president Randi Weingarten supported much or most of the “Science of Reading” but she “doesn’t advocate for what we have found so disrespectful: scripted curricula or ‘teacher proof’ programs.”
And we face new threats when, as is happening in Oklahoma,” the ideology-driven, reward-and-punish parts of the “Mississippi Miracle,” are combined with the Moms for Liberty’s focus on “back to basics” foundational skills and phonics.
But, I would remind the Chamber of its call for recruiting and retaining high-quality teachers in order to attract and retain business investors for Oklahoma. After all, the best way to attract high-quality teachers, and parents of students, is to allow for high-quality, holistic teaching and learning, not make them work in a 21st century version of a Model T assembly line.