Archives for category: NAEP

The biggest lie about American school kids is that most are “below grade level.” This lie is repeated so often by prominent figures that it is widely believed. But it’s not true. Those who believe it are wrong. Those who repeat it, knowing it’s not true, are liars.

The source of the lie and the confusion is clear: the achievement levels in which NAEP scores are reported. The levels are “advanced,” “proficient,” “basic,” and “below basic.” When the media write about the latest release of NAEP scores, they frequently treat “proficient” as “grade level.”

But “proficient” is NOT “grade level.” It represents solid achievement, a rigorous aspirational goal. “Proficient” is equivalent to a solid A.

Every NAEP report on test scores says clearly in a footnote that “proficiency” is not the same as grade level. For example: “NAEP Proficient does not signify meeting grade-level expectations.” Yet the media and prominent commentators who should know better repeat the lie that most students are below grade level. The fact is that most students will never reach the high bar of “proficient.”

In 2023, as Bruce Lesley points out, Biden’s Secretary of Education–Miguel Cardona–testified to a Congressional committee that only one-third of American students were reading “at grade level.” I was flabbergasted. I couldn’t believe he said something so outrageous. I called Dr. Peggy Carr, who at that time was the Commissioner of Education Statistics. She was as surprised as I was that Secretary Cardona repeated the erroneous statistic. I asked Dr. Carr whether she had ever briefed him on understanding NAEP results; she had not.

I gave her an idea. Propose a change in name for “proficiency.” Change the name to “mastery.” No one would claim that “mastery” was the same as “grade level.” She liked the idea and promised to take it to the board. Whether she did, I don’t know. But nothing changed.

Bruce Lesley wrote this open letter to the National Assessment Governing Board, which oversees NAEP testing. Lesley is president of First Focus on Children and its partner organization First Focus Campaign for Children, bipartisan advocacy organizations dedicated to making children a priority in federal, state, and international policy. He has led both organizations since 2006 and 2009, respectively, building them into recognized national voices on child health, education, early childhood, economic security, budget and tax policy, immigration, children’s rights, and more recently, international child policy.

He wrote:

To the National Assessment Governing Board, the National Center for Education Statistics, and the leadership of the National Assessment of Educational Progress:

Every institution whose work affects children should begin with one question: “Is this good for children?”

By that standard, the National Assessment of Educational Progress (NAEP) has some important issues that deserve to be resolved. First and foremost, your achievement-level labels — “Basic,” “Proficient,” and “Advanced” — are being weaponized against the very children NAEP exists to serve, and you know it, because your own staff has been saying so for twenty-five years.

To be clear, this open letter is not a claim that NAEP’s underlying data is necessarily wrong, and it is not an argument against NAEP. The argument and request is narrower: you have a real and critically important ethical responsibility to correct the public misuse of your own data. NAEP should defend its credibility against those currently diminishing it.

This Week’s House Mark-Up Provides Another Example

On July 15, 2026, the House Education and Workforce Committee marked up a ten-bill package to facilitate the dismantling of the U.S. Department of Education.

In his opening statement, Chairman Tim Walberg (R-MI) argued that “too many children can’t read or do math at grade level,” and used that claim as a central justification for several of the bills. That claim is false.

Chairman Walberg was drawing on NAEP data — the statistic that roughly two-thirds of American fourth-graders do not score “Proficient” in reading, which is wrongly cited as evidence of failing to meet grade-level reading levels. For some, this is done out of confusion and, for others, to promote a political agenda to undermine public schools. In reality, NAEP proficient is aspirational and reflects a standard that is well above grade level.

Unfortunately, during the markup, multiple members of Congress repeated the same error. But again, NAEP Proficient is not grade level. It has never been grade level.

When the public, the press, the administration, and Congress repeatedly miscite this fact, the National Assessment Governing Board (NAGB) must do much more to clarify and correct misstatements about what it means.

Education expert Peter Greene explains:

The problem is two fold. One part of the problem is that “proficient” is used on many state and local assessments to mean “at grade level,” or what once upon a time would have been called a gentleman’s C; this leads to some honest confusion for some folks. The other part of the problem is folks who are invested in the narrative that public schools are failing and who benefit from the confusion surrounding the term.

Greene adds:

And every time NAEP scores are released, education journalists write piece after piece explaining “proficient” all over again, usually in the wake of some prominent person decrying the large number of students not “at grade level.”

That confusion is NAGB’s responsibility to address, and it has deserved attention for years, but all the more NOW.

This Is Not a Partisan Problem

Chairman Walberg and his colleagues’ misstatements are only the most recent officials to make this mistake (whether unintentionally out of confusion or internationally), and the pattern runs through both political parties.

Secretary Betsy DeVos, in the first Trump Administration, told the public that two-thirds of American students could not read at grade level— the same inaccurate conflation Chairman Walberg and his colleagues made yesterday.

Secretary Miguel Cardona, testifying before Congress in April 2023 under the Biden Administration, told lawmakers directly that only one-third of students were reading “on Grade level,” treating a NAEP proficiency figure as if it were a grade-level statistic, in nearly identical language.

Potential Democratic Party presidential candidate Rahm Emanuel is doing it as part of his tour of early primary states

And Secretary Linda McMahon, in the current Trump Administration, has used more careful wording — noting that nearly 70% of eighth graders are “not proficient” in reading — but has paired that technically accurate phrase with language implying total system failure. A Snopes piece by Rae Deng described this claim as lacking its own level of reading comprehension because, again, it completely mischaracterizes what NAEP’s “proficient” standard means.

Outside advocacy groups have been considerably less careful than any of them.

Moms for Liberty has publicly proclaimed that 68% of children cannot read at grade level, a direct misstatement of NAEP data. Here is just one of many examples. 

Furthermore, one of the organization’s co-founders has separately misread a state’s NAEP proficiency rate as that state’s overall literacy rate. Wrong again.

Corey DeAngelis, a leading advocate for school privatization, vouchers, and against public education, has cited NAEP proficiency figures directly, without qualification, as evidence that public schools are a system-wide “disgrace.”

Greene captures these types of political misuse of NAEP data in this Substack post.

Curmudgucation The Most Misused Statistics In Education.If someone is telling you that some extraordinary percentage of students can’t read at grade level, they’re probably wrong…Read more3 years ago · 2 likes · 1 comment · Peter Greene

This confusion is intentional by people arguing for both the dismantling of public education and federal investments in children.

Unfortunately, NAGB’s silence has allowed that rhetorical usefulness to go unchecked under Republican and Democratic administrations alike, and it is being used right now, this week, on Capitol Hill to justify eliminating the very agency that funds and safeguards the data NAGB produces.

NAGB’s Own Experts Have Been Saying This for Years

In 2001, Mary Lynne Bourque and Susan Loomis — a staff member and a board member of the National Assessment Governing Board itself — wrote plainly that the Proficient achievement level “does not refer to ‘at grade’ performance,” and that performance at Proficient is not the same as being “proficient” in a subject as any ordinary person would use that word.

Chester “Checker” Finn, Jr., who chaired the panel that adopted the achievement levels in 1992, has been candid that the levels were designed to be aspirational — a description of where students should ideally arrive, not a diagnosis of where most currently stand.

NCES itself has attached a caution to NAEP score reports for years: the Proficient level “does not represent grade level proficiency as determined by other assessment standards.”

If NAEP’s own architects and NAGB’s own website already say this, it is past time to be diligent in correcting the record when people misuse and misstate what it means. It is also on NAGB to stop publishing results in a format that predictably, foreseeably, and repeatedly gets misread as a verdict on grade-level performance, especially when you can see exactly how that misreading gets used again and again.

The clearest confirmation of all of this comes from NCES’s own data. Researchers Gina Cervetti and Kathleen Hinchman mapped every state’s definition of fourth-grade “grade-level” reading proficiency directly onto the NAEP scale and found that, as of the most recent analysis, nearly every state’s own standard for grade-level reading lines up with NAEP’s Basic level, not NAEP’s Proficient level. That means the honest translation of the data runs the opposite direction from how Chairman Walberg and others use it: by the states’ own definitions of grade level, roughly two-thirds of American fourth graders are reading at or above grade level, not below it.

Cervetti and Hinchman are also blunt about what actually is a crisis in the data: not a reading crisis, but an equity crisis. In 2022, only 48% of students eligible for free or reduced-price lunch scored at or above NAEP Basic, compared with 76% of students who were not eligible — a 28-point gap that has persisted, largely unchanged, for decades.

That is a story about generational wealth and unequal access to housing, healthcare, and school resources, not a story about failing classrooms, and NAEP’s own framing continues to let people tell the wrong story with your numbers.

What Education Writers and Researchers Have Been Saying

Diane Ravitch, who served seven years on the National Assessment Governing Board under President Clinton, has called out the confusion between NAEP Proficient and grade level as one of the most damaging and persistent falsehoods in American education discourse, noting that NAEP itself explicitly warns against the equivalence you continue to permit others to make.

Greene has argued that cut scores like “Proficient” function as scaled, curved judgments dressed up as fixed standards — noting that if every child scored above a cut, the establishment reaction would be to declare the cut too easy, not to celebrate the achievement. That is not how a genuine, fixed criterion is supposed to behave, and it is worth NAGB’s honest reckoning.

Mark Weber, a New Jersey teacher and education researcher, has done careful public work mapping state proficiency standards onto the NAEP scale, and his conclusion undercuts a favorite talking point of your critics-turned-allies in this fight: there is no empirical evidence that closing the so-called “honesty gap” between state and NAEP proficiency rates does anything to improve student achievement. If setting state cut scores to match yours were actually the lever for better outcomes, we would expect to see it in the data. We do not. That matters because it means the standard is being imported into state accountability systems on faith, not evidence — exactly the kind of unsupported claim NAGB should be correcting rather than allowing to spread.

The Brookings Institution’s Brown Center on Education Policy has been making this same case for nearly two decades. Tom Loveless, the Brown Center’s longtime director, authored a 2007 report concluding bluntly that NAEP’s cut scores were set too high. 

His 2016 Brookings piece, “The NAEP Proficiency Myth,” went further, noting that the achievement levels came under critical review from the U.S. Government Accountability Office, the National Academy of Sciences, and the National Academy of Education shortly after they were adopted — with the National Academy of Sciences review concluding the achievement levels were fundamentally flawed.

Loveless adds:

Advocates of the NAEP proficient standard want it to be for all students. That is ridiculous. Another way to think about it: proficient for today’s eighth graders reflects approximately what the average twelfth grader knew in mathematics in 1990. Someday the average eighth grader may be able to do that level of mathematics. But it won’t be soon, and it won’t be every student.

That is not a stray outside critique. That is respectable experts in the field, writing for decades, about the very categories NASB is still using today without correction.

One Point Should Not Separate “Failing” from “Successful”

NAGB also owes the public an honest accounting of what a cut score actually is. A cut score is a single point on a continuous scale, chosen somewhat arbitrarily by a panel, above which a child is declared “Proficient” and below which the same child, one point lower, is declared “Basic,” which is actually grade level.

Two children who are functionally indistinguishable in what they know and can do are sorted into entirely different public categories — one used as evidence that a school, a state, or a federal agency is failing, the other treated as evidence of success — because of a single point set by a committee, not because of any meaningful difference in the children themselves.

That is not a rounding error. It is the mechanism by which your data gets converted into political ammunition.

If NAGB cannot explain, in terms parents can understand, why the child who scores one point below the line is a different kind of learner than the child one point above it, then the line is doing rhetorical work the data was never built to support.

As the psychiatrist and educator William Glasser warned schools decades ago, chasing a point or two of movement on a test score is precisely the wrong institutional goal — and yet that is the goal NAEP’s cut scores hand every state, district, and school in the country by default.

Researcher Andrew Ho makes a similar point. He has identified proficiency cut scores as arbitrary markers, set through what he calls an “overwrought, judgmental, and ultimately political process,” not derived from any fixed line in human learning.

Ho has also documented a specific illusion that follows from that arbitrariness: because a large cluster of students always sits near the middle of the score distribution, a cut score placed close to that cluster will make small, ordinary shifts in performance look like dramatic gains or losses, purely as an artifact of how many students happen to sit right at the line — not because anything real changed in how much they learned. A researcher with no stake in the politics of this issue is describing the identical mechanism that turns your data into a rhetorical weapon: the closer the line sits to where children actually cluster, the more your data will appear to swing wildly for reasons that have nothing to do with children’s learning.

Criterion-Referenced in Name, Arbitrary in Practice

NAEP describes itself as a criterion-referenced assessment, distinct from norm-referenced tests like the SAT that simply rank students against one another. That distinction matters, and I want to represent it accurately rather than overstate it — NAEP does not “grade on a curve” in the way the SAT’s percentile scoring does.

However, the practical effect on families is not so different as the label suggests. NAEP’s cut scores were set by hand-picked panels making judgment calls about what students “should” know, not derived from an external, agreed-upon standard of competence, and independent evaluators — including a National Academies review in 2017 — have called for stronger evidence connecting NAEP performance levels to any real-world outcome at all.

A test that is criterion-referenced in name but whose criteria were set arbitrarily, and whose results still track family income and race as tightly as any norm-referenced test on the market, produces the same practical harm as the norming bias critics have long raised: it tells us more about a child’s zip code than about a fixed, meaningful standard of what that child knows.

Notably, NAGB has conceded the point this year. The 2026 NAEP reading framework — administered to students for the first time this spring — now explicitly disaggregates racial and ethnic subgroup results by socioeconomic status, on the premise, well documented for decades, that apparent racial differences in test scores largely track socioeconomic differences. That is a welcome and overdue acknowledgment.

But it is also, in effect, NAGB admitting in 2026 what critics have argued for years: that the results have been measuring wealth and family circumstance as much as they measure “proficiency,” all along. If that acknowledgment is real, it should extend backward, to how NAGB talks about every score ever published, not just forward, to a single new breakdown in the data tables.

The Test Itself Is Not Neutral

Even setting the cut scores aside, the content of the test carries its own bias, and NAEP’s own commissioned reviewers have said so. The NAEP Validity Studies Panel — a technical review body NCES itself created and funds — published an analysis by Gerunda Hughes in 2023 documenting that the statistical methods used to build NAEP-style test items can systematically disadvantage the very students the test is supposed to serve fairly.

When an item is answered correctly by nearly every student, it gets treated as a poor “discriminator” between high and low performers and is typically cut from the test in favor of harder items, even though that easy item may represent exactly the content that should be mastered.

This is not a hypothetical risk. Education researcher Wayne Au, in Unequal by Design: High Stakes Testing and the Standardization of Inequality, documents exactly how this mechanism has played out on the SAT, a test built using the same basic pretesting logic NAEP relies on.

In his book, Au cites researchers Kidder and Rosner, who examined more than 300,000 SAT test-takers and the pool of trial questions used to build future exams and found that some trial questions were answered correctly by Black students, or by Latino students, more often than by White students. Those questions were then discarded — not because they were poor measures of the content, but because they failed to reproduce the racial score gap the rest of the test already produced. A question only “counted” as valid if high-scoring test-takers, who are disproportionately White, tended to get it right in pretesting.

My mother has verified the same process when she was asked to be on a panel to evaluate whether the item questions were “fair”. The publishers of the Texas State assessment at the time ran through the questions and kept throwing out questions as biased toward Black or Hispanic children if they scored the same or close to the scores of White children

In contrast, questions in which there was a substantial gap in favor of White students were not flagged – thus, “norming” the disparity in test score outcomes into subsequent tests. Although my mother repeatedly objected, she was overruled throughout the day and, not surprisingly, never asked back to be a reviewer.

The result, as Au describes it, is a self-reinforcing loop: item selection is calibrated to match existing racial score gaps, which locks those same gaps into every future version of the test, all without anyone ever explicitly considering race in the selection criteria. 

NAEP is a different test administered by a different organization, and I am not asserting that NAEP’s item-selection process has been documented to work in the same way. But NAEP uses the same category of item statistics that made this outcome possible on the SAT, and NAGB’s own validity panel has already flagged the risk. Given what is now documented on a test as consequential as the SAT, NAGB owes the public a direct, public answer to a direct question: has anyone checked whether NAEP’s item-selection process does the same thing?

There is also cultural and geographic bias. As the son of an English teacher and a math teacher, it should be no surprise that I did fairly well on standardized tests throughout my life. But I vividly recall a reading passage from the PSAT that focused on nautical issues and the definition of a “flotilla.” 

Having grown up in El Paso, Texas, a city located hundreds of miles from any coastline, the passage and vocabulary word were unfamiliar to any of us taking the test in the desert borderlands. On the other hand, we would crush a passage referring to “tortillas.” NAEP’s own reviewers have a name for this: cultural validity, the idea that a test cannot cleanly separate what a child knows from what a child has been exposed to.

Research that NAEP’s own validity panel cites has found that when students are allowed to choose among reading passages on different topics, rather than being assigned a single passage that may be unfamiliar or uninteresting to them, some groups of students — including Black eighth graders and Hispanic twelfth graders in the panel’s own cited study — score much higher. That is evidence that some of what NAEP currently measures is exposure and familiarity, not just reading ability, and it argues for reform in how passages and vocabulary are chosen, not just in how results are labeled.

Again, the validity panel’s report contains proof that this is a design choice, not a fact of nature. In 1972, the psychologist Robert Williams built a test called the Black Intelligence Test of Cultural Hegemony, using vocabulary and content drawn from Black American culture instead of the dominant culture’s frame of reference. When Black and White teenagers took it, Black students substantially outscored White students by substantial margins.

Nothing about the underlying children changed between that test and the SAT. What changed was whose knowledge and cultural fluency the test happened to be built around.

That single fact should end, permanently, any claim that a test’s outcomes reveal some fixed truth about which children “can” or “cannot” read, think, or reason. What these tests reliably measure is often which cultural and economic frame of reference a child was raised in, and how well that frame matches the one test-makers chose to build around — which is another way of describing accumulated wealth, school funding, and generational inequity, not a verdict on a child’s mind.

That is real, and policies that address school finance inequity, child poverty, childhood hunger, and adverse childhood experiences (ACEs) deserve real policy attention. These issues would undoubtedly do more to improve educational outcomes in this country rather than privatization of public schools or the elimination of the Department of Education.

Claims that two-thirds of American children cannot read at grade level are simply false, and their interpretation by policymakers and advocates is harming children. There is an old warning that was popularized by author Mark Twain but attributable to British Prime Minister Benjamin Disraeli about three kinds of falsehood — “lies, damned lies, and statistics.”

In this case, even a true number, presented without its context, can mislead more effectively than an outright fabrication. NAEP’s “proficiency” level is an aspirational one, but the grade-level story built on top of it is doing real harm. NAGB is a position to explain the difference, and the public is not, until you tell them.

The Damage Is Not Abstract: What Gets Tested Is What Gets Taught

This is not a technical quibble.

Every time “below Proficient” gets reported to the public as “can’t read” or “can’t do math,” it becomes ammunition for defunding public schools and for portraying millions of children — disproportionately low-income children and children of color — as failures because of a label your board itself has said should not be read that way.

It also reshapes what happens inside the classroom. When reading and math scores on tests built around NAEP cut points become the metric by which schools, teachers, and even state superintendents are judged, instructional time follows the incentive:

  • Short, decontextualized passages crowd out real books — my children were taught how to write a brief constructive response (BCR) before they were even taught what a paragraph was.
  • Science, government, history, the arts, and physical education are pushed to the margins of the elementary school day because they are not tested and therefore not rewarded.

Children end up narrower, not better educated, in the very subjects that make them informed citizens — and NAEP’s own cut-score architecture is a direct contributor to that narrowing, whether or not that was your intent.

NAGB tried a partial fix in 2018, adding the word “NAEP” before each level — “NAEP Proficient” rather than “Proficient” — so people would stop equating your terms with generic ones.

James Harvey, executive director of the National Superintendents Roundtable, was right to call that gesture insufficient at the time. Harvey said:

…the American people should understand that the misleading term “proficient” sets a performance benchmark beyond the reach of most students in the world.

Harvey argued “proficient” should be changed to something like “high” to avoid being “fooled.”

His point has been proven many times, including this week when a sitting congressional committee chairman, citing NAEP-adjacent data to justify eliminating a federal agency, still used the word “grade level” as if it meant what NAEP’s Proficient level does not mean.

The Perverse Incentive NAEP Has Inspired: Grade Retention As Score Manipulation

The clearest evidence that NAEP’s cut scores create perverse incentives and “manufactured” crises, rather than honest information, is what states have started doing in response to them: holding back third-graders who miss an early-literacy cut score, in order to produce a fourth-grade NAEP cohort that looks better on paper.

Education professor and researcher Paul Thomas has documented this closely in states such as Mississippi, where fourth-grade reading gains celebrated as a “Mississippi miracle” tracked closely with a mandatory third-grade retention policy.

Paul ThomasCounter-Narratives: Mississippi Reading ReformEmily Hanford has profited from two very compelling stories…Read more2 days ago · 1 like · Paul Thomas

A child who is nine years old competing against classmates who are eight will predictably score higher on a test built around the same content; that is a fact about test administration, not about literacy. Furthermore, those same “gains” have been shown to fade by eighth grade, once the retained cohort catches up in age to its peers without having genuinely caught up in learning.

This is worth NAGB’s own honest reckoning, not because the research on retention is unanimous — reasonable analysts, including some closely tied to NAEP’s own governing board, dispute how much of Mississippi’s gain is genuine instructional improvement versus retention’s effect on cohort composition — but because NAEP’s achievement levels are the mechanism creating the incentive either way.

States are not retaining eight-year-olds because it is good for those children. They are retaining them because a single cut score on a single test has been elevated to a measure of whether a state’s education policy is working. The cost of that incentive falls on children: retained students who show a short-term score bump can, over time, experience the opposite of what was intended — greater disengagement, higher rates of dropping out before graduation, and the well-documented psychological toll of being told, at eight or nine years old, that they failed.

William Glasser spent much of his career, in Schools Without Failure, and later in The Quality School, explaining exactly why this backfires. He argued that standardized testing reduces learning to disconnected, memorized facts at the expense of critical thinking and real application — and that the “right answer, wrong answer” format of a multiple-choice test teaches children that education is a hunt for a single predetermined answer rather than a process of genuine understanding.

In the schools Glasser held up as models, closed-book tests were replaced with open-book, collaborative assessments that actually resembled the problems students would face outside school. His deeper claim, grounded in what he called Choice Theory, is that people — including children — are driven by needs for freedom, power, and simple enjoyment in their work, and that using test scores to rank, shame, or coerce students destroys the very motivation that produces quality work in the first place.

Labels matter. When children repeatedly hear that two-thirds of them “cannot read at grade level,” many internalize failure that is not supported by the evidence. Parents lose confidence in neighborhood schools. Teachers become demoralized. Policymakers propose increasingly radical structural changes to fix a crisis that has been inaccurately described.

Glasser also warned explicitly against making small, arbitrary numerical gains — his example was raising a test score by a point or two — the primary institutional goal of a school, insisting instead on building a genuine culture of quality. That is precisely the trap a single-point NAEP cut score sets for states, and it is the trap third-grade retention policies walk students directly into.

That is the opposite of what an assessment meant to serve children should produce, and it deserves your acknowledgment, not your silence.

What We Are Asking You To Do, Now

  1. Issue a direct, public correction each time a federal official misstates NAEP Proficient as “grade level,” the way you would correct any other material misuse of your data. Silence is not neutrality; it is acquiescence in the misuse.
  2. Publish, prominently and alongside every score release, the state-by-state mapping showing that “grade level” as states themselves define it corresponds to NAEP Basic, not NAEP Proficient — the analysis your own data already supports and that outside researchers have had to do on your behalf.
  3. Publish a plain-language document — “What NAEP Proficient Does, and Does Not, Mean” — and require it alongside every score release, every webpage, every press briefing, and every congressional testimony that cites NAEP data. Most of this letter’s argument could be prevented by a single page NAGB.
  4. Stop using “Basic,” “Proficient,” and “Advanced” as headline labels without their NAEP qualifier in every release, chart, and public statement — not as a footnote, but as a mandatory part of the label itself, displayed with the same prominence as the number.
  5. Retire “Basic,” “Proficient,” and “Advanced” altogether in favor of terms that do not already carry a plain-English meaning your data does not support. If the words themselves are the problem, changing a modifier in front of them has not been enough.
  6. Extend the honesty of the 2026 reading framework’s socioeconomic disaggregation backward, not just forward. If you now accept that racial gaps in your data are substantially explained by family socioeconomic status, say so plainly every time a racial achievement gap is reported, and stop letting that gap be cited as evidence of school failure without that context.
  7. Commission and publish the external validity evidence the National Academies asked for in 2017 — a transparent accounting of what your cut scores do and do not predict, so the public can evaluate the standard rather than take your word for its meaning.
  8. Publicly acknowledge the perverse incentive your cut scores have created for third-grade retention policies, and commission independent, longitudinal research — tracking students well past eighth grade, through high school graduation — before any state is permitted to point to NAEP gains as proof that retaining eight-year-olds is good policy.
  9. Act on your own validity panel’s 2023 findings, and answer the question the SAT evidence now raises. Publicly disclose whether NAEP’s item-selection process has ever been audited for the same self-reinforcing bias documented on the SAT — where trial questions that marginalized students answered correctly were discarded for failing to reproduce the existing score gap — and commit to an independent audit if it has not. Explain how tests are “normed” from one year to the next and made comparable in a manner that is understandable to the public.

Kids can’t wait for another year of this same correction being offered and ignored, or for another cohort of eight-year-olds to be held back so a state’s chart can look better. NAGB has the power to end the confusion your own board identified more than two decades ago. Please do so. Kids deserve it.

Paul Thomas taught in public high schools for many years, before becoming a professor at Furman College in South Carolina. He is a persistent critic of the “Mississippi Miracle.” He uses data to check on state claims. In this post, he fact-checks Florida.

He wrote:

Reading proficiency is a powerful data point despite it being a moving target.

When anyone refers to “reading proficiency,” that usually means a percentage of students who have met or exceeded an established score on a standardized test of reading.

However, “proficiency” is not a standard term. States tend to use “proficient” as grade level expectations while NAEP uses “proficient” as an aspirational achievement level (and “basic” more closely correlates with state grade-level proficiency).

To further complicate “reading proficiency,” not only does the measurement vary from state to state, but also the expectations for what percentage of students should be proficient at any grade is more a debate than an established fact.

How many students should be proficient in reading? Sometimes it is 90%sometimes it is 95%—and then there are state goals, for example, in Florida, as reported by Aldeman:

A 10-part video series produced by the Children’s Literacy Project tells what happened. It makes a compelling case that these results are attributable to a distinctive public-private partnership between the district and a nonprofit called The Learning Alliance. The story starts with two moms, Liz Woody-Remington and Barbara Hammond, whose children were struggling to read. In 2010, they asked themselves: What would it take to get 90% of the district’s children reading on grade level by the end of third grade?

I find these statistics troubling, similar to concerns raised by Hansford:

Over the years, I have on numerous occasions seen the claim that 95% of students can learn how to read proficiently, so long as they are provided adequate tier 1/2 instruction. Truthfully, it has always stuck out to me as a strange figure, for three reasons. First, most academic research does not typically use percentages in this sort of manner. Second, I often see this figure unaccompanied by a citation. And third, it seems low; I find it hard to believe that 5% of students just cannot learn how to read. …For this figure to have scientific validity, it would need experimental research demonstrating it to be true. Ideally, I would want to see multiple large scale studies, due to the universality of the claim. Intrigued by the discussion, I put out a public call on twitter asking if anyone had a citation for the figure.

Hansford walks us through the research (thin at best) and reaches an interesting conclusion:

This all said, it does seem there is some level of support for 96% being a benchmark goal, for reading proficiency rates. While some might argue, this is too high, I worry it’s too low, as it is clearly possible to achieve better than 96%. For example, in the Torgesen 2003 paper, 98.4% of students were able to read at grade level. When I asked for research on this topic, I was given an anecdote about a school using EBLI that went from 87% proficiency rates to 100%, within a matter of years. Well this is just an anecdote. I do think 100% proficiency is—in many cases—possible and should always be the goal.

I think the points here that must not be missed are the role of “anecdote” in claims about reading proficiency as well as claims about surprising gains and outlier “miracle” evidence, such as, again, Aldeman highlights:

Even more impressively, low-income third graders at Indian River schools scored better than the statewide average for all students. And, perhaps not surprisingly, when we went looking for high-poverty schools that were nevertheless getting good outcomes in reading, we identified three of the district’s schools — Rosewood Magnet, Fellsmere Elementary and Pelican Island Elementary — for our “Bright Spots” list. Fellsmere in particular stood out: Based on its 99% poverty rate, our calculations predicted that it would have a third grade reading rate of just 29%. But its actual rate was much higher, at 53%.

Indian River County was never exactly a failing district, but a decade ago it was performing a bit worse than the state as a whole. It has since begun to pull away, especially in third grade. Coming out of the pandemic, 60% of district third graders scored proficient in reading in 2023. That figure rose to 63% in 2024 and then jumped again, to 69%, in 2025.

This reporting fits into a “beating the odds”approach that frames outlier evidence as the normfor an entire population.

The evidence [1] is overwhelming in education that outlier “miracle” evidence is usually misleading or false, and even more problematic, outlier success, when valid, is rarely scalable.

In short, “beating the odds” stories make for compelling journalism and politics but not for effective or reasonable education reform.

These stories from Florida also raise some red flags.

The organization promoting this reform, Children’s Literacy Project, is faith-based.

Like other Republican states such as Oklahoma and Texas, Florida is seeking ways to erode the separation of church and state, specifically in public schools.

Schools partnering with organizations to promote and support reform is not necessarily a problem, but the outside help does create tensions about ideology as well as erodes the likelihood reforms are scalable.

Another few aspects of Florida are not highlighted in the reporting but deserve attention.

Returning to measurements of reading proficiency, Florida is in the bottom quartile of states in terms of the standard for “proficient”:

Florida, like Mississippi, is also a state where relative success in grade 4 reading quickly evaporates by grade 8:

Again like Mississippi, Florida is in the top of states for grade 4 reading on NAEP, but drops to the bottom quartile in grade 8:

https://radicalscholarship.com/wp-content/uploads/2025/06/image-6.png

Finally, the media and political story most often focuses on reforms in reading programs, teacher training, school leadership, and school expectations; however, outlier and surprising gains in grade 4 reading are likely driven by grade retention (a harmful punishment) and not the celebrated reforms.

Notably, high-grade retention states like Florida and Mississippi are also the states with significant decreases from grade 4 to grade 8.

Florida has a long history of aligning itself with “miracle” education reform that proves to be a mirage.

Beware the current numbers game about reading proficiency—a measurement that changes with the political wind.


[1] Thomas, P.L. (2016). Miracle schools or political scam? In W.J. Mathis & T.M. Trujillo, Learning from the Federal Market-Based Reforms: Lessons for ESSA. Charlotte, NC: IAP.


Atlanta Journal-Constitution

https://share.google/NXK2OD6xIFegOuWoe

By David Reinking and Peter Smagorinsky

Every day we read about people asking, “At what grade level does my child read?” “Is it true that 54% of adults in the U.S. read below a sixth-grade level?” “Have reading scores dropped an entire grade level since the pandemic?”

The assumption behind these questions is test scores are precise indicators of reading ability, like scientific laboratory measurements. But like blood pressure levels — in which there is agreement about what’s being measured — they are variable and open to interpretation. 

Despite the subjectivity and lack of agreement in defining reading, grade level and ability, grade-level reading ability is often mistakenly viewed as determined by a precise, stable test score, one that does not take into account factors such as students’ health and hunger in their testing performance. Not everyone agrees on what is salient at a particular grade level, leading to subjectivity in weighting phonics knowledge, vocabulary, comprehension, the ability to synthesize a theme and recognize an author’s point of view in a given passage, or some combination of such things. 

Standardized tests are typically the basis for establishing grade level. But test scores themselves don’t indicate grade level, which is a creation of an interpreter. That’s why different tests don’t always produce the same grade level. A student who tests at fourth grade in one state may test at the third or fifth grade when moving to another state using a different test. In short, different tests or standards can produce different grade levels.

David Reinking is a retired professor at Clemson and the University of Georgia. He is an inductee in the Reading Hall of Fame, and a former co-editor of Reading Research Quarterly and Journal of Literacy Research. (Courtesy)

David Reinking is a retired professor at Clemson and the University of Georgia. He is an inductee in the Reading Hall of Fame, and a former co-editor of Reading Research Quarterly and Journal of Literacy Research. (Courtesy)

The National Assessment of Educational Progress calls itself “the nation’s report card” even to the point of using the phrase on its website and then having it repeated as if it is an established fact. It is often invoked in commentaries on grade levels. But it wasn’t designed for that purpose. NAEP itself states that “NAEP Proficient achievement level does not represent grade level proficiency as determined by other assessment standards (e.g., state or district assessments). NAEP achievement levels are to be used on a trial basis and should be interpreted and used with caution.”

But that hasn’t stopped many policy makers and journalists from trying to connect a NAEP test score to a grade level. Giving in to political pressure and rejecting recommendations from authorities in developing tests, in 1990 NAEP officials did introduce four tiers of reading achievement: Below Basic, Basic, Proficient and Advanced. These levels were established solely using subjective judgment about what’s expected of children tested at a specific grade level.

Peter Smagorinsky is a retired professor at UGA, an inductee in the Reading Hall of Fame, and a former co-editor of Research in the Teaching of English. (Courtesy)

Peter Smagorinsky is a retired professor at UGA, an inductee in the Reading Hall of Fame, and a former co-editor of Research in the Teaching of English. (Courtesy)

Then, they set equally subjective cut scores to establish boundaries between these four levels. States often use a parallel model for their own tests. In Virginia, student performance is measured on a 0–600 scale, with proficiency set at 400-499 and advanced at 500 or above. It’s hard to imagine that a meaningful difference exists between a student scoring 499 (proficient) and 500 (advanced).

Much confusion is also centered in interpreting whether “basic” is acceptably normal or if it is reasonable to expect all students to be “proficient.” Many commentators, some of whom have a vested interest in arguing that there is a reading crisis, argue the latter. Some have promoted the false idea that “proficient” is grade-level reading, which it absolutely is not. Then, they wrongly argue that two-thirds of American students are reading below grade level by counting “basic” scores as below grade level.

Another way to illustrate the problem is to simply rename NAEP’s subjective categories as “below average,” “average,” “above average” and “far above average.” Then, approximately 60% of students are reading at or above an average score, and only 40% (instead of the usually expected 50%) of students are below average. Presto, much of the reading crisis disappears. As further evidence against a crisis, there has been relatively little variation in NAEP reading scores since 1992, even if an upward trend began retreating around 2015 with many plausible but unconfirmed explanations.

A number of educators have debunked the conclusions of NAEP misinterpreters. Yet, the dogged belief persists that everything can be reduced to subjective interpretations of test scores divided into hierarchical categories that can be falsely, if conveniently, converted to grade levels. 

We are concerned whenever we encounter all-too-common misinformation about grade-level reading ability. When misinformation becomes disinformation offered by those who use grade-level reading ability to advance political, polemical or ideological agendas, we become concerned about how faith in test scores lends them to manipulation and deception to help create the crisis that critics have historically claimed is engulfing schools, only to be saved by their favorite solutions. 


David Reinking is a retired professor at Clemson and the University of Georgia, an inductee in the Reading Hall of Fame, and a former co-editor of Reading Research Quarterly and Journal of Literacy Research. Peter Smagorinsky is a retired professor at UGA, an inductee in the Reading Hall of Fame, and a former co-editor of Research in the Teaching of English.

Paul L. Thomas of Furman University has been a persistent critic of the narrative about the “Mississippi Miracle.” The story gained great traction when New York Times‘ columnist Nicholas Kristof took it national on September 1, 2023, in an article titled: “America Has a Reading Problem. Mississippi Has a Solution.” The “miracle” supposedly was accomplished without doing anything to improve the lives of children and their families, without even raising teachers’ salaries. The “science of reading” did the trick; that, plus holding back third graders who didn’t pass the final reading test.

Many articles have been written since then recycling the claim that the “science of reading” was largely responsible for the impressive growth in Mississippi’s fourth grade reading scores on NAEP (the National Assessment of Educational Progress), which is administered every two years. If only states forced teachers to teach the “science of reading,” there would be no failure in reading (except, of course, for the students who were retained in third grade and not participants in the fourth grade testing.)

The “Mississippi Miracle” allegedly occurred within the context of a “Southern Surge,” where low-spending, non-union states like Alabama and Louisiana also participated in a miraculous increase in reading scores. These professors complexified that claim recently.

The most recent article confirming the “miracle” appeared in The Atlantic and was written by Rachel Canter, who participated in the Mississsippi reforms as leader of Mississippi First and is now at the Progressive Policy Institute in Washington, D.C.

Paul Thomas writes on his Substack blog:

“No story has caught the imagination of education reformers this decade quite like the ‘Mississippi miracle,’” Rachel Canter asserts in The Atlantic, adding:

Other states are now trying to emulate what Mississippi did. Those efforts largely revolve around adopting what’s known as the “science of reading”— a set of principles and teaching techniques, including phonics, that are grounded in decades of empirical research.

Canter, the Director of Education Policy at the Progressive Policy Institute, released as well a report on Mississippi reading and education reform, noting:

I personally spent 17 years helping state leaders run that race. As the head of Mississippi First, a nonprofit I founded in 2008, I played a hand in, and sometimes led, many of the state’s key education policy conversations with the legislature while also working with the Mississippi Department of Education to implement the reform agenda. This is my insider’s view of what policymakers, philanthropists, and pundits should know about what really happened.

Both Canter’s article and her report are lessons themselves in how education reform in the US works, specifically during this cycle driven by the “science of reading” and “science of learning.”

Notably, Canter mentions “empirical research,” yet neither a magazine article nor a think tank report meet the standards of “scientific” championed by “science of” reformers—experimental/quasi-experimental research published in peer-reviewed journals [1].

Also, Canter’s article introduces on a larger scale one of the many multiverses of the “science of reading” existing currently.

The article and report express what Mississippi officials have been arguing for a while: Mississippi reform is not a miracle; it is many years of hard and complex work.

Canter, in fact, seems to double-down on Mississippi reform is effective due to high-stakes accountability (the core of education reform since Reagan, reform that has never worked but perpetuated a permanent cycle of crisis and reform in the US).

I will return to Canter’s argument about Mississippi’s reform success, but I think the criticism of overly simplistic stories about the Mississippi “miracle” are valid and many are beginning to acknowledge that news articles and podcasts have driven reductive and misguided reading reform, policy, and classroom practice [2].

In short, a lesson we should learn, finally, is to reject “miracle” narratives in education. 

Lessons Ignored (And Questions Unanswered)

The problem with Canter’s article and report (beyond that they lack experimental rigor) is that her claims are just as misleading and often just as incomplete as the media stories being sold.

One lesson ignored in the Mississippi story is that it suffers from “the moment” syndrome. I have been asking since the start of the “miracle” narrative: Why haven’t we looked at the historical increase in grade 4 NAEP reading scores, including an ignored spike well before the 2019 christening of “miracle”?:

A bigger lesson, however, is taking greater care when deciding if reforms work as well as what causes that success. Related, as well, is assuring that the data used to decide success or failure represents learning.

Here the Mississippi story is much different that the media “miracle” or Cantor’s argument that high-stakes accountability has worked in the state.

Several questions must be answered.

If Mississippi’s reform has worked, why does the state have the same wealth and race gaps as in 1998?

If Mississippi’s reform has worked, why does the state continue to retain about 9000 K-3 students per year?

  • 2014-2015 – 3064 (grade 3) – 12,224 K-3 retained/ 32.2% proficiency
  • 2015-2016 – 2307 (grade 3) – 11,310 K-3 retained/ 32.3% proficiency
  • 2016-2017 – 1505 (grade 3) – 9834 K-3 retained / 36.1 % proficiency
  • 2017-2018 – 1285 (grade 3) – 8902 K-3 retained / 44.7% proficiency
  • 2018-2019 – 3379 (grade 3) – 11,034 K-3 retained / 48.3% proficiency
  • 2021-2022 – 2958 (grade 3) – 10,388 K-3 retained / 46.4% proficiency
  • 2022-2023 – 2287 (grade 3) – 9,525 K-3 retained/ 51.6% proficiency
  • 2023-2024 – 2033 (grade 3) – 9,121 K-3 retained/ 57.7% proficiency
  • 2024-2025 – 2132 (grade 3) – 9250 K-3 retained/ 49.4% proficiency

And most significantly, if Mississippi reform has worked, do the test score increases in grade 4 represent greater student learning?

There is little scientific evidence on this important question, but the evidence is suggesting a principle by Gerald Bracey: “Rising test scores do not necessarily mean rising achievement.”

First, an analysis of reading reform and a statistical analysis of Mississippi test score increases suggest that those increases are statistical manipulations caused by grade retention and not student learning.

When grade 8 data are compared to grade 4, those analyses seem accurate since states behind Mississippi in grade 4 catch and pass by grade 8 (include the subgroup of Black students):

The irony here is that in 2019 when Hanford declared Mississippi reading reform a “miracle,” many uncritically jumped on that bandwagon.

The Atlantic article is receiving the same uncritical and effusive response—although it is no more credible.

Canter offers just a different compelling but ultimately misleading story.

As of 2026, there simply is no empirical evidence Mississippi’s reading reform has worked.

There remains no “science” in the multiverse of “science of reading” stories.


[1] One frustrating aspect of the “science of reading” movement has been the demand for “science” while advocates tend to use anecdotes, cherry pick evidence, and ignore research counter to their stories. Note the expectations, often ignored, for “scientific” by The Reading League:

https://radicalscholarship.com/wp-content/uploads/2022/08/scientifically-based-research.jpg

[2] I have four open-access articles in English Journal, documenting with research that the media stories (specifically by Emily Hanford) are misleading and inaccurate.


If you want to help with the costs of keeping my public work open access and free, please DONATE.


Subscribe to Paul Thomas

Launched 4 months ago

P.L. Thomas, Professor of Education (Furman University, Greenville SC), is the poetry editor for English Journal. NCTE named Thomas the 2013 George Orwell Award winner. Follow his work @plthomasEdD.

John Thompson, retired teacher and historian in Oklahoma, considers ideas about how to improve Oklahoma’s schools, but insists that one overlooked cause of lower academic progress, was the torrent of misguided mandates written in Washington, D.C., such as No Child Left Behind and Race to the Top.

Thompson writes:

Despite our disagreements on some policies and research methodologies, I have respect for Adam Tyner, the executive director of the Oklahoma Center for Education Policy  He earned a doctorate in Political Science, and was the National Research Director at the Thomas B. Fordham Institute.Tyner is the author of The Fall to 48th: Documenting Oklahoma’s Educational Decline, which draws upon NAEP scores, and cites Diane Ravitch as to their reliablity. While I agree that Oklahoma schools can come back, I’m troubled by the title of his NonDoc piece, “The ‘Southern Surge’ suggests Oklahoma’s education system can bounce back.” 

Being a retired inner-city teacher, I am pleased by Tyner’s rejection of cheap, quick, and simple solutions. But, as a historian, I would focus on different NAEP test scores, and the way that No Child Left Behind (NCLB); Race to the Top (RttT); and budget cuts undermined teaching and learning.

To his credit, Tyner linked to Matt Barnum’s analysis of both the potential benefits and harms of the “Southern Surge,” and the “Mississippi Miracle.” Barnum acknowledged the gains in 4th grade test scores by states that drew upon the “Science of Reading.” But, he concluded:

Eighth graders’ results “have been less impressive for these Southern exemplars.” Alabama’s eighth grade reading scores have been falling and are among the lowest in the country. Louisiana’s eight grade reading scores remain at the 2002 level. And, Mississippi’s eighth grade reading scores are about the same as they were in 1998.

I believe that Tyner’s history of the last three decades should be read in conjunction of his recent commentary in the Oklahoman. 
He starts it with Phonics instruction being “a first step towards teaching literacy.” But, he adds, “Background knowledge is key to reading comprehension.”

Tyner then explains:

To become a strong reader in middle school and beyond, students need a firm foundation of core knowledge, and that comes not just from practicing reading, but from developing a broad vocabulary and an understanding of a large range of topics — from geography and history to literature and science.

He then critiques many Oklahoma schools for efforts to improve comprehension by mainly:

Having students practice so-called “comprehension skills and strategies,” such as finding the main idea in a passage and making inferences. These Chromebook-based exercises often resemble test prep. Although some of this practice is fine, hours spent on it crowd out history, geography, science and literature.

This is very consistent with a scholarly paper by the SRI, Report: Beyond the Surface: Leveraging High-Quality Instructional Materials for Robust Reading Comprehension Learning brief, funded by Tulsa’s Schusterman Family Foundation. As reported by the 74, Katrina Woodworth, the director at SRI’s Center for Education Research & Improvement, explained. “The point is to both teach reading and to build students’ knowledge base so that they have more scaffolding for future learning of both content and meaning.” But even the most promising Science of Reading programs they studied, may be “unintentionally encouraging teachers to focus on surface-level goals.”

One of the lead authors, Dan Reynolds, asked, “Are we teaching our K-4 kids that reading is just tasks? Are we teaching them that they just need to label stuff and fill out graphic organizers?”

Reynolds said the “Surface-level” instruction they discovered, “weakens instruction for students and can later manifest as a skills disadvantage.” 

And, getting back to Tyner, he wrote that an “important caveat to the undeniable successes of Mississippi and Louisiana in raising fourth-grade reading is that those states have seen little improvement in eighth-grade reading.”

While I very much agree with his position on the harm done by the failure to focus on background information, educators didn’t voluntarily undermine the teaching of history, the arts, and critical thinking. After all, the SRI study finds hope in the evidence that students and teachers prefer deep reading instruction.

But, I wish he had explained how the decline of holistic instruction was the predictable result of the NCLB’s and RttT’s test-driven mandates. During that time, for example, I served on a team assembled by our outstanding State Superintendent Sandy Garrett, in order to minimize the harm we knew was coming with NCLB.

Due to the demand that schools meet impossible testing goals, schools were forced to cut back on social studies, history, science, and the arts, as well as critical thinking. They inflicted the worst harm on schools serving the poorest children of color. Being a history teacher in extremely high-challenge high schools, I was horrified by the hundreds of stories I was told by students who said they were “robbed of an education.”

And those experiences explain why I’m worried by Tyner’s call for “deliberate efforts to improve instruction and accountability.” I would communicate with many thousands of teachers, and students, and I can’t remember anyone who lived through those “reforms” and didn’t see test-driven, accountability-driven instruction as a failure.

Moreover, while Tyner calls for solid funding of the infrastructure necessary to implement the Southern Strategy, he is less clear about the harms that retaining students can have. Given the lies perpetrated by rightwingers who claimed Oklahoma failed to improve reading because Joy Hofmeister quickly ended retentions, I wish he would be more explicit in fact-checking them.  

A history of 21st century education in Oklahoma should also explicitly include the reasons why Oklahoma backed off from passing four End of Instruction tests. Rep. Joe Eddins explained in 2005, “Based on test data, the House of Representatives staff estimates 89,000 failed tests each year.”

So, Oklahomans focused on win-win policies, and NAEP 8th grade test scores, stopped declining in 2005, and went up from 2009 to 2013.  (2013 was the year when national 8th grade reading and math scores also peaked.) 

I taught in an alternative school, in 2012, when new End-of-Instruction tests were being piloted. I resigned after being required to give the vast majority of my students’ worksheets, and focus on tutoring a few students who had a chance of passing the test, and graduate. Fortunately, under the leadership of Superintendent Joy Hofmeister, that law was repealed in 2016.

A history of what went wrong in Oklahoma schools should also address the budget cuts that killed those successes.

As the Oklahoma Policy Institute reported in 2016:

Oklahoma’s per pupil funding of the state aid formula for public schools has fallen 26.9 percent after inflation between FY 2008 and FY 2017. These continue to be the deepest cuts in the nation, and Oklahoma’s lead is growing. On a percentage basis, we’ve cut nearly twice as much as the next worst state, Alabama.

Moreover, Mississippi’s cuts ( -9.2) were about a third of Oklahoma’s, and Florida’s and Louisiana’s cuts were a little less than 20% and about 10%. Tennessee increased its funding by 9.8%.

After Nearly a Decade, School Investments Still Way Down in Some StatesPublic investment in K-12 schools — crucial for communities to thrive and the U.S. economy to offer broad opport…

Although I would have written a different history on Oklahoma education’s decline, I do believe we can rebuild our education systems.

But, I would have liked to read more of Tyner’s thoughts about the damage teachers witnessed by accountability-driven reforms that were imposed on Oklahoma schools, and huge funding cuts. My main response to his history, however, is that this is the time to be more blunt in terms of what it would  really take to achieve equitable levels of reading for comprehension.  

Given the lack of evidence that the “Southern Surge” is improving reading comprehension, providing long-term benefits, and doing more good than harm, we should find a more holistic way to reverse the harm inflicted on our schools by top-down mandates of the last quarter of a century. 

Paul L. Thomas was a high school teacher in South Carolina for nearly twenty years, then became an English professor at Furman University, a small liberal arts college in South Carolina. He is a clear thinker and a straight talker.

He wrote this article for The Washington Post. He tackles one of my pet peeves: the misuse and abuse of NAEP proficiency levels. Politicians and pundits like to use NAEP “proficiency” to mean”grade level.” There is always a “crisis” because most students do not score “proficient.” Of course not! NAEP proficient is not grade level! NAEP publications warn readers not to make that error. NAEP proficient is equivalent to an A. If most students were rated that high, the media would complain that the tests were too easy. NAEP Basic is akin to grade level.

He writes:

After her controversial appointment, U.S. Education Secretary Linda McMahon posted this apparently uncontroversial claim on social media: “When 70% of 8th graders in the U.S. can’t read proficiently, it’s not the students who are failing — it’s the education system that’s failing them.”

Americans are used to hearing about the nation’s reading crisis. In 2018, journalist Emily Hanford popularized the current “crisis” in her article “Hard Words,” writing, “More than 60 percent of American fourth-graders are not proficient readers, according to the National Assessment of Educational Progress, and it’s been that way since testing began in the 1990s.”

Five years later, New York Times columnist Nicholas Kristof repeated that statistic: “One of the most bearish statistics for the future of the United States is this: Two-thirds of fourth graders in the United States are not proficient in reading.”

Each of these statements about student reading achievement, though probably well-meaning, is misleading if not outright false. There is no reading crisis in the U.S. But there are major discrepancies between how the federal government and states define reading proficiency.

At the center of this confusion is the National Assessment of Educational Progress, a congressionally mandated assessment of student performance known also as the “nation’s report card.” The NAEP has three achievement levels: “basic,” “proficient” and “advanced.”

The disconnect lies with the second benchmark, “proficient.” According to the NAEP, students performing “at or above the NAEP Proficient level … demonstrate solid academic performance and competency over challenging subject matter.” But this statement includes a significant clarification: “The NAEP Proficient achievement level does not represent grade level proficiency as determined by other assessment standards (e.g., state or district assessments).”

In almost every state, “grade level” proficiency on state testing correlates with the NAEP’s “basic” level; in 2022, 45 states set their standard for reading proficiency in the NAEP’s “basic” range. Therefore, it is inaccurate to say that nearly two-thirds of fourth-graders are not capable readers.

The NAEP has been a key mechanism for holding states accountable for student achievement for over 30 years. Yet, educators have expressed doubt over the assessment’s utility. In 2004, an analysis by the American Federation of Teachers raised concerns about the NAEP’s achievement levels: “The proficient level on NAEP for grade 4 and 8 reading is set at almost the 70th percentile,” the union wrote. “It would not be unreasonable to think that the proficiency levels on NAEP represent a standard of achievement that is more commonly associated with fairly advanced students.”

The NAEP has set unrealistic goals for student achievement, fueling alarm about a reading crisis in the United States that is overblown. The common misreading of NAEP data has allowed the country to ignore what is urgent: addressing the opportunity gap that negatively impacts Black and Brown students, impoverished students, multilingual learners, and students with disabilities.

To redirect our focus to these vulnerable populations, the departments of education at both the federal and state levels should adopt a unified set of achievement terms among the NAEP and state-level testing. For over three decades, one-third of students have been below NAEP “basic” — a figure that is concerning but does not constitute a widespread reading crisis. The government’s challenge will be to provide clearer data — instead of hyperbolic rhetoric — to determine a reasonable threshold for grade-level proficiency.

What’s more, federal and state governments should consider redesigning achievement terms altogether. Identifying strengths and weaknesses in student reading would be better served by achievement levels determined by age, such as “below age level,” “age level” and “above age level.”

Age-level proficiency might be more accurate for policy and classroom instruction. As an example, we can look to Britain, where phonics instruction has been policy since 2006. Annual phonics assessments show score increases by birth month, suggesting the key role of age development in reading achievement.

In the United States, only the NAEP Long-Term Trend Assessment is age-based. Testing by age avoids having the sample of students corrupted by harmful policies such as grade retention, which removes the lowest-performing students from the test pool and then reintroduces them when they are older. Grade retention is punitive: It is disproportionately applied to students of color, students in poverty, multilingual learners and students with disabilities — the exact students most likely to struggle as readers.

Some evidence suggests that grade retention correlates with higher test scores. In a study of U.S. reading policy, education researchers John Westall and Amy Cummings concluded states that mandated third-grade retention based on state testing saw increases in reading scores.

However, the pair acknowledge that these were short-term benefits: For example, third-grade retention states such as Mississippi and Florida had exceptional NAEP reading scores among fourth-graders but scores fell back into the bottom 25 percent of all states among eighth-graders.

The researchers also caution that the available data does not prove whether test score increases are the result of grade retention or other state-sponsored learning interventions, such as high-dosage tutoring. Without stronger evidence, states might be tempted to trade higher test scores for punishing vulnerable students, all without permanent improvement in reading proficiency.

Hyperbole about a reading crisis ultimately fails the students who need education policy grounded in more credible evidence. Reforming achievement levels nationwide might be one step toward a more accurate and useful story about reading proficiency.

The article has many links. Rather than copying each one by hand, tedious process, I invite you to open the link and read the article.

As I was writing up this article, Mike Petrilli sent me the following graph from the 2024 NAEP. There was a decline in the scores of White, Black, and Hispanic fourth grade students “above basic.”

70% of White fourth-graders scored at or above grade level.

About 48% of Hispanics did.

About 43% of Blacks did.

The decline started before the pandemic. Was it the Common Core? Social media? Something else?

Should we be concerned? Yes. Should we use “crisis” language? What should we do?

Reduce class sizes so teachers can give more time to students who need it.

Do what is necessary to raise the prestige of the teaching profession: higher salaries, greater autonomy in the classroom. Legislators should stop telling teachers how to teach, stop assigning them grades, stop micromanaging the classroom.

Oklahoma’s State Superintendent, Ryan Walters, changed last years’ testing cut scores, redefining the term “proficient” in the state’s accountability data. Fortunately, there has been a bipartisan backlash against Walters’ lack of transparency when making the change, which looked like an effort to trick Oklahomans into believing that he had improved student outcomes.

But, this month, the Oklahoma Commission for Educational Quality and Accountability brought back a misleading, inappropriate, and destructive definition of the term proficiency for accountability purposes.

In doing so, the Commission revitalized the use of one of the most effective weapons for privatizing public education. They perpetuated the lie that “proficiency” is “grade level,” thus making it sound like public schools are irrevocably broken. 

We need to remember the history of this propaganda which took off during the Reagan Administration, which misused data in its “A Nation at Risk” to push high-stakes testing.

The National Assessment of Educational Progress (NAEP) scores are the best estimate of students’ outcomes, but they should be used for diagnostic, not accountability purposes.   But, as the Tulsa World reported, in 2011, Jeb Bush’s Foundation for Excellence in Education (FEE) high-jacked NAEP’s terminology when writing and editing then State Superintendent Janice Barresi’s new accountability-driven A-F school report card. The World presented evidence that the FEE was engaged in a “pay-to-play” scheme to reap profits while influencing policy.

As The Washington Post reported in 2013, FEE was at the nexus of rightwing political influence in K-12 education and corporate interests seeking to profit from the nation’s schools. It claimed that raising “expectations” for students would advance their learning. In fact, NAEP scores provide evidence that starting in 2012 , when corporate reforms were in place, the opposite happened, as NAEP scores declined, reversing decades of incremental growth.

It did, however, advance the privatization of public education.

At the 2024 Oklahoma conferenceBush’s new think tank, ExcelinEd used misleading and misconstrued data from the National Assessment of Educational Progress (NAEP), to conflate NAEP “proficiency” with “grade level.”

In fact, as Oklahoma Watch’s Jennifer Palmer explained, Oklahoma’s 8th grade reading proficiency grade requires that “students demonstrate mastery over even the most challenging grade-level content and are ready for the next grade, course or level of education.” That definition of mastery of grade level skills included critical thinking, interpretation, evaluation, analysis, and synthesis when reading across multiple texts, and writing.

But, Palmer noted, “8th graders who didn’t score proficient, but are in the ‘basic’ category, can still do all this.”

Moreover, as Jan Resseger further explained, the nation’s NAEP proficiency grade “represents A level work, at worst an A-.” She asks, “Would you be upset to learn that “only” 40% of 8th graders are at an A level in math and “only” 1/3rd scored an A in reading?”

Resseger also cited the huge body of research explaining why School Report Cards aren’t a reliable tool for measuring school effectiveness.

We need a better understanding how and why the word “proficiency” has been weaponized against schools. To do so, we must master the huge body of research which explains why standardized tests aren’t fair, reliable, or valid measures of how well schools are performing.

In 2013, after surveying national experts about “misnaepery,” Education Week explained that NAEP “is widely viewed as the most accurate and reliable yardstick of U.S. students’ academic knowledge … But when it comes to many of the ways the exam’s data are used, researchers have gotten used to gritting their teeth.”

Also in 2013, James Heckman, a Nobel Prize laureate who lived in Oklahoma City as a child, warned of the dangers of misusing test data. In 2025, Heckman and his co-author, Alison Baulos, published “Instead of Panicking over Test Scores, Let’s Rethink How We Measure Learning and Student Success.” They urge us to “pause some tests and redirect resources toward more meaningful ways to promote and assess student learning.”

They don’t oppose the use of tests as one measure when used for diagnostic purposes; those metrics “may be valuable for tracking large-scale trends — such as monitoring recovery from the COVID-19 pandemic.” However, “the current overreliance on tests is costly in many ways and is not an effective strategy for improving education as a whole.” And, “standardized tests often conceal more than they reveal.” 

Getting back to recent headlines, I appreciate the press’ reporting on Ryan Walters’ lack of transparency. I’m even more impressed with their reporting on the lack of evidence to support his claims that his administration has improved outcomes. But they now need to report on the reasons why the Commission made a terrible mistake, apparently based on the alt facts generated by corporate reformers’ false public relations spin.

On May 10, Dana Goldstein wrote a long article in The New York Times about how education disappeared as a national or federal issue. Why, she wondered, did the two major parties ignore education in the 2024 campaign? Kamala Harris supported public schools and welcomed the support of the two big teachers’ unions, but she did not offer a flashy new program to raise test scores. Trump campaigned on a promise to privatize public funding, promote vouchers, charter schools, religious schools, home schooling–anything but public schools, which he regularly attacked as dens of iniquity, indoctrination, and DEI.

Goldstein is the best education writer at The Times, and her reflections are worth considering.

She started:

What happened to learning as a national priority?

For decades, both Republicans and Democrats strove to be seen as champions of student achievement. Politicians believed pushing for stronger reading and math skills wasn’t just a responsibility, it was potentially a winning electoral strategy.

At the moment, though, it seems as though neither party, nor even a single major political figure, is vying to claim that mantle.

President Trump has been fixated in his second term on imposing ideological obedience on schools.

On the campaign trail, he vowed to “liberate our children from the Marxist lunatics and perverts who have infested our educational system.”Since taking office, he has pursued this goal with startling energy — assaulting higher education while adopting a strategy of neglect toward the federal government’s traditional role in primary and secondary schools. He has canceled federal exams that measure student progress, and ended efforts to share knowledge with schools about which teaching strategies lead to the best results. A spokeswoman for the administration said that low test scores justify cuts in federal spending. “What we are doing right now with education is clearly not working,” she said.

Mr. Trump has begun a bevy of investigations into how schools handle race and transgender issues, and has demanded that the curriculum be “patriotic” — a priority he does not have the power to enact, since curriculum is set by states and school districts.

Actually, federal law explicitly forbids any federal official from attempting to influence the curriculum or textbooks in schools.

Education lawyer Dan Gordon wrote about the multiple laws that prevent any federal official from trying to dictate, supervise, control or interfere with curriculum. There is no sterner prohibition in federal law than the one that keeps federal officials from trying to dictate what schools teach.

Of course, Trump never worries about the limits imposed by laws. He does what he wants and leaves the courts to decide whether he went too far.

Goldstein continued:

Democrats, for their part, often find themselves standing up for a status quo that seems to satisfy no one. Governors and congressional leaders are defending the Department of Education as Mr. Trump has threatened to abolish it. Liberal groups are suing to block funding cuts. When Kamala Harris was running for president last year, she spoke about student loan forgiveness and resisting right-wing book bans. But none of that amounts to an agenda on learning, either.

All of this is true despite the fact that reading scores are the lowest they have been in decades, after a pandemic that devastated children by shuttering their schools and sending them deeper and deeper into the realm of screens and social media. And it is no wonder Americans are increasingly cynical about higher education. Forty percent of students who start college do not graduate, often leaving with debt and few concrete skills.

“Right now, there are no education goals for the country,” said Arne Duncan, who served as President Barack Obama’s first secretary of education after running Chicago’s public school system. “There are no metrics to measure goals, there are no strategies to achieve those goals and there is no public transparency.”

I have been writing about federal education policy for almost fifty years. There are things we have learned since Congress passed the Elementary and Secondary Education Act in 1965. That law was part of President Lyndon B. Johnson’s agenda. Its purpose was to send federal funds to the schools enrolling the poorest students. Its purpose was not to raise test scores but to provide greater equity of resources.

Over time, the federal government took on an assertive role in defending the rights of students to an education: students with disabilities; students who did not speak English; and students attending illegally segregated schools.

In 1983, a commission appointed by President Reagan’s Secretary of Education Terrell Bell declared that American schools were in crisis because of low academic standards. Many states began implementing state tests and raising standards for promotion and graduation.

President George H.W. Bush convened a meeting of the nation’s governors, and they endorsed an ambitious set of “national goals” for the year 2000. E.g., the U.S. will be first in the world by the year 2000; all children will start school ready to learn by 2000. None of the goals–other than the rise of the high school graduation rate to 90%–was met.

The Clinton administration endorsed the national goals and passed legislation (“Goals 2000”) to encourages states to create their own standards and tests. President Clinton made clear, however, that he hoped for national standards and tests.

President George W. Bush came to office with a far-reaching, unprecedented plan called “No Child Left Behind” to reform education by a heavy emphasis on annual testing of reading and math. He claimed that because of his test-based policy, there had been a “Texas Miracle,” which could be replicated on a national scale. NCLB set unreachable goals, saying that every school would have 100% of their students reach proficiency by the year 2014. And if they were not on track to meet that impossible goals, the schools would face increasingly harsh punishments.

In no nation in the world have 100% of all students ever reached proficiency.

Scores rose, as did test-prep. Many untested subjects lost time in the curriculum or disappeared. Reading and math were tested every year from grades 3-8, as the law prescribed. What didn’t matter were science, history, civics, the arts, even recess.

Some schools were sanctioned or even closed for falling behind. Schools were dominated by the all-important reading and math tests. Some districts cheated. Some superintendents were jailed.

In 2001, there were scholars who warned that the “Texas Miracle” was a hoax. Congress didn’t listen. In time the nation learned that there was no Texas Miracle, never had been. But Congress clung to NCLB because they had no other ideas.

When Obama took office in 2009, educators hoped for relief from the annual testing mandates but they were soon disappointed. Obama chose Arne Duncan, who had led the Chicago schools but had never been a teacher. Duncan worked with consultants from the Gates and Broad Foundations and created a national competition for the states called Race to the Top. Duncan had a pot of $5 billion that Congress had given him for education reform.

Race to the Top offered big rewards to states that applied and won. To be eligible, states had to authorize the creation of charter schools (almost every state did); they had to agree to adopt common national standards (that meant the Common Core standards, funded wholly by the Gates Foundation and not yet completed); sign up for one of two federally funded standardized tests (PARCC or Smarter Balanced) ; and agree to evaluate their teachers by the test scores of their students. Eighteen states won huge rewards. There were other conditions but these were the most consequential.

Tennessee won $500 million. It is hard to see what, if anything, is better in Tennessee because of that audacious prize. The state put $100 million into an “Achievement School District,” which gathered the state’s lowest performing schools into a new district and turned them into charters. Chris Barbic, leader of the YES Prep charter chain in Houston was hired to run it. He pledged that within five years, the lowest-performing schools in the state would rank among the top 20% in the state. None of them did. The ASD was ultimately closed down.

Duncan had a great fondness for charter schools because they were the latest thing in Chicago; while superintendent, he had launched a program he called Renaissance 2010, in which he pledged to close 80 public schools and open 100 charter schools. Duncan viewed charters as miraculous. Ultimately Chicago’s charter sector produced numerous scandals but no miracles.

I have written a lot about Race to the Top over the years. It was layered on top of Bush’s NCLB, but it was even more punitive. It targeted teachers and blamed them if students got low scores. Its requirement that states evaluate teachers by student test scores was a dismal failure. The American Statistical Association warned against it from the outset, pointing out that students’ home life affected test scores more than their teachers.

Duncan’s Renaissance 2010 failed. It destroyed communities. Its strategy of closing neighborhood schools and dispersing students encountered growing resistance. The first schools that Duncan launched as his exemplars were eventually closed. In 2021, the Chicago Board of Education voted unanimously to end its largest “school turnaround” program, managed by a private group, and return its 31 campuses to district control. Duncan’s fervent belief in “turnaround” schools was derided as a historical relic.

Race to the Top failed. The proliferation of charter schools, aided by a hefty federal subsidy, drained students and resources from public schools. Charter schools close their doors at a rapid pace: 26% are gone in their first five years; 39% in their first ten years. In addition, due to lax accountability, charters have demonstrated egregious examples of waste, fraud, and abuse.

The Common Core was supposed to lift test scores and reduce achievement gaps, but it did neither. Conservative commentator Mike Petrilli referred to 2007-2017 as “the lost decade.” Scores stagnated and achievement gaps barely budged.

So what have we learned?

This is what I have learned: politicians are not good at telling educators how to teach. The Department of Education (which barely exists as of now) is not made up of educators. It was not in a position to lead school reform. Nor is the Secretary of Education. Nor is the President. Would you want the State legislature or Congress telling surgeons how to do their job?

The most important thing that the national government can do is to ensure that schools have the funding they need to pay their staff, reduce class sizes, and update their facilities.

The federal government should have a robust program of data collection, so we have accurate information about students, teachers, and schools.

The federal government should not replicate its past failures.

What Congress can do very effectively is to ensure that the nation’s schools have the resources they need; that children have access to nutrition and medical care; and that pregnant women get prenatal care so that their babies are born healthy.

A few years back, the story of the “Mississippi Miracle” in reading was all the rage. The increase in scores of fourth grade students on NAEP scores was hailed as miraculous, a testament to the dramatic power of the “science of reading.” New York Times’ columnist Nicholas Kristof wrote a column praising Mississippi for raising the test scores of its fourth graders without spending any more money. Anyone could do it!

I was critical of Kristof’s enthusiasm and pointed out that the scores of fourth graders soared but the reading scores of eighth graders did not. The scores of the older students were among the lowest in the nation. What kind of “miracle” dissolves as students get older?

Thomas Ultican reviews the “Mississippi Miracle” and also finds it to be hype. But he sees it as good reason to kill NAEP, which Trump is now doing.

I don’t often disagree with Tom, who is a relentless researcher of scams and hoaxes perpetrated by the critics of public schools.

I oppose the misuse of high-stakes standardized tests to hold teachers, students, and schools “accountable,” because the tests are loaded with errors and inevitably reflect family income and family education, not the ability of students or teachers. I have written about the inherent flaw of standardized tests in my last three books.

What I like about NAEP is that it is a no-stakes test. It too reflects family income and family education, like all standardized tests. But no one is punished or rewarded for their test scores.

NAEP shows trends by states, cities, gender, race, ethnicity, special ed status, income, etc.

It is NAEP that reveals the lie behind the “Mississippi Miracle.” NAEP shows that fourth graders made dramatic progress and minimal sleuthing demonstrates that the lowest performing students were held back in third grade, excluded from the testing pool.

It’s NAEP that reveals that eighth graders placed 43rd of 50 states. The Miracle didn’t persist.

I think NAEP should remain and the federal mandate for testing every child every year in every school should be abandoned.

Trump or Musk or a bunch of kids who work for DOGE decided that the U.S. doesn’t need to collect statistics or conduct research about the condition of education. So they wiped out the National Center for Education Statistics at the U.S. Department of Education. This is akin to closing down the Bureau of Labor Statistics. NCES is literally the only reliable, nonpartisan source of information about U.S. education. It is not partisan.

NCES is the heart of the U.S. Department of Education. Its purpose is to study “the progress and condition” of American education. It collects data and statistics about every aspect of American education. A bill was passed in 1867 to create an agency with that mission, and that was the beginning of NCES. At first, it was called the Department of Education, but two years later, it was renamed the Office of Education and placed in the Department of the Interior. In 1939, it was shifted to the Federal Security Agency, and in 1953 it became part of the newly created Departnent of Health Dducation and Welfare. In 1979, President Carter signed legislation creating the U.S. Department of Education, and in 1980, the Department began to function.

NCES has always been nonpartisan. It publishes an annual report called The Condition of Education, which is a valuable compendium of facts and trends that covers almost every aspect of education, from preschool through graduate studies. If you want to know the high school graduation rate over the past century, that’s the source. If you want to compare the graduation rates by gender or race, that’s there too.

NCES also oversees the National Assessment of Educational Progress (NAEP), the federal testing program known as “the nation’s report card.” NAEP has a bipartisan governing board, which is appointed by the Secretary of Education and serves as a policymaking body.

During my time as Assistant Secretary of Education for the Office of Education Research and Innovation from 1991-93, NCES was in my domain. In 1998, Secretary Richard Riley appointed me to serve on the governing board of NAEP, which I did for seven years. There were parts of my domain that I might have offloaded, but with a scalpel, not a chainsaw.

Musk and his DOGE team just eviscerated not only the Department of Education by firing half its employees, but they laid waste to NCES.

Jill Barshay of The Hechinger Report has the story. The staff of NCES has been reduced from about 100 to 3. Three! I think that’s called a death certificate.

She began:

President Donald Trump promises he’ll make American schools great again. He has fired nearly everyone who might objectively measure whether he succeeds.

This week’s mass layoffs by his secretary of Education, Linda McMahon, of more than 1,300 Department of Education employees delivered a crippling blow to the agency’s ability to tell the public how schools and federal programs are doing through its statistics and research branch. The Institute of Education Sciences (IES) is now left with fewer than 20 federal employees, down from more than 175 at the start of the second Trump administration, according to my reporting. It’s not clear how the institute can operate or even fulfill its statutory obligations set by Congress. 

IES is modeled after the National Institutes of Health and was established in 2002 during the administration of former President George W. Bush to fund innovations and identify effective teaching practices. Its largest division is a statistical agency that dates back to 1867 and is called the National Center for Education Statistics (NCES), which collects basic statistics on the number of students and teachers. NCES is perhaps best known for administering the National Assessment of Educational Progress, which tracks student achievement across the country. The layoffs  “demolished” the statistics agency, as one former official characterized it, from roughly 100 employees to a skeletal staff of just three. 

“The idea of having three individuals manage the work that was done by a hundred federal employees supported by thousands of contractors is ludicrous and not humanly possible,” said Stephen Provasnik, a former deputy commissioner of NCES who retired early in January. “There is no way without a significant staff that NCES could keep up even a fraction of its previous workload…”

The mass firings and contract cancellations stunned many. “This is a five-alarm fire, burning statistics that we need to understand and improve education,” said Andrew Ho, a psychometrician at Harvard University and president of the National Council on Measurement in Education, on social media.  

Former NCES Commissioner Jack Buckley, who ran the education statistics unit from 2010 to 2015, described the destruction as “surreal.” “I’m just sad,” said Buckley. “Everyone’s entitled to their own policy ideas, but no one’s entitled to their own facts. You have to share the truth in order to make any kind of improvement, no matter what direction you want to go. It does not feel like that is the world we live in now.”

The deepest cuts

While other units inside the Education Department lost more employees in absolute numbers, IES lost the highest percentage of employees — roughly 90 percent of its workforce. Education researchers questioned why the Trump administration targeted research and statistics. “All of this feels like part of an attack on universities and science,” said an education professor at a major research university, who asked not to be identified for fear of retaliation. 

The future of NAEP is up in the air. The staff to oversee contracts for data collection, testing, and analysis of results is gone.

Please open the article and read it. This is a deliberate death-blow to the most important function of the U.S. Departnent of Education: the collection and dissemination of facts, data, statistics, and trends in the states and the nation.