About a dozen years ago, I heard from Ted Dintersmith. I didn’t know anything about him other than that he was a venture capitalist who was interested in education reform. My hackles went up, because the term “venture capitalist who was interested in education reform” immediately conjured up images of Democrats for Education Reform,” a faux group that is dedicated to charter schools, Teach for America, evaluating teachers by student test scores, and union-busting.
When I thought about the people in this category, I thought about very wealthy men like Ravenal Boykin Curry IV, Whitney Tilson, John Perry, Ken Griffin, Jeff Yass, Daniel Loeb, Arthur Rock, Joel Greenblatt, Bruce Rauner, and Paul Tudor Jones. They are financiers who have committed millions to the cause of privatizing public schools. They are united in their contempt for public schools, teachers, and unions.
So I was wary of Ted and suspicious of his motives. And I pushed him aside, thinking he was part of the cabal of know-it-all entrepreneurs with a zeal to reform (privatize) public schools.
So, I now publicly admit that I was wrong. Ted has traveled the country, visiting schools and listening. He knows how hard teachers work, and he admires them. He has teamed up with some savvy people, like Tony Wagner and the late Sir Ken Robinson, to produce films and write books.
When I recently read his latest article, I realized that I had completely misjudged Ted. He doesn’t want to destroy public schools or privatize them. He wants to make them better. He wants students to love school and learn what they need to know to make informed choices.
He has seen some good charter schools (as I have), but he knows that a few hundred charter schools won’t lead to the transformation of tens of thousands of public schools. And–my thought–to the extent that charter schools defund public schools, they create more problems, more inequities.
So, I hereby apologize to Ted Dintersmith for prejudging him and for lumping him with the Billionaire Boys Club, those who see school transformation as a hobby.
When you read this post you will see that he doesn’t have all the answers, but he is asking the right question.
He writes:
About these notes: Always Monday. Always short. Always free. Always about civil society’s future. At times, interesting. If this was forwarded to you, subscribe here.
First, thanks for your feedback on my recent posts. My goal of late is to explore the various factors causing America to unravel. Toward a systems model of the forces undermining America’s democracy, moral compass, and social cohesion. Today, education.
For five decades, U.S. education policy has been driven by two goals: test scores and college. The No Child Left Behind Act in 2002. Then in 2010, Common Core. Obama’s rallying cry: “By 2020, the United States will once again lead the world in college completion.” Our all-consuming education focus. With massive collateral damage.
Let’s start with the funnel. While the college path works out well for some, most end up on the wrong side of education’s Bell Curve. Over half go directly to career, lacking a hirable skill and often with gutted self-esteem. Of those heading off to college, half don’t finish. Of those who graduate, half don’t get any kind of dream job. A system that’s produced some 43 million adults burdened by $1.8 trillion of student loan debt.
With a boost from Harvard-educated JFK, America grew to equate a prestigious college degree with human worth. This equivalence worked its way into news stories, movies, television shows, commercials . . . and labor markets. Prestige employers (that’s you, Goldman Sachs and McKinsey) made an elite college degree the essential hiring criterion. Less glamorous companies (that’s you, Sherwin-Williams) made a college degree mandatory for promotions. Rich families push their kids from the womb toward an Ivy+ degree. Poor families want the same for their kids. University research teams (that’s you, Georgetown) produce skewed studies touting a massive $1.2 million earnings lift for college grads. The Federal government opens the spigot for readily-obtained, impossible-to-shake student loans.
We’d like to think that America’s education system levels the playing field, but it does the opposite. The SCOTUS 1973 San Antonio Independent School District v. Rodriguez says it’s just fine for U.S. public schools to be largely funded by local property taxes. Rich kids are showered with resources’; poor kids show up to school hungry. Test scores point to an education ‘achievement gap’ that gets pinned on our teachers – not America’s staggering income inequality. The system’s winners feed elite colleges that educate more students from the richest 1% of families than from the bottom 50%. For more, read Paul Tough’s book The Inequality Machine.
With college placements as the measure of success, America’s K12 schools became college-ready test-prep institutions. Pit kids in dog-eat-dog competitions: class rank, SAT scores, AP scores, college placement. A barrage of state- and Federally-mandated multiple-choice exams that rank kids, teachers, schools, districts, and states. Clear marching orders for K12 schools – train kids to perform well on multiple-choice exams and get them off to some college.
To this day, America relies on two numbers to define the education quality of a school, district, state, nation, and human. Math and reading scores. That’s it. Don’t believe me? Just note what journalists write and what policymakers cite. We’ve pushed generations of young Americans to get those damn test scores up. No worries if we short-change career-based learning, civics, history, and art. No worries if we disenage kids and demoralize teachers. No worries if a relentless focus on rote learning crushes a child’s curiosity, creativity, joy, and agency. No worries if kids get the soul-sucking message that what matters in life is outcompeting others on pointless tasks. Education priorities that encapsulate the classic definition of insanity – doing the same thing over and over, expecting different results.
So many reforms, so little change
During their formative K12 years, America’s kids take more than 100 standardized exams. These cost-minimizing exams are expressly designed to be graded by a computer. And if a computer can grade a task, it can do that task. In every sense, we push kids to develop rote skills that machines handle flawlessly. Already, AI aces every assignment and exam required to graduate magna cum laude from Harvard. Reading exams call for regurgitating what’s put in front of you – perfect training for conspiracy-absorbing adults. Math exams are tied to esoteric micro-procedures (e.g., factoring polynomials) that adults never use. Yet we chase these scores, for decades, making no progress, ignoring Goodhart’s Law: “Any measure that becomes a target ceases to be a good measure.”
There you have it. America’s colossal education botch. Train kids to perform rote tasks. Crush out the essential human traits of curiosity, audacity, creativity, agency, and joy. Standardize what’s taught and tested to produce data that ranks and sorts — rather than helping each child develop their distinctive human potential. Most get a chronic ‘you lack proficiency’ message in a system that heaps advantage on the affluent. Botched priorities, paving the way to our current nightmare. A populace unable and unwilling to critically analyze the torrent of lies coming from the White House. A populace that holds its nose to vote for a convicted felon who rails on the ‘educated elite.’ A boiling cauldron of resentment, corruption, and anti-science vitriol. As to how this has worked out, read David Brooks’ fascinating How the Ivy League Broke America.
I don’t offer these views on education casually. For the past fifteen years, I’ve immersed myself in this world. I’ve visited hundreds of schools, and convened thousands of conversations, all across America. I celebrate the bright spots in my books (What School Could Be: Insights and Inspiration from Teachers Across America) and films (Most Likely to Succeed and Multiple Choice). But I’ve seen the failure. Failure caused by an obsolete system and failed accountability measures. Our classroom teachers are heroes. It’s not easy to fight the system, but many do. Give special thanks to the teachers in your life, who persevere despite being shortchanged on pay, respect, and trust. They endure in an entrenched system that has churned out generations of young adults ill-prepared for career, citizenship, and life. A system sowing the seeds for democracy’s collapse.
Three Bold Ideas
Career-Based Learning for All: My recent film Multiple Choice showcases a mainstream public school district that immerses all high-school kids in career-based learning across a range of skills (e.g., carpentry, welding, healthcare, cybersecurity, digital media, AI). Better for the career-bound and, a bit surprisingly, better for the college-bound. The key is making career-based learning essential for all kids, not just ‘those’ kids who don’t resonate with academic curriculum.
Finland: America can learn much from Finland. Their remarkable education progress resulted from a budget crisis that forced them to choose between data and educator excellence. They dumped all standardized tests and focused on developing outstanding classroom teachers – better training, more pay, loads of respect. Finnish kids spend comparatively little time on ‘school’; even high-schoolers have just six hours of class time and homework during the school week. Lots of time for play and exploration. A well-educated populace. For more, read Pasi Sahlberg’s Finnish Lessons.
Accountability: Base accountability on a child’s growing ability to create and carry out initiatives that help make their world better. Portfolios of purpose, not a checklist of counterproductive numbers. Check out New York’s Performance Standards Consortium or read Tony Wagner’s Mastery.
Bonus Idea: Long talked about, America would be well-served if every high school graduate engaged in public service — lifting up their community while developing important skills. Those going on to college bring maturity and real-world experience to years often wasted on beer pong. Those eschewing the college path are off and running with purpose and skills. Young Americans across all demographics in a melting pot that bridges our current demographic and political divides.
The biggest lie about American school kids is that most are “below grade level.” This lie is repeated so often by prominent figures that it is widely believed. But it’s not true. Those who believe it are wrong. Those who repeat it, knowing it’s not true, are liars.
The source of the lie and the confusion is clear: the achievement levels in which NAEP scores are reported. The levels are “advanced,” “proficient,” “basic,” and “below basic.” When the media write about the latest release of NAEP scores, they frequently treat “proficient” as “grade level.”
But “proficient” is NOT “grade level.” It represents solid achievement, a rigorous aspirational goal. “Proficient” is equivalent to a solid A.
Every NAEP report on test scores says clearly in a footnote that “proficiency” is not the same as grade level. For example: “NAEP Proficient does not signify meeting grade-level expectations.” Yet the media and prominent commentators who should know better repeat the lie that most students are below grade level. The fact is that most students will never reach the high bar of “proficient.”
In 2023, as Bruce Lesley points out, Biden’s Secretary of Education–Miguel Cardona–testified to a Congressional committee that only one-third of American students were reading “at grade level.” I was flabbergasted. I couldn’t believe he said something so outrageous. I called Dr. Peggy Carr, who at that time was the Commissioner of Education Statistics. She was as surprised as I was that Secretary Cardona repeated the erroneous statistic. I asked Dr. Carr whether she had ever briefed him on understanding NAEP results; she had not.
I gave her an idea. Propose a change in name for “proficiency.” Change the name to “mastery.” No one would claim that “mastery” was the same as “grade level.” She liked the idea and promised to take it to the board. Whether she did, I don’t know. But nothing changed.
Bruce Lesley wrote this open letter to the National Assessment Governing Board, which oversees NAEP testing. Lesley is president of First Focus on Children and its partner organization First Focus Campaign for Children, bipartisan advocacy organizations dedicated to making children a priority in federal, state, and international policy. He has led both organizations since 2006 and 2009, respectively, building them into recognized national voices on child health, education, early childhood, economic security, budget and tax policy, immigration, children’s rights, and more recently, international child policy.
He wrote:
To the National Assessment Governing Board, the National Center for Education Statistics, and the leadership of the National Assessment of Educational Progress:
Every institution whose work affects children should begin with one question: “Is this good for children?”
By that standard, the National Assessment of Educational Progress (NAEP) has some important issues that deserve to be resolved. First and foremost, your achievement-level labels — “Basic,” “Proficient,” and “Advanced” — are being weaponized against the very children NAEP exists to serve, and you know it, because your own staff has been saying so for twenty-five years.
To be clear, this open letter is not a claim that NAEP’s underlying data is necessarily wrong, and it is not an argument against NAEP. The argument and request is narrower: you have a real and critically important ethical responsibility to correct the public misuse of your own data. NAEP should defend its credibility against those currently diminishing it.
This Week’s House Mark-Up Provides Another Example
On July 15, 2026, the House Education and Workforce Committee marked up a ten-bill package to facilitate the dismantling of the U.S. Department of Education.
In his opening statement, Chairman Tim Walberg (R-MI) argued that “too many children can’t read or do math at grade level,” and used that claim as a central justification for several of the bills. That claim is false.
Chairman Walberg was drawing on NAEP data — the statistic that roughly two-thirds of American fourth-graders do not score “Proficient” in reading, which is wrongly cited as evidence of failing to meet grade-level reading levels. For some, this is done out of confusion and, for others, to promote a political agenda to undermine public schools. In reality, NAEP proficient is aspirational and reflects a standard that is well above grade level.
Unfortunately, during the markup, multiple members of Congress repeated the same error. But again, NAEP Proficient is not grade level. It has never been grade level.
When the public, the press, the administration, and Congress repeatedly miscite this fact, the National Assessment Governing Board (NAGB) must do much more to clarify and correct misstatements about what it means.
The problem is two fold. One part of the problem is that “proficient” is used on many state and local assessments to mean “at grade level,” or what once upon a time would have been called a gentleman’s C; this leads to some honest confusion for some folks. The other part of the problem is folks who are invested in the narrative that public schools are failing and who benefit from the confusion surrounding the term.
Greene adds:
And every time NAEP scores are released, education journalists write piece after piece explaining “proficient” all over again, usually in the wake of some prominent person decrying the large number of students not “at grade level.”
That confusion is NAGB’s responsibility to address, and it has deserved attention for years, but all the more NOW.
This Is Not a Partisan Problem
Chairman Walberg and his colleagues’ misstatements are only the most recent officials to make this mistake (whether unintentionally out of confusion or internationally), and the pattern runs through both political parties.
Secretary Miguel Cardona, testifying before Congress in April 2023 under the Biden Administration, told lawmakers directly that only one-third of students were reading “on Grade level,” treating a NAEP proficiency figure as if it were a grade-level statistic, in nearly identical language.
And Secretary Linda McMahon, in the current Trump Administration, has used more careful wording — noting that nearly 70% of eighth graders are “not proficient” in reading — but has paired that technically accurate phrase with language implying total system failure. A Snopes piece by Rae Deng described this claim as lacking its own level of reading comprehension because, again, it completely mischaracterizes what NAEP’s “proficient” standard means.
Outside advocacy groups have been considerably less careful than any of them.
Corey DeAngelis, a leading advocate for school privatization, vouchers, and against public education, has cited NAEP proficiency figures directly, without qualification, as evidence that public schools are a system-wide “disgrace.”
Greene captures these types of political misuse of NAEP data in this Substack post.
This confusion is intentional by people arguing for both the dismantling of public education and federal investments in children.
Unfortunately, NAGB’s silence has allowed that rhetorical usefulness to go unchecked under Republican and Democratic administrations alike, and it is being used right now, this week, on Capitol Hill to justify eliminating the very agency that funds and safeguards the data NAGB produces.
NAGB’s Own Experts Have Been Saying This for Years
In 2001, Mary Lynne Bourque and Susan Loomis — a staff member and a board member of the National Assessment Governing Board itself — wrote plainly that the Proficient achievement level “does not refer to ‘at grade’ performance,” and that performance at Proficient is not the same as being “proficient” in a subject as any ordinary person would use that word.
Chester “Checker” Finn, Jr., who chaired the panel that adopted the achievement levels in 1992, has been candid that the levels were designed to be aspirational — a description of where students should ideally arrive, not a diagnosis of where most currently stand.
NCES itself has attached a caution to NAEP score reports for years: the Proficient level “does not represent grade level proficiency as determined by other assessment standards.”
If NAEP’s own architects and NAGB’s own website already say this, it is past time to be diligent in correcting the record when people misuse and misstate what it means. It is also on NAGB to stop publishing results in a format that predictably, foreseeably, and repeatedly gets misread as a verdict on grade-level performance, especially when you can see exactly how that misreading gets used again and again.
The clearest confirmation of all of this comes from NCES’s own data. Researchers Gina Cervetti and Kathleen Hinchman mapped every state’s definition of fourth-grade “grade-level” reading proficiency directly onto the NAEP scale and found that, as of the most recent analysis, nearly every state’s own standard for grade-level reading lines up with NAEP’s Basic level, not NAEP’s Proficient level. That means the honest translation of the data runs the opposite direction from how Chairman Walberg and others use it: by the states’ own definitions of grade level, roughly two-thirds of American fourth graders are reading at or above grade level, not below it.
Cervetti and Hinchman are also blunt about what actually is a crisis in the data: not a reading crisis, but an equity crisis. In 2022, only 48% of students eligible for free or reduced-price lunch scored at or above NAEP Basic, compared with 76% of students who were not eligible — a 28-point gap that has persisted, largely unchanged, for decades.
That is a story about generational wealth and unequal access to housing, healthcare, and school resources, not a story about failing classrooms, and NAEP’s own framing continues to let people tell the wrong story with your numbers.
What Education Writers and Researchers Have Been Saying
Diane Ravitch, who served seven years on the National Assessment Governing Board under President Clinton, has called out the confusion between NAEP Proficient and grade level as one of the most damaging and persistent falsehoods in American education discourse, noting that NAEP itself explicitly warns against the equivalence you continue to permit others to make.
Greene has argued that cut scores like “Proficient” function as scaled, curved judgments dressed up as fixed standards — noting that if every child scored above a cut, the establishment reaction would be to declare the cut too easy, not to celebrate the achievement. That is not how a genuine, fixed criterion is supposed to behave, and it is worth NAGB’s honest reckoning.
Mark Weber, a New Jersey teacher and education researcher, has done careful public work mapping state proficiency standards onto the NAEP scale, and his conclusion undercuts a favorite talking point of your critics-turned-allies in this fight: there is no empirical evidence that closing the so-called “honesty gap” between state and NAEP proficiency rates does anything to improve student achievement. If setting state cut scores to match yours were actually the lever for better outcomes, we would expect to see it in the data. We do not. That matters because it means the standard is being imported into state accountability systems on faith, not evidence — exactly the kind of unsupported claim NAGB should be correcting rather than allowing to spread.
The Brookings Institution’s Brown Center on Education Policy has been making this same case for nearly two decades. Tom Loveless, the Brown Center’s longtime director, authored a 2007 report concluding bluntly that NAEP’s cut scores were set too high.
His 2016 Brookings piece, “The NAEP Proficiency Myth,” went further, noting that the achievement levels came under critical review from the U.S. Government Accountability Office, the National Academy of Sciences, and the National Academy of Education shortly after they were adopted — with the National Academy of Sciences review concluding the achievement levels were fundamentally flawed.
Loveless adds:
Advocates of the NAEP proficient standard want it to be for all students. That is ridiculous. Another way to think about it: proficient for today’s eighth graders reflects approximately what the average twelfth grader knew in mathematics in 1990. Someday the average eighth grader may be able to do that level of mathematics. But it won’t be soon, and it won’t be every student.
That is not a stray outside critique. That is respectable experts in the field, writing for decades, about the very categories NASB is still using today without correction.
One Point Should Not Separate “Failing” from “Successful”
NAGB also owes the public an honest accounting of what a cut score actually is. A cut score is a single point on a continuous scale, chosen somewhat arbitrarily by a panel, above which a child is declared “Proficient” and below which the same child, one point lower, is declared “Basic,” which is actually grade level.
Two children who are functionally indistinguishable in what they know and can do are sorted into entirely different public categories — one used as evidence that a school, a state, or a federal agency is failing, the other treated as evidence of success — because of a single point set by a committee, not because of any meaningful difference in the children themselves.
That is not a rounding error. It is the mechanism by which your data gets converted into political ammunition.
If NAGB cannot explain, in terms parents can understand, why the child who scores one point below the line is a different kind of learner than the child one point above it, then the line is doing rhetorical work the data was never built to support.
As the psychiatrist and educator William Glasser warned schools decades ago, chasing a point or two of movement on a test score is precisely the wrong institutional goal — and yet that is the goal NAEP’s cut scores hand every state, district, and school in the country by default.
Researcher Andrew Ho makes a similar point. He has identified proficiency cut scores as arbitrary markers, set through what he calls an “overwrought, judgmental, and ultimately political process,” not derived from any fixed line in human learning.
Ho has also documented a specific illusion that follows from that arbitrariness: because a large cluster of students always sits near the middle of the score distribution, a cut score placed close to that cluster will make small, ordinary shifts in performance look like dramatic gains or losses, purely as an artifact of how many students happen to sit right at the line — not because anything real changed in how much they learned. A researcher with no stake in the politics of this issue is describing the identical mechanism that turns your data into a rhetorical weapon: the closer the line sits to where children actually cluster, the more your data will appear to swing wildly for reasons that have nothing to do with children’s learning.
Criterion-Referenced in Name, Arbitrary in Practice
NAEP describes itself as a criterion-referenced assessment, distinct from norm-referenced tests like the SAT that simply rank students against one another. That distinction matters, and I want to represent it accurately rather than overstate it — NAEP does not “grade on a curve” in the way the SAT’s percentile scoring does.
However, the practical effect on families is not so different as the label suggests. NAEP’s cut scores were set by hand-picked panels making judgment calls about what students “should” know, not derived from an external, agreed-upon standard of competence, and independent evaluators — including a National Academies review in 2017 — have called for stronger evidence connecting NAEP performance levels to any real-world outcome at all.
A test that is criterion-referenced in name but whose criteria were set arbitrarily, and whose results still track family income and race as tightly as any norm-referenced test on the market, produces the same practical harm as the norming bias critics have long raised: it tells us more about a child’s zip code than about a fixed, meaningful standard of what that child knows.
Notably, NAGB has conceded the point this year. The 2026 NAEP reading framework — administered to students for the first time this spring — now explicitly disaggregates racial and ethnic subgroup results by socioeconomic status, on the premise, well documented for decades, that apparent racial differences in test scores largely track socioeconomic differences. That is a welcome and overdue acknowledgment.
But it is also, in effect, NAGB admitting in 2026 what critics have argued for years: that the results have been measuring wealth and family circumstance as much as they measure “proficiency,” all along. If that acknowledgment is real, it should extend backward, to how NAGB talks about every score ever published, not just forward, to a single new breakdown in the data tables.
The Test Itself Is Not Neutral
Even setting the cut scores aside, the content of the test carries its own bias, and NAEP’s own commissioned reviewers have said so. The NAEP Validity Studies Panel — a technical review body NCES itself created and funds — published an analysis by Gerunda Hughes in 2023 documenting that the statistical methods used to build NAEP-style test items can systematically disadvantage the very students the test is supposed to serve fairly.
When an item is answered correctly by nearly every student, it gets treated as a poor “discriminator” between high and low performers and is typically cut from the test in favor of harder items, even though that easy item may represent exactly the content that should be mastered.
In his book, Au cites researchers Kidder and Rosner, who examined more than 300,000 SAT test-takers and the pool of trial questions used to build future exams and found that some trial questions were answered correctly by Black students, or by Latino students, more often than by White students. Those questions were then discarded — not because they were poor measures of the content, but because they failed to reproduce the racial score gap the rest of the test already produced. A question only “counted” as valid if high-scoring test-takers, who are disproportionately White, tended to get it right in pretesting.
My mother has verified the same process when she was asked to be on a panel to evaluate whether the item questions were “fair”. The publishers of the Texas State assessment at the time ran through the questions and kept throwing out questions as biased toward Black or Hispanic children if they scored the same or close to the scores of White children.
In contrast, questions in which there was a substantial gap in favor of White students were not flagged – thus, “norming” the disparity in test score outcomes into subsequent tests. Although my mother repeatedly objected, she was overruled throughout the day and, not surprisingly, never asked back to be a reviewer.
The result, as Au describes it, is a self-reinforcing loop: item selection is calibrated to match existing racial score gaps, which locks those same gaps into every future version of the test, all without anyone ever explicitly considering race in the selection criteria.
NAEP is a different test administered by a different organization, and I am not asserting that NAEP’s item-selection process has been documented to work in the same way. But NAEP uses the same category of item statistics that made this outcome possible on the SAT, and NAGB’s own validity panel has already flagged the risk. Given what is now documented on a test as consequential as the SAT, NAGB owes the public a direct, public answer to a direct question: has anyone checked whether NAEP’s item-selection process does the same thing?
There is also cultural and geographic bias. As the son of an English teacher and a math teacher, it should be no surprise that I did fairly well on standardized tests throughout my life. But I vividly recall a reading passage from the PSAT that focused on nautical issues and the definition of a “flotilla.”
Having grown up in El Paso, Texas, a city located hundreds of miles from any coastline, the passage and vocabulary word were unfamiliar to any of us taking the test in the desert borderlands. On the other hand, we would crush a passage referring to “tortillas.” NAEP’s own reviewers have a name for this: cultural validity, the idea that a test cannot cleanly separate what a child knows from what a child has been exposed to.
Research that NAEP’s own validity panel cites has found that when students are allowed to choose among reading passages on different topics, rather than being assigned a single passage that may be unfamiliar or uninteresting to them, some groups of students — including Black eighth graders and Hispanic twelfth graders in the panel’s own cited study — score much higher. That is evidence that some of what NAEP currently measures is exposure and familiarity, not just reading ability, and it argues for reform in how passages and vocabulary are chosen, not just in how results are labeled.
Again, the validity panel’s report contains proof that this is a design choice, not a fact of nature. In 1972, the psychologist Robert Williams built a test called the Black Intelligence Test of Cultural Hegemony, using vocabulary and content drawn from Black American culture instead of the dominant culture’s frame of reference. When Black and White teenagers took it, Black students substantially outscored White students by substantial margins.
Nothing about the underlying children changed between that test and the SAT. What changed was whose knowledge and cultural fluency the test happened to be built around.
That single fact should end, permanently, any claim that a test’s outcomes reveal some fixed truth about which children “can” or “cannot” read, think, or reason. What these tests reliably measure is often which cultural and economic frame of reference a child was raised in, and how well that frame matches the one test-makers chose to build around — which is another way of describing accumulated wealth, school funding, and generational inequity, not a verdict on a child’s mind.
That is real, and policies that address school finance inequity, child poverty, childhood hunger, and adverse childhood experiences (ACEs) deserve real policy attention. These issues would undoubtedly do more to improve educational outcomes in this country rather than privatization of public schools or the elimination of the Department of Education.
Claims that two-thirds of American children cannot read at grade level are simply false, and their interpretation by policymakers and advocates is harming children. There is an old warning that was popularized by author Mark Twain but attributable to British Prime Minister Benjamin Disraeli about three kinds of falsehood — “lies, damned lies, and statistics.”
In this case, even a true number, presented without its context, can mislead more effectively than an outright fabrication. NAEP’s “proficiency” level is an aspirational one, but the grade-level story built on top of it is doing real harm. NAGB is a position to explain the difference, and the public is not, until you tell them.
The Damage Is Not Abstract: What Gets Tested Is What Gets Taught
This is not a technical quibble.
Every time “below Proficient” gets reported to the public as “can’t read” or “can’t do math,” it becomes ammunition for defunding public schools and for portraying millions of children — disproportionately low-income children and children of color — as failures because of a label your board itself has said should not be read that way.
It also reshapes what happens inside the classroom. When reading and math scores on tests built around NAEP cut points become the metric by which schools, teachers, and even state superintendents are judged, instructional time follows the incentive:
Short, decontextualized passages crowd out real books — my children were taught how to write a brief constructive response (BCR) before they were even taught what a paragraph was.
Science, government, history, the arts, and physical education are pushed to the margins of the elementary school day because they are not tested and therefore not rewarded.
Children end up narrower, not better educated, in the very subjects that make them informed citizens — and NAEP’s own cut-score architecture is a direct contributor to that narrowing, whether or not that was your intent.
NAGB tried a partial fix in 2018, adding the word “NAEP” before each level — “NAEP Proficient” rather than “Proficient” — so people would stop equating your terms with generic ones.
James Harvey, executive director of the National Superintendents Roundtable, was right to call that gesture insufficient at the time. Harvey said:
…the American people should understand that the misleading term “proficient” sets a performance benchmark beyond the reach of most students in the world.
Harvey argued “proficient” should be changed to something like “high” to avoid being “fooled.”
His point has been proven many times, including this week when a sitting congressional committee chairman, citing NAEP-adjacent data to justify eliminating a federal agency, still used the word “grade level” as if it meant what NAEP’s Proficient level does not mean.
The Perverse Incentive NAEP Has Inspired: Grade Retention As Score Manipulation
The clearest evidence that NAEP’s cut scores create perverse incentives and “manufactured” crises, rather than honest information, is what states have started doing in response to them: holding back third-graders who miss an early-literacy cut score, in order to produce a fourth-grade NAEP cohort that looks better on paper.
Education professor and researcher Paul Thomas has documented this closely in states such as Mississippi, where fourth-grade reading gains celebrated as a “Mississippi miracle” tracked closely with a mandatory third-grade retention policy.
A child who is nine years old competing against classmates who are eight will predictably score higher on a test built around the same content; that is a fact about test administration, not about literacy. Furthermore, those same “gains” have been shown to fade by eighth grade, once the retained cohort catches up in age to its peers without having genuinely caught up in learning.
This is worth NAGB’s own honest reckoning, not because the research on retention is unanimous — reasonable analysts, including some closely tied to NAEP’s own governing board, dispute how much of Mississippi’s gain is genuine instructional improvement versus retention’s effect on cohort composition — but because NAEP’s achievement levels are the mechanism creating the incentive either way.
States are not retaining eight-year-olds because it is good for those children. They are retaining them because a single cut score on a single test has been elevated to a measure of whether a state’s education policy is working. The cost of that incentive falls on children: retained students who show a short-term score bump can, over time, experience the opposite of what was intended — greater disengagement, higher rates of dropping out before graduation, and the well-documented psychological toll of being told, at eight or nine years old, that they failed.
William Glasser spent much of his career, in Schools Without Failure, and later in The Quality School, explaining exactly why this backfires. He argued that standardized testing reduces learning to disconnected, memorized facts at the expense of critical thinking and real application — and that the “right answer, wrong answer” format of a multiple-choice test teaches children that education is a hunt for a single predetermined answer rather than a process of genuine understanding.
In the schools Glasser held up as models, closed-book tests were replaced with open-book, collaborative assessments that actually resembled the problems students would face outside school. His deeper claim, grounded in what he called Choice Theory, is that people — including children — are driven by needs for freedom, power, and simple enjoyment in their work, and that using test scores to rank, shame, or coerce students destroys the very motivation that produces quality work in the first place.
Labels matter. When children repeatedly hear that two-thirds of them “cannot read at grade level,” many internalize failure that is not supported by the evidence. Parents lose confidence in neighborhood schools. Teachers become demoralized. Policymakers propose increasingly radical structural changes to fix a crisis that has been inaccurately described.
Glasser also warned explicitly against making small, arbitrary numerical gains — his example was raising a test score by a point or two — the primary institutional goal of a school, insisting instead on building a genuine culture of quality. That is precisely the trap a single-point NAEP cut score sets for states, and it is the trap third-grade retention policies walk students directly into.
That is the opposite of what an assessment meant to serve children should produce, and it deserves your acknowledgment, not your silence.
What We Are Asking You To Do, Now
Issue a direct, public correction each time a federal official misstates NAEP Proficient as “grade level,” the way you would correct any other material misuse of your data. Silence is not neutrality; it is acquiescence in the misuse.
Publish, prominently and alongside every score release, the state-by-state mapping showing that “grade level” as states themselves define it corresponds to NAEP Basic, not NAEP Proficient — the analysis your own data already supports and that outside researchers have had to do on your behalf.
Publish a plain-language document — “What NAEP Proficient Does, and Does Not, Mean” — and require it alongside every score release, every webpage, every press briefing, and every congressional testimony that cites NAEP data. Most of this letter’s argument could be prevented by a single page NAGB.
Stop using “Basic,” “Proficient,” and “Advanced” as headline labels without their NAEP qualifier in every release, chart, and public statement — not as a footnote, but as a mandatory part of the label itself, displayed with the same prominence as the number.
Retire “Basic,” “Proficient,” and “Advanced” altogether in favor of terms that do not already carry a plain-English meaning your data does not support. If the words themselves are the problem, changing a modifier in front of them has not been enough.
Extend the honesty of the 2026 reading framework’s socioeconomic disaggregation backward, not just forward. If you now accept that racial gaps in your data are substantially explained by family socioeconomic status, say so plainly every time a racial achievement gap is reported, and stop letting that gap be cited as evidence of school failure without that context.
Commission and publish the external validity evidence the National Academies asked for in 2017 — a transparent accounting of what your cut scores do and do not predict, so the public can evaluate the standard rather than take your word for its meaning.
Publicly acknowledge the perverse incentive your cut scores have created for third-grade retention policies, and commission independent, longitudinal research — tracking students well past eighth grade, through high school graduation — before any state is permitted to point to NAEP gains as proof that retaining eight-year-olds is good policy.
Act on your own validity panel’s 2023 findings, and answer the question the SAT evidence now raises. Publicly disclose whether NAEP’s item-selection process has ever been audited for the same self-reinforcing bias documented on the SAT — where trial questions that marginalized students answered correctly were discarded for failing to reproduce the existing score gap — and commit to an independent audit if it has not. Explain how tests are “normed” from one year to the next and made comparable in a manner that is understandable to the public.
Kids can’t wait for another year of this same correction being offered and ignored, or for another cohort of eight-year-olds to be held back so a state’s chart can look better. NAGB has the power to end the confusion your own board identified more than two decades ago. Please do so. Kids deserve it.
The assumption behind these questions is test scores are precise indicators of reading ability, like scientific laboratory measurements. But like blood pressure levels — in which there is agreement about what’s being measured — they are variable and open to interpretation.
Despite the subjectivity and lack of agreement in defining reading, grade level and ability, grade-level reading ability is often mistakenly viewed as determined by a precise, stable test score, one that does not take into account factors such as students’ health and hunger in their testing performance. Not everyone agrees on what is salient at a particular grade level, leading to subjectivity in weighting phonics knowledge, vocabulary, comprehension, the ability to synthesize a theme and recognize an author’s point of view in a given passage, or some combination of such things.
Standardized tests are typically the basis for establishing grade level. But test scores themselves don’t indicate grade level, which is a creation of an interpreter. That’s why different tests don’t always produce the same grade level. A student who tests at fourth grade in one state may test at the third or fifth grade when moving to another state using a different test. In short, different tests or standards can produce different grade levels.
David Reinking is a retired professor at Clemson and the University of Georgia. He is an inductee in the Reading Hall of Fame, and a former co-editor of Reading Research Quarterly and Journal of Literacy Research. (Courtesy)
The National Assessment of Educational Progress calls itself “the nation’s report card” even to the point of using the phrase on its website and then having it repeated as if it is an established fact. It is often invoked in commentaries on grade levels. But it wasn’t designed for that purpose. NAEP itself states that “NAEP Proficient achievement level does not represent grade level proficiency as determined by other assessment standards (e.g., state or district assessments). NAEP achievement levels are to be used on a trial basis and should be interpreted and used with caution.”
But that hasn’t stopped many policy makers and journalists from trying to connect a NAEP test score to a grade level. Giving in to political pressure and rejecting recommendations from authorities in developing tests, in 1990 NAEP officials did introduce four tiers of reading achievement: Below Basic, Basic, Proficient and Advanced. These levels were established solely using subjective judgment about what’s expected of children tested at a specific grade level.
Peter Smagorinsky is a retired professor at UGA, an inductee in the Reading Hall of Fame, and a former co-editor of Research in the Teaching of English. (Courtesy)
Then, they set equally subjective cut scores to establish boundaries between these four levels. States often use a parallel model for their own tests. In Virginia, student performance is measured on a 0–600 scale, with proficiency set at 400-499 and advanced at 500 or above. It’s hard to imagine that a meaningful difference exists between a student scoring 499 (proficient) and 500 (advanced).
Much confusion is also centered in interpreting whether “basic” is acceptably normal or if it is reasonable to expect all students to be “proficient.” Many commentators, some of whom have a vested interest in arguing that there is a reading crisis, argue the latter. Some have promoted the false idea that “proficient” is grade-level reading, which it absolutely is not. Then, they wrongly argue that two-thirds of American students are reading below grade level by counting “basic” scores as below grade level.
A number of educators have debunked the conclusions of NAEP misinterpreters. Yet, the dogged belief persists that everything can be reduced to subjective interpretations of test scores divided into hierarchical categories that can be falsely, if conveniently, converted to grade levels.
We are concerned whenever we encounter all-too-common misinformation about grade-level reading ability. When misinformation becomes disinformation offered by those who use grade-level reading ability to advance political, polemical or ideological agendas, we become concerned about how faith in test scores lends them to manipulation and deception to help create the crisis that critics have historically claimed is engulfing schools, only to be saved by their favorite solutions.
David Reinking is a retired professor at Clemson and the University of Georgia, an inductee in the Reading Hall of Fame, and a former co-editor of Reading Research Quarterly and Journal of Literacy Research. Peter Smagorinsky is a retired professor at UGA, an inductee in the Reading Hall of Fame, and a former co-editor of Research in the Teaching of English.
Mike DeGuire, retired Denver educator, warned Coloradans that the usual billionaires are lining up behind Mike Bennett for the Democratic nomination for Governor. Bennett is currently a Senator but previously was Superintendent of Schools in Denver, where he promoted the NCLB agenda of test-and-punish, charters schools, and corporate reform. He never was an educator so he swallowed corporate reform hook, line, and sinker.
Colorado’s Democratic primary for governor between Attorney General Phil Weiser and U.S. Sen. Michael Bennet is heating up. TV ads are everywhere, and social media is abuzz with supporters extolling their favorite candidate’s strengths or the opponent’s weaknesses. Colorado has elected only one Republican governor in 50 years, so many pundits believe whoever wins the Democratic primary will likely win the November election.
Money is becoming a big factor in this campaign. Bennet has a distinct advantage thus far, primarily due to one group of funders: billionaires. More than half of Bennet’s super PAC donations are from billionaires, individuals and groups affiliated with organizations run by billionaires, and from a “dark money” group. Research shows that billionaires “are swaying elections all across America.”
As of the May 18 filing deadline, Bennet had over $11.5 million in total donations compared to Weiser’s $7 million. Over $7 million of Bennet’s money is from his super PAC, Rocky Mountain Way, which includes over $1 million from an independent expenditure dark money organization called Brighter Future for Colorado. Weiser has $1.1 million from his super PAC, Fighting for Colorado, and just over $6 million from individual donations.
Michael Bloomberg is the 18th richest man in the world with a net worth of over $109 billion, and he is the largest individual donor to Bennet’s super PAC, giving $2.5 million thus far. But he is not the only billionaire donor in Bennet’s camp. These billionaires also contributed to Bennet’s super PAC: Steve Mandel and his wife ($175,000,); Tench Coxe and his wife ($100,000); Edythe Broad ($3,000); Marc Heising ($75,000); Eric Mindich ($25,000); Deborah Simon ($25,000); and Robert Fanch ($25,938).
In addition to the billionaires’ money, over a dozen hedge fund managers and venture capitalists contributed between $10,000 and $100,000 each to Bennet’s super PAC. The ultra-wealthy use their donations to gain loyalty from candidates who will enact policies that align with their values and protect their wealth through tax breaks, financial incentives and limited regulations on their corporations. They also use nonprofit foundations to fund organizations they support philosophically.
Tax filings published by ProPublica for the years 2022-24 show that billionaires Reed Hastings and John Arnold used their nonprofit, City Fund, to give money to Denver Families for Public Schools, which contributed $45,000 to Bennet. The former CEO of City Fund, Neerav Kingsland, donated $2,000. The Bloomberg Family Foundation donated millions to the Charter School Growth Fund. That nonprofit also funds the Colorado League of Charter Schools which, along with 50Can and Stand for Children, gave $470,000 to Bennet’s super PAC. Bloomberg’s dark money group, the American Opportunity Action, gave $45,000. The total investment from Bloomberg and other billionaire-funded nonprofits surpasses $3 million.
Bloomberg’s support for Bennet’s candidacy reflects a relationship and shared philosophy on education reform that stretches back nearly two decades. Before Bennet entered the U.S. Senate, he served as Denver’ school superintendent from 2005 to 2009, the same time that Bloomberg was serving as New York mayor, where he had control of the city’s schools. Like Bennet, Bloomberg promoted corporate education reforms, oversaw the expansion of charter schools, test-based accountability systems, and market-oriented policies.
Both Bennet and Bloomberg ran for president in 2020. Bloomberg spent over $37 million of his own money on his unsuccessful campaign. Bennet received money for his candidacy from over 32 billionaires who were hedging their bets on who would eventually win the party’s nomination. Several billionaires supporting Bennet for president included some of the richest people in Colorado: the Ergen family, Pat Stryker and Ken Tuchman.
While Bloomberg often wins when he donates money to candidates, there are exceptions. Last year, Bloomberg joined with 26 other billionaires to support former Gov. Andrew Cuomo in the New York mayoral race, donating $13 million to his campaign. New Yorkers resoundingly said no to the billionaire money and elected Zohran Mamdani.
The money involved so far in this year’s gubernatorial Democratic primary pales in comparison to the $34 million spent in the last contested Colorado Democratic primary, in 2018.Many observers believe that Gov. Jared Polis basically bought the governor’s seat by contributingmore than $22 million of his own money to defeat three other candidates. Bloomberg was also involved in the 2018 gubernatorial race, donating $2 million to Mike Johnston who came in third to Polis. Five years later, Bloomberg helped Johnston win his 2023 race for Denver mayor when he and another billionaire, Reid Hoffman, donated nearly $2 million to Johnston’s election.
Ballots drop June 8 for the June 30 Democratic primary. Will the independent and Democratic voters buck the trend of billionaires swaying elections and elect Weiser, or will this billionaire investment pay off for Bennet?
More than 600 faculty in STEM fields at the University of California signed a letter asking for the restoration of the SAT or ACT for students who want to major in STEM fields, according to the Chronicle of Higher Education. They complained that too many students enroll in STEM classes without adequate preparation.
Absent a test requirement, the faculty said, too many severely unprepared students were choosing STEM majors, where they were certain to fail.
It calls on university leaders to reinstate the requirement that applicants for STEM-intensive majors submit SAT or ACT math scores. In 2020, under legal pressure and equity concerns, the system eliminated that requirement and urged public colleges to start accepting more students from impoverished high schools. Critics said the testing requirement unfairly favored privileged students and wasn’t the best predictor of college success.
“The SAT/ACT mathematics requirement is not an obstacle to equity; rather, it is a prerequisite for it,” the letter, which was distributed by faculty members in the math department at the University of California at Berkeley but signed by faculty members systemwide, said.
“Failing to measure preparation gaps does not remove barriers; it moves them into the classroom, where they become harder to overcome. An admissions process that ignores foundational readiness does a disservice to the most vulnerable students.”
Without standardized-test results or other reliable readiness measures, it’s hard to know which students are actually prepared for STEM majors, the letter says.
For those of us who have criticized standardized tests, based on their inherent flaws and their current overuse, this is a reminder that these instruments are valuable for some purposes. In highly competitive fields, like the STEM subjects, it makes no sense to admit college students whose skills are inadequate to the challenge. College professors should not be expected to teach midddle-school math.
Those colleges that choose an open-admission policy are free to do so.
But where the field of study requires a certain level of preparation, students should demonstrate that they are ready and prepared as a condition of admission.
Universities that don’t like standardized tests could offer their own test.
Which brings us back to the opening of the 20th century, when a large number of colleges created the College Entry Examination Board to devise a common test that would demonstrate whether or not students were ready for college.
The Board administered a test each year that assessed students’ knowledge and ability in courses. The “college boards,” as they were known, required full answers to thoughtful questions. They were not standardized and machine-scored. Students were told in advance which works of literature would be assessed and read them to be prepared.
The “college boards” were read and scored by college and high school faculty.
The hand-written exams were replaced by the standardized exams in 1941, on Pearl Harbor day. The leaders of the CEEB sacrificed the old style exams with the onset of the war. It was a move they had wanted to make, to save money and time.
Ever since, we have struggled with the reality that some kind of test was necessary to demonstrate college readiness, alongside the awareness that the standardized tests are biased in favor of students with higher family incomes. They are also biased in favor of students who attended good schools with experienced teachers, advanced classes, and ample resources.
A group of activists in Colorado speaks out against standardized tests. Angela Engels’ article was printed in the Colorado Times Recorder.
A Message from Judy Solano, Chair, A4PEP (Advocates for Public Education Policy).
“It takes courage to speak out about the injustices in the world, especially when policy-makers funded by wealthy education reform organizations hold the power. May we all be warriors in the battle for truth.”
COLORADO TIMES RECORDER
Test-Based Accountability Is Failing Colorado’s Children
As the Colorado General Assembly wraps up the 2026 session, lawmakers once again failed to confront one of the most costly and disruptive features of public education: high-stakes standardized testing.
Key legislative proposals that would have addressed the burden of high-stakes testing were defeated again. SB26-068would have reduced CMAS standardized tests to the minimum extent possible, and HB26-1291 would have reduced teacher evaluations for non-probationary teachers from annually to every three years. Both had bipartisan sponsorship and were supported by teachers, parents, and community members, but opposed by billionaire-backed education reform lobbyists.
A memorial backed by the Advocates for Public Education Policy (A4PEP), urging Congress to return authority to locally elected school boards, as provided in the eww hearing.
These decisions deserve more than a procedural explanation.
The issue of high-stakes testing is neither marginal nor new. Since the passage of No Child Left Behind in 2002, policymakers have imposed lengthy and expensive criterion-referenced standardized tests on students, then incorrectly used that data to make high-stakes decisions about teacher pay and school closures. Because test scores are most closely correlated with income, policymakers have tolerated the practice of closing schools in low-income neighborhoods at the expense of our most vulnerable students.
A 2014 study showed that Colorado’s testing requirements cost up to $78 million annually in combined state and district expenditures. Adjusting for inflation, that would equal more than $100 million today. Approximately 450,000 students spend an average of 20 hours of classroom instruction each year on these assessments — more than 9 million hours diverted away from teaching and learning.
Meanwhile, the number of students identified as at-risk has increased by 118%, more than doubling since 2000. After more than two decades, the results of this approach are clear. Achievement gaps have not closed. Teacher attrition continues to rise. Families increasingly question a system that prioritizes testing over learning, with many choosing to opt out altogether.
This is not a policy in need of minor adjustment. It is an accountability structure that has failed to deliver on its core promises — and it deserves fundamental sreplacement. And yet, it remains firmly in place. Not because the evidence is unclear, but because the accountability system itself is protected.
Standardized testing is no longer just an educational tool. It is embedded in a network of contracts, compliance requirements, testing vendors, consulting firms, and political interests backed by well-funded lobbying efforts. There are real financial and political incentives to preserve it, regardless of outcomes.
Spending more than $100 million annually on a system that continues to produce the same disparities while costing students millions of hours of learning time reinforces a governance model that rewards compliance and discourages challenge.
This climate — where profits and political interests are prioritized before children — did not emerge on its own. It has been shaped over time by a campaign finance system that rewards candidates who support policies centered on test-based school accountability.
When you see something wrong, something inhumane, don’t just say something. Do something. It’s time policymakers stop protecting harmful practices and confront the consequences of policies that continue to waste taxpayer dollars, diminish learning opportunities, and drive many of our most talented and experienced teachers from the profession. After twenty-five years, the ramifications of inaction are impossible to ignore.
Public education should not revolve around protecting systems, contracts, corporate profits, or political interests. It should revolve around children. Colorado students deserve a well-rounded education that values critical thinking, creativity, and civic engagement over excessive testing and data collection. For too long, corporate education groups and privatizers have robbed students of a meaningful education and carefree childhood. This accountability model offers the illusion of control while costing Colorado a future of empowered, well-educated leaders.
Jan Resseger, social justice warrior, strongly dissents from those who want to bring back the test-based accountability of No Child Left Behind and Race to the Top.
Defining schools by their achievement test scores is reductive. Of course we want our children to learn to read, to enjoy and understand literature, to master math, and to study history and the sciences, but a fixation on comparing school districts’ test scores blinds us to the human relations that constitute a classroom, to the social formation of children that happens at school, and to myriad other ways of thinking about what students are accomplishing at school. The temptation then is to define schoolteachers as producers of test scores and forget about all the other ways they help our children learn and grow.
Because test scores provide a simple, universal measure, we grab onto it and give it more weight than all the other factors we can’t so easily measure. Kevin Welner, a professor of education policy at the University of Colorado and director of the National Education Policy Center identifies family income, a factor entirely outside of school, as the most significant variable affecting a school district’s aggregate test scores: “Those of us who work in or with schools never question the enormous impact that a teacher or school can have on a student. But this essential truth coexists with another truth: that differences between schools account for a relatively small portion of measured outcome differences. That is, opportunity gaps in the U.S arise primarily outside of schools. This should not be a surprise. Poverty, concentrated poverty, and racialized poverty are pervasive features of America. School improvement efforts cannot directly help children and their families overcome decades of policies that perpetuate systemic racism and economical inequality.”
Last week, the NY Times’ Claire Cain Miller, Frencesca Paris and Sarah Mervosh reported on a major new demographic study documenting a widespread decline over the past decade in U.S. students’ standardized test scores: “Something troubling is happening in U.S. education. Almost everywhere in America, students are performing worse than their peers were 10 years ago… A report on the new data describes a decade-long ‘learning recession.’… Education experts say there is no single reason for the declines. But the timing provides some clues. Students’ test scores had been increasing since 1990—then abruptly stopped in the mid-2010s. That coincided with two events: an easing of federal school accountability under No Child Left Behind (NCLB), which was replaced in 2015, and the rise of smartphones, social media and personalized school laptops. The pandemic then accelerated learning declines, especially for the poorest students. Some pandemic effects have lingered. Student absenteeism, for example, remains higher than pre-pandemic… Test scores in low-income districts fell furthest, but affluent districts—the types of places families move to for the schools—also lost ground.”
The reporters do acknowledge a number of factors that may correlate with dropping scores, but they seem to lean toward blaming a lot of the problem on the end of No Child Left Behind. They are mistaken when they declare that the Every Student Succeeds Act (ESEA), NCLB’s replacement, ended test-based school accountability. In fact that 2015 law just made the states, not the federal government, agree to impose sanctions on the schools that had been unable significantly to raise test scores. The reporters quote Brian A. Jacob, a professor at the University of Michigan, who believes NCLB’s fading influence has been one cause of test score decline: “It was not a cure-all, but I think it really did improve student achievement… There’s evidence that school accountability does change behaviors of teachers and administrators and probably parents and students.”
A prominent retired professor of education, Diane Ravitch pushed back immediately on what she understood as the bias of the recent NY Times article: “I reject the claim that scores have stagnated because of the easing of No Child Left Behind-Race to the Top pressures. Sure, they increased the pressure on students, teachers, and principals, but their negative effects undermined the quality of education. Picking the right bubble on a standardized test became the goal of education. Campbell’s Law says that when a measure becomes the goal, it loses its value as a measure. Social scientist Donald Campbell wrote that ‘the more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor.’ “
Ravitch names a number of experts who have evaluated the damage wrought by the No Child Left Behind Act’s strategy: to punish schools and teachers who, supposedly, weren’t working hard enough to make all students reach test-score proficiency by 2014. The most prominent is Daniel Koretz, the Harvard University expert on standardized testing, who, in 2017, published The Testing Charade: Pretending to Make Schools Better. Koretz not only explains Campbell’s Law, but he shows how the pressure of test-based accountability corrupted what happened public schools across the country when the federal government threatened mandatory closure, or mandatory privatization or charterization of so-called “failing schools.”
Koretz reminds us that in places where test scores did rise under No Child Left Behind, it may not have reflected students’ academic growth. Test score gains were in many places artificially produced through test prep, the narrowing of the school curriculum, and even cheating: “Cheating—by teachers and administrators, not by students—is one of the simplest ways to inflate scores, and if you aren’t caught, it’s the most dependable.” (The Testing Charade, p. 73) His book covers the tragic Atlanta cheating scandal, and other examples when teachers read the tests in advance and prepared students to answer specific questions. Koretz describes various kinds of test prep coaching and drilling that were widespread in the NCLB era. And, “(Teachers) reported that they reduced—sometimes very substantially—the amount of time devoted to teaching science, which was not tested, in order to make additional time for prepping kids in math and reading.” (The Testing Charade, pp. 95-96)
Last week’s NY Times report on the possible causes of an overall drop in test scores over the recent decade also names two other possible causes. First, a decade ago, as schools began to provide laptops or electronic tablets to their students for online learning, students’ widespread dependence on their smartphones also became epidemic: “Something happened globally around the same time: the proliferation of devices, at home and in school. Nearly half of American teenagers now say they are online ‘almost constantly,’ compared with just under a quarter who said that a decade ago, according to Pew Research Center.” Due to the proliferation of devices, our classrooms operate differently, and our children are doing less reading of books for study and enjoyment.
Second, the reporters, explain, there was massive and well documented learning loss during the COVID pandemic: “Immediately after the pandemic, there was hope that students would recover quickly. The new data shows that scores inched upwards in reading last year, and have climbed more steadily in math since 2022. But it has been nowhere near enough to make up for lost ground…. The biggest losses have been among the lowest-achieving students.”
I have never heard anyone who has been able to trace the extent of long term damage during COVID, when students’ schools were closed and many children were left while their parents were at work to learn remotely on computers. Chronic absence has been a greater problem since COVID, and something schools have struggled to overcome. No one has been able to assess how long COVID will keep affecting children who were preschoolers and young elementary students back in 2019.
Finally there is one other big factor that could also be related to falling test scores over time: states have been perpetually reducing funding for public schools. According to the most recent research from the Albert Shanker Institute: “There are 42 states (including the District of Columbia) that devote a smaller share of their economies to their K-12 schools than they did before the 2007-2009 recession. This seems to be a permanent disinvestment in public education.” “(U)nequal opportunity is (also) universal in the U.S. In all states, higher-poverty districts are funded less adequately than lower-poverty districts… We find that 37 percent of white students attend districts with negative adequacy gaps, compared with 75 percent of African American students and 62 percent of Hispanic students. In other words, African American students are about twice as likely as their white peers to attend school in a district with below-adequate funding, while Hispanic students are almost 70 percent more likely to do so, and Native American… students are 50 percent more likely. Similarly, African American students are over 3 times more likely than white students to attend chronically underfunded districts….” These economic factors are likely to have affected students’ learning over time.
Our society will not be able to address our economic, social, and educational injustices through No Child Left Behind-style, test-based public school accountability.
This week, a report by the Education Scorecard, led by Sean Reardon at the Stanford group; Thomas Kane at the Center for Education Policy Research at Harvard; and Douglas Staiger at Dartmouth proclaimed that we are in a decade-long “learning recession.” It found that 83% of state reading scores declined from 2015 to 2025.
While I respect the Scorecard’s skills in compiling test score patterns, due to my time as an academic historian, an education researcher, and an inner city teacher, who witnessed the extreme harm done to students by the No Child Left Act of 2001 and the 2010 Race to the Top, I must challenge many of the conclusions that are being drawn from the test score patterns that Reardon, Kane, Staiger, and their partners present.
Kane then claimed that accountability-driven mandates due to the NCLB and the RttT produced gains that “‘may be one of the most important social policy successes of the last half-century that nobody knows about.’” That statement has been refuted by numerous studies including RAND’s research which concluded that the failure of attempts to improve learning through high-stakes testing added to the proof, “that one does not fatten a hog by weighing it.”I believe the test-driven teacher evaluations that Kane pushed were the most destructive education policy that I’ve ever heard of, and were a major factor in undermining teaching background information and reading for comprehension.Their test results patterns, I argue, actually support the opposite of the defense of NCLB and the RttT; it was the full implementation of high stakes testing, not the rejection of those failed policies, that was one of the top two causes of the sharp decline in literacy.
On the other hand, I agree that a main reason for the decline is the failure to manage social media, and that chronic absenteeism is a major factor.
But, first, I want to explain the political reasons why reading outcomes in the Tulsa Public Schools (TPS), and the Oklahoma City Public School System (OKCPS) fell so far. Secondly, I want to help defuse the “blame game,” and push back against the ramping up of unfair criticism of urban schools that is likely to get worse.
TPS students had gained only 3.8 years of learning over five years. Moreover, the OKCPS students only gained 4.4 years.
The TPS had had better schools than Oklahoma City, and we repeatedly visited Tulsa to learn from them. But, in 2010 they received a Gates Foundation grant for evaluating teachers, that Kane and Staiger helped create. Then, I frequently visited Tulsa and listened to both teachers and frustrated consultants as they complained about the damage being done to teaching and learning. Not surprisingly, it became much harder to recruit or retain teachers.
Also, data from American Enterprise Institute’s Nat Malkus showed that the TPS’s chronic absenteeism rate was 48.2%, compared to the nation’s 31.9% chronic absenteeism rate for similar schools.
But, before Oklahoma City’s educators in high-challenge schools are blamed, the extreme segregation they face must be taken into account. Oklahoma County has 14 school districts. along with magnet, charter, and private schools. School choice resulted in neighborhood schools with intense concentrations of students from extreme, generational poverty, who have endured multiple traumas (known as ACEs), thus driving down the OKCPS’s test scores.
Consequently, in 2015, suburban and exurban schools Edmond, Mustang, Moore, and Yukon were ranked higher than the national average by 1.6; .6; 1; and .8 years. By 2024, their scores declined by the same or by lower rates as similar national schools. So, it’s hard to make the case that the lack of teacher accountability, as opposed to segregation by choice, drove those drops in reading.
At the risk of sounding too nerdy, the historian in me needs to recall the chronologies for test score gains and decreases. I argue that the most meaningful reading metric is the 8th grade NAEP, which had been improving incrementally from 255 in 1971, to 263 in 2012, before it fell to 260 in 2020, and to 256 in 2023.
Both my experiences in the classroom, and the reading of the data, support the narrative that it took a while for the destructive policies of both interconnected reforms to be put in place, but when that happened, both laws drove meaningful learning down.
On the other hand, some claim that the reversal of the most punitive parts of RttT caused that decline. But those changes didn’t occur until 2015, after 8th grade reading scores were already in decline. Even so, in Oklahoma, the conservative Oklahoma Council of Public Affairs (OCPA) blamed State Superintendent Joy Hofmeister for the drop in state reading scores because she ended the practice that made us second in the nation in retentions.
Getting back to today’s national discussion about literacy, one data-driven scholar, Brian Jacobs, was cited for supporting NCLB despite its problematic features. He said, “It was not a cure-all, but I think it really did improve student achievement.”
But, if you follow the link to his research, it concludes, “Our results suggest that NCLB had no impact on reading achievement for 4th or 8th graders.” And it gives virtually no evidence that it didn’t undermine learning about science, history, arts, and music.
Reading the news coverage of the Education Scorecard brings me back to three sets of memories. During the early 1990’s, our school superintendent bragged about implementing the Reagan administration’s A Nation at Risk. So many of my students who grew up in that era would thank me for teaching in a meaningful manner, and then complain that they had previously been “robbed of an education” by its testing.
Secondly, at the turn of the century, I repeatedly talked with smart, sincere data experts about methodological problems when using their metrics for real world policies, as opposed to economic theory. I repeatedly heard the reply that their job was to show that data-driven accountability can improve teaching. If I’m right, they would say, they would run some more controls (presumably after the policies were in place). But it wasn’t their job to predict what will happen if those policies are adopted.
Thirdly, as the RttT was implemented, my students from the poorest elementary and middle schools would repeatedly thank me for showing them respect by teaching them in a meaningful manner. And, they kept volunteering that they had been “robbed of an education.”
It is also important to remember that the majority of OKCPS students are Hispanic, and remember that the OKCPS probably would have collapsed if it had not been for immigration. Now, when ICE is terrorizing immigrants, we must come together in support of our threatened students in order to reduce its contribution to chronic absenteeism.
Moreover, I don’t recall talking to a parent who doesn’t see the need to help young people control, and not be controlled, by their digital devices.
And I almost never talk to a parent, a student, or an educator who doesn’t want to cut back on high-stakes testing and test prep.
So, I agree we need to take the Education Scorecard seriously, but we should use it as a diagnostic tool to help us come together for the team efforts required for bringing back the joy of reading.
For instance, I agree with Elaine Allensworth, the executive director of the Chicago Consortium on School Research, who responded to the Scorecard saying we should not panic, but “We need to really start asking questions about what we can do to support students so they feel engaged in school.”
John Thompson, retired teacher and historian in Oklahoma, was stunned by some survey results released about parents’ opinions on education. He took a deep dive, read the raw data, and discovered that the survey was conducted by ExcelInEd, Jeb Bush’s organization. Excel promotes high-stakes accountability for public schools but no accountability whatsoever for voucher schools, which they also promote.
ExcelinEd has familiar game plan: they use inaccurate NAEP statistics to defame public schools, demand more accountability to crush the morale of principals, teachers, and parents, then insist that vouchers and charters are the way forward. As Josh Cowen showed in his book The Privateers, voucher schools get far worse results than public schools, and numerous studies have shown that charter schools are usually no better than public schools and often much worse.
Thompson writes:
Patricia Levesque, the executive director of ExcelinEd, recently wrote a commentary about a survey of 500 Oklahoma parents, claiming that more than 80% of them want “a state testing and accountability system to measure student achievement, and they expect honesty and accuracy about their children’s grade level performance.”
So, I took a dive into the survey. My reading of it was very different than Levesque’s.
In some ways, the survey she described is consistent with the Education Department parent survey that State Superintendent Lindel Fields released. But Levesque’s interpretation of the results was very different than Fields’ analysis of the state’s parent feedback.
The survey Levesque cited found that 74% of parents want a pay raise for teachers, and another 74% say we spend too little on education. Her study found that 80% of parents were very or somewhat satisfied with their school but, for some reason, it adds, “While overall positive, this fails to hit the common 95% satisfaction sought in commercial endeavors.”
While 78% of parents support retention by 3rd grade of students who don’t read on “level,” parents estimate that about 83% students read at or above grade level; and 78% are confident in the way their schools teach reading.
FYI, in 1998, 80% of Oklahoma 8th graders read at that level, but now about 59% do. My reading of the research, and classroom experience, attributes the subsequent decline to the way that No Child Left Behind and the Race to the Top undermined the teaching of History, Science, Arts, and of the background knowledge that is essential for reading comprehension; huge funding cuts; COVID; and Ryan Walters; as well as the rise of social media.
Yes, 83% of the survey are supportive of student testing, which is no surprise. But, the study doesn’t dig into the difference between testing for tracking student progress, as opposed to high-stakes testing. After all, there is great support for testing for diagnostic purposes, as opposed to the reward-and-punish testing that has been rampant since the NCLB was enacted.
Conversely, the Education Departments’ parent survey seems to be calling for schools to tackle the crucial issues that they were forced to ignore, as districts invested in high-stakes test-prep.
When Superintendent Fields explained that the results of statewide surveys of educators and parents, informed the budget priorities he is seeking. Superintendent Fields reported, “Early literacy, support systems to improve behavior and mental health resources and teacher recruitment and retention are among the top three concerns for all groups surveyed.”
The parents survey included repeated calls for teaching critical thinking skills, and media literacy; identifying misinformation; and early grade emphasis on literacy.
It explained that parents “highlighted the importance of both academic and life skills, emphasizing the need for students to be well-prepared for real-world challenges.”
Parents said that misinformation is very prevalent, and children need to be taught how to tell fact from fiction. They understand that learning how to be critical consumers of information is “literally the foundation of a successful life.” They know that social media and A.I. can make kids “susceptible to conspiracy theories and propaganda.”
What I didn’t see in the parents’ responses was calls for data-driven accountability; online, as opposed to personal tutoring for 3rd graders; or simple “miracles.”
What I saw was a desire to return to personal connections. I saw goals that would require more support for educators, as well as requiring cooperation with social workers, health providers, and mentors that are necessary for preparing children for a full life in the 21st century.
When I wrote a history of public schools in the 20th century (Left Back: A Century of Failed School Reforms), I couldn’t help but notice a consistent pattern: an infatuation with fads and panaceas, not by teachers but by pundits and education professors.
Teachers struggled with large class sizes, obsolete textbooks, and low pay, but the buzz was all too often focused on the latest magical reform. At one extreme was militaristic discipline, at the other was the romantic idea of letting children learn when they wanted and whatever they wanted to. Phonics or whole language? Interest or effort?
Every reform had some truth in it, but the extremes must have been very frustrating to teachers. There is no single method that’s just right for every child all the time.
The latest fad is Ed-tech, the belief that children will learn more and more efficiently if they spend a large part of their time on a computer.
My views were influenced by something I read in 1984. The cover story of Forbes was about “The Coming Revolution in Education.” The stories in the issue was about the promise of technology. Curiously, the magazine’s technology editor wrote a dissent. In 1984 Forbes published an article about the promise of computers in the schools. He wrote: “The computer is a tool, like a hammer or a wrench, not a philosophers’ stone. What kind of transformation will computers generate in kids? Just as likely as producing far more intelligent kids is the possibility that you will create a group of kids fixated on screens — television, videogame or computer.” He predicted that “in the end it is the poor who will be chained to the computer; the rich will get teachers.”
For the past few decades, Ed-tech has been the miracle elixir that will solve all problems..
But now, writes Jennifer Berkshire, there is a backlash against Ed-tech among parents and teachers.
They may have realized that the most fervent promoters of Ed-tech are vendors of Ed-tech products.
Stories about parents rebelling against big tech are everywhere right now. They’re sick of the screens, the hoovering up of their children’s data, and they view AI and its rapid incursion into schools as a menace, not a ‘co-pilot’ for their kids’ education. This is a positive development, in my humble opinion, especially since the backlash against the tech takeover of schools crosses partisan lines. Meanwhile, pundits and hot takers are weighing in, declaring the era of edtech, not just a failure, but the cause of our failing schools.
Which raises a not insignificant question. Now that everyone who is anyone agrees that handing schools over to Silicon Valley was big and costly mistake, how did the nation’s teachers and students end up on the receiving end of this experiment in the first place? And here is where our story grows murky, dear reader. In fact, if you’re old enough to remember the absolute mania around ‘personalized learning’ that took hold during the Obama era, count yourself as fortunate. Because lots of the same influential, not to mention handsomely compensated, folks who were churning out ‘reports’about our factory-era schools 15 minutes ago, suddenly seemed cursed by failing memories.
The not-so-wayback-machine
If you need a refresher to summon forth the 2010-era ed tech frenzy, proceed directly to Audrey Watters’ unforgettable write-up: “The 100 Worst Ed-Tech Debacles of the Decade.” Watters’ has moved on to a new newsletter and AI refusal, but her once lonely voice as the ‘Cassandra’ of education technology remains as essential as ever. Her tally of “ed-tech failures and fuck-ups and flawed ideas” is studded with now tarnished silver bullets that promised to transform our factory-era schools into futuristic tech centers, making a pretty penny in the process: AltSchool, inBloom, Rocketship, Amplify, DreamBox, Summit… The names have changed or been forgotten but the throughline—a fundamental misunderstanding of schools and teaching combined with the promise of hefty returns—remains constant.
My own introduction to the ed tech hustle came back in 2015. Jeb Bush’s annual convening for his group, the Foundation for Excellence in Education, or FEE, to use its comically apt acronym, came to Boston. To which I said, ‘sign me up!’ Always an early adapter (see, for example, school vouchers in Florida), FEE was unabashedly pro technology, as I wrote in a story for the Baffler.
It’s one of FEE’s articles of faith that the solutions to our great educational dilemmas are a mere click away—if, that is, the schools and the self-interested dullards who run them would just accept the limitless possibilities of technology. Of course, these gadgets don’t come cheap. And this means that, like virtually all the other innovations touted by our postideological savants of education reform, the vision of a tech-empowered American student body calls for driving down our spending on teaching (labor costs account for the lion’s share of the $600 billion spent on public education in the United States each year) and pumping up our spending on gizmos.
In virtually every session I attended, someone would relate a story about a device that was working education miracles, followed by a familiar lament: if only the teachers, or their unions, or the education ‘blob’ would get out of the way.
False profits
In a recent piece for Fortune, reporter Sasha Rogelberg offers an interesting origin story for the tech takeover of public education. And you don’t need to read past the title to get where she’s going: ‘American schools weren’t broken until Silicon Valley used a lie to convince them they were—now reading and math scores are plummeting.’ I’d make the header even clunkier and add ‘the education reform industry’ to the mix. While the push to get tech into classrooms predates Obama-era education reform (check out Watters’ fantastic history of personalized learning, Teaching Machines, for the extended play version), it was the reformers’ zeal, when married to Silicon Valley’s profit optimization, would prove so irresistible.
In the last hundred years, the base of the United States economy has shifted from industry to knowledge—but the average American classroom operates in much the same way it always has: one teacher, up to thirty same-age students, four walls. This report from StudentsFirst argues that this one-size-fits-all approach doesn’t cut it in the modern world, in which mastery of higher-order knowledge and skills ought to matter more than time spent in front of a teacher—and that what we need is competency-based education. This approach, also known as the “personalized model,” is characterized by advancing students through school based on what they know and can do, using assessments to give them timely, differentiated support, made easier by the introduction of learning technology.
StudentsFirst, the hard-charging school reform org started by Michelle Rhee, has since been eaten by 50CAN, which now advocates for school vouchers, but the fare they offered up was standard. Indeed, here’s a fun activity for you. Revisit any prominent reform group, individual, or cause and you will find the same argument about our factory-era schools, followed, inevitably, by the same sales pitch for a tech-centric solution.
Race to the Top, Obama’s signature education reform initiative, didn’t just bribe cash-strapped states to overhaul their teacher evaluation systems. It also ‘encouraged’ states to shift their standardized tests online. And Arne Duncan and Obama’s Department of Education actively courted the tech industry, encouraging them to think of schools as a space ripe for disruption. “Many of today’s young people will be working at jobs that don’t currently exist,” warned the XQ Institute, the reform org started by Steve Jobs’ widow, Laurene Powell Jobs. Today Powell Jobs presides over the Atlantic, where new panic pieces regarding young, tech addled dumb dumbs appear seemingly every day.
Warning signs
My obsessive interest in the intersection of education and politics began back in 2012, when my adopted home state of Massachusetts came down with a serious—and well-funded—case of education reform fever. At a time when red states were crushing the collective bargaining rights of teachers (Wisconsin, anyone?), I was struck by how often reform-minded Democrats ended up repurposing the right’s anti-union, anti-teacher, anti-public-school rhetoric for their own righteous cause. Ed tech sat right smack in the center of this queasy juncture—beloved by liberal reformers, ensorcelled by press releases promising higher test scores, and conservatives who liked the idea of spending less on schools by replacing teachers with machines.
Recall, if you will, Rocketship charter schools, whose innovative blended learning model caused the test scores of its students—almost all poor and minority—to go up like a rocket. Richard Whitmire’s fawning 2013 book, On the Rocketship: How Top Charter Schools Are Pushing the Envelope, is a veritable time capsule of the era. Unlike the fusty Model-T schools of yore, Rocketship schools were tech forward. Students spent a chunk of each day in so-called Learning Labs, taking, retaking or practicing taking tests, a practice that had a measurable impact, especially since 50 percent of teachers’ pay was tied to test scores ascending. All that clicking also translated into dollar signs, wrote Whitmire. “A major cost-saving solution was for students to spend significant time working on laptops in large groups supervised by noncertified, lower-paid “instructional lab specialists.”
Rocketship has since fallen back to earth, in part because of stellar reporting like this from Anya Kamenetz, documenting the chain’s less savory practices. But it’s hard to overstate just how excited the reform world was about this stuff. Next time you hear an edu-pundit bemoaning the take over of kindergarten classrooms by big tech, remember that Rocketship got there first. “[K]indergarten teachers are spending less time making letter sounds,” co-founder Preston Smith told Kamenetz. And reformers couldn’t get enough.
Whodunit?
Investigative reporter Amy Littlefield has an intriguing-sounding new book out in which she uses the model of an Agatha Christie novel to suss out who killed abortion rights in the US. I imagine that taking a similar approach to the question of how big tech conquered public education would end up in Murder on the Orient Express territory. That’s the classic Christie whodunit in which everyone on the train ends up having ‘dunit.’ These days, there is a comical effort underway by reformers to distance themselves from the tech takeover—what train? I’ve never been on a train! But the idea that Silicon Valley had the cure for all that ailed the nation’s public schools was absolutely central to Obama-era education reform.
I’d locate the zenith of the reform/tech love affair in 2017 when New Schools Venture Fund, a reform org that funds all of the other orgs, laid down a challenge, or rather, a big bet. At its annual summit, backed by a who’s who of tech funders—Gates, Zuckerberg, Walton, NSVF called for big philanthropy to bet big on tech-based personalized learning. “The world has changed dramatically … and our schools have struggled to keep up,” then CEO Stacey Childress warned the crowd. But not all the news was bad. Going all in on education innovation would also pay off handsomely, claimed NSVF, producing an estimated 200 to 500 percent return on investment. And lest parents, teachers and students failed to adequately appreciate the various reimaginings they were in for, NSVF had an answer for that too: a $200 million ad campaign to “foster understanding and demand.”
As I was preparing to type a sentence about how poorly NSVF’s “Big Bet on the Future of American Education” has aged, a press release popped up in my inbox, announcing that Netflix founder Reed Hastings is joining forces with Democrats for Education Reform or DFER. “Just as Netflix replaced a one-size-fits-all broadcast model with something more personal and responsive, Hastings believes public education can make the same leap.”
AI is a once-in-a-thousand-year shift, and what happens in K-12 is at the center of it. The schools that figure out how to combine individualized software with teachers focused on social-emotional development are going to unlock something we’ve never seen before.
Of course, transforming “a school system in desperate need of reinvention” the way that Hastings reinvented home entertainment will require “governance innovation and political will.” No doubt an ad campaign is in the works too. And convincing education ‘consumers’ that individualized software = school is going to be a tough sell as the Great Big Tech Backlash accelerates.