The Evidence Article, and Why It Has to Be Careful
Ask a room of education policy veterans what No Child Left Behind accomplished and the answers will arrive with unusual confidence and unusual disagreement. One side will point to rising proficiency rates and closing gaps, and will describe the statute as the first federal law that made schools answerable for the children they had been overlooking. The other side will point to narrowed curricula, gaming, and the vast machinery of sanctions that punished schools without fixing them, and will describe the statute as an accountability system that measured everything except what mattered. Both sides can produce citations. Both sides can produce charts. That is exactly what makes this article hard to write and necessary to read. The literature on this statute is not a shortage of evidence. It is a surplus of partial evidence, drawn from instruments that disagree with one another, interpreted by analysts whose conclusions were often formed before the data arrived.

The One Test for this article is a practical one. A reader who finishes it should be able to state what the evidence shows about test scores, achievement gaps, curriculum, and school improvement under the 2001 accountability regime, to distinguish the findings that replicate from the ones that do not, and to understand the measurement problem that makes several popular claims about this statute untestable. The test is deliberately pitched at the level of judgment rather than recall. Knowing a single number is not enough, because the literature rewards anyone who can quote one favorable figure while ignoring the instrument that produced it. What matters is knowing which numbers survive contact with an external yardstick and which do not, and why the distinction decides nearly every argument about this law.
The organization is built around five findings, each attributed to named research with its period and its method. The first finding covers national assessment trends, where mathematics scores for younger students rose across the period while reading gains were smaller and high school performance stayed flat, and where the hardest question is attribution, because the trends began before the statute was enacted. The second finding is the state test problem: state proficiency rates rose far faster than national assessment results in many states, and the divergence is the strongest available evidence that some of the apparent improvement reflected standard-setting rather than learning. The third finding covers subgroup effects, where research using the staggered introduction of accountability systems in the states, before and after the federal law, generally finds positive effects on mathematics achievement concentrated among lower-performing students and schools, with weaker effects in reading. The fourth finding is curriculum narrowing, where survey evidence documents increased instructional time in tested subjects and reduced time in untested ones, a real cost that honest discussion has to acknowledge. The fifth finding is the school improvement machinery: the choice and supplemental service provisions had very low take-up rates and the restructuring interventions were implemented inconsistently, so the consequence side of the statute is largely untested rather than proven ineffective.
Running through all five findings is the namable claim of this article, the two-ruler problem. The statute measured progress with instruments the measured parties controlled, which means the central empirical question about it cannot be answered from state data alone, and every credible finding in this literature depends on an external yardstick. A state that sets its own proficiency bar, administers its own test, and reports its own results has not measured learning in the same sense that a disinterested observer would. The federal assessment, administered to samples of students under uniform conditions, plays that external role, and the literature that survives is the literature that respects the difference. The statute’s defenders and its detractors both have moments when they prefer the ruler that flatters their case. The article’s job is to refuse the preference and apply the same standard to both.
A word about horizon is necessary before the substance begins. The law described here is the law as it stood in November 2014, the date on this article, and the evidence is the evidence available by that date. The accountability regime established by Public Law 107-110, with requirements phasing in from the 2002 to 2003 school year, was effectively unwound in most states by the waiver process that began with the administration’s offer of flexibility on September 23, 2011, and the first approvals on February 9, 2012, and the statute remained the law of the land until its replacement was signed on December 10, 2015. Nothing in this article is extended to any policy after its horizon, and no finding here is offered as guidance about what any legislature should do. The question is what the evidence showed about a particular statute over a particular period, and that question is large enough.
The Statute Being Measured
The effects this article evaluates flow from Public Law 107-110, signed in January 2002, the reauthorization of the Elementary and Secondary Education Act that carried the short title No Child Left Behind. The provisions under evaluation are described in full at no-child-left-behind-act-2001-guide, which this article treats as the companion document for what the law required, so that this article can concentrate on what the law produced. For the reader who needs the compressed version, the structure was this: annual testing in reading and mathematics in grades three through eight and once in high school; state-set proficiency standards defining what counted as adequate performance; adequate yearly progress determinations for every school and district, disaggregated by subgroups including racial and ethnic groups, students with disabilities, English learners, and economically disadvantaged students; a fixed trajectory requiring that all students reach proficiency not later than twelve years after the end of the 2001 to 2002 school year, which set the deadline in the 2013 to 2014 school year, as section 1111(b)(2)(F) of the reauthorized act provided; and an escalating sequence of consequences for schools that missed their targets, from public school choice, to supplemental educational services, to corrective action, to restructuring.
The underlying program predates the 2001 law by more than three decades. Title I of the Elementary and Secondary Education Act of 1965, described at elementary-secondary-education-act-1965, is the federal funding stream for schools serving concentrations of low-income students, and the 2001 statute was the eighth reauthorization of that 1965 law. The distinction matters for the evidence because it sets the baseline. Title I money, and the schools that received it, existed before the accountability regime, and some of the trends the evidence literature measures were in motion long before the testing mandates began. Any evaluation that treats 2002 as year zero of American school improvement is measuring the statute against a fiction.
Three design features of the statute determine what can and cannot be learned from the record. First, the statute set the target but let each state define the ruler. Proficiency was a state-defined threshold on a state-designed or state-chosen test, so the same word, proficient, meant different levels of demonstrated knowledge in different states, and in many cases the meaning changed within a single state across years. Second, the statute imposed the trajectory nationally while the measurement stayed local, which meant the federal government demanded convergence on a deadline while delegating the definition of arrival. Third, the statute’s consequence ladder applied to schools, not to states, which created an incentive structure in which state officials, who controlled the ruler, faced political pressure from the schools their rulers were failing. A state that held its standard high watched its schools cascade into sanctions; a state that eased its standard watched the same schools make their targets. The system did not formally punish the honest measurer, but the politics of the ladder did. These three features are the reason the two-ruler problem is not an abstract worry but the central fact of the evidence.
The timeline matters for interpreting every trend this article discusses. The statute was signed in January 2002, but its requirements phased in: accountability reporting began with the 2002 to 2003 school year, as the Congressional Research Service documented in report RL31284, and the adequate yearly progress machinery reached full operation over the following years. The waiver process that effectively unwound the regime in most states began with the administration’s offer of flexibility on September 23, 2011, and the first approvals came on February 9, 2012, for ten states, with New Mexico’s approval following on February 15. The statute remained the law of the land until its replacement was signed on December 10, 2015. The evidence window this article evaluates therefore runs roughly from the 2002 to 2003 school year through 2011, with some series extending further, and every trend attributed to the statute has to be read against that window rather than against the calendar year of enactment alone.
Two boundaries keep this article honest. The first is the period. The federal accountability regime ran from the 2002 to 2003 school year through the waiver years beginning in 2011 and 2012, and the evidence below is organized around that window. Anything a study measures outside that window is named as such. The second is the evidence standard. A score trend counts only when its assessment and period are named. A research claim counts only when its authors, publication and method are named. A figure that cannot be attributed to a source is dropped rather than softened. These are not decorative rules. They are the reason the findings below differ from the versions that circulate in political argument, where state proficiency rates are routinely cited as evidence of learning, pre-existing trends are routinely attributed to the statute, and low take-up of a service is routinely treated as proof that consequences failed. Each of those moves is a recurring error in public discussion of this statute, and each is addressed where it appears.
How to Read the Evidence That Follows
Three distinctions will keep the findings below from being misread, and they are worth stating before the evidence begins because each one corrects a specific habit of public argument about this statute.
The first distinction is between the federal layer and the state systems underneath it. No Child Left Behind did not invent test-based accountability. Through the 1990s, many states built their own systems, with annual testing, public reporting and consequences of varying severity, and they did so at different times and with different intensity. The federal statute imposed a common federal layer on top of that uneven landscape in 2002: every state had to test every year in the specified grades, report by subgroup and move all students toward proficiency on a fixed timeline. When this article speaks of the statute’s effects, it means the effects of that federal layer and its interaction with the state systems beneath it, not the effects of accountability as an idea. The distinction matters because the strongest research designs in the literature exploit exactly this layering. States that already had strong systems in the 1990s received a weaker federal dose than states that had little or nothing, and comparing their trajectories is what lets researchers estimate something closer to a cause than a national before-and-after comparison can deliver.
The second distinction is between a trend and an effect. A trend is what happened. An effect is what the statute caused. Finding one reports trends on the national assessment, and those trends are real whether or not the statute caused them. Findings two through five report research that tries to separate the statute’s contribution from the trend, and they use different tools: external yardsticks, staggered adoption, surveys and implementation audits. Public argument constantly slides between the two, treating a rise that began in 1973 as proof that a 2002 statute worked, or treating flat high school scores as proof that it did nothing. Both slides are fallacies, and this article marks them where they occur.
The third distinction is between absolute achievement and movement against a target. The national assessment measures the first: how much students know, on a fixed scale. Adequate yearly progress measured the second: whether a school’s scores moved enough against a bar that rose every year. A school could teach better every year and still fail, if the bar rose faster than its scores, and a school could teach worse and still pass, if its state’s bar moved down. The late years of the regime, when nearly half the nation’s schools missed their targets, are unintelligible without this distinction, and much public commentary about those years reads the miss rates as evidence that learning collapsed. It did not. The bar moved. Keep these three distinctions in hand and the five findings read as one coherent account rather than five separate claims.
One more piece of context will help the findings land. Before the federal mapping studies appeared, the public debate about the statute ran almost entirely on state data, and the two camps talked past each other with numbers that were never comparable. The October 2009 mapping study changed what counted as evidence by putting the fifty state rulers on one scale for the first time, and the August 2011 follow-up extended the comparison across a longer window. After those studies, any honest account had to reckon with the divergence, and the accounts that did not were visibly selecting their ruler. This article is written on the far side of that shift. Every score claim below names its assessment, because the mapping studies made it impossible to pretend the assessment does not matter.
The Two-Ruler Problem
Every credible finding about this statute depends on an external yardstick, because the statute’s own measurement system was controlled by the parties being measured. That sentence deserves a slower treatment, because it is the claim on which the article’s five findings are organized, and because misunderstanding it is the most common error in public discussion of the law.
Consider what a state proficiency rate actually is. It is the share of tested students scoring above a cut score on a particular examination, where the state chose the examination, the state set the cut score, and the state decided which students counted as tested and how their scores were reported. None of those choices is illegitimate in itself; a state must set a cut score somewhere, and must choose a test. But every one of those choices is also a policy lever that moves the reported number without moving student knowledge. Lower the cut score and the proficiency rate rises. Exclude more students from the denominator and the rate rises. Narrow the tested content to what is most easily taught and the rate rises. The reported number is therefore a joint product of two things, how much students know and how the state chose to measure them, and the number alone cannot tell a reader how much of its movement came from each source. This is not a special suspicion about the honesty of state officials. It is a mathematical property of any system in which the measured party controls the instrument.
The external yardstick is the National Assessment of Educational Progress, the federal assessment program that administers uniform tests to representative samples of students. The long-term trend component of that program, which tracks the performance of nine, thirteen, and seventeen year olds on assessments designed to remain comparable across decades, provides the longest continuous external series. The main assessment component, which tests grades four, eight, and twelve against frameworks that change more often, provides shorter series with more curricular detail. Both components share the property that matters: the states being evaluated do not set the cut scores, choose the test, or control the reporting. When the external yardstick and the state ruler agree, the finding is strengthened. When they disagree, the disagreement is itself the finding, because it tells the reader that something other than learning moved the state number.
The distinction between the two rulers runs through every section that follows, and it is stated wherever scores are discussed, as the brief for this article requires. State proficiency rates are evidence of what the state system reported, not evidence of what students learned, unless corroborated by an external measure. Federal assessment trends are evidence of what students demonstrated under uniform conditions, with the limitation that no national test captures everything a state test captures. Research designs that compare states that adopted accountability systems at different times, the staggered-adoption literature, are valuable precisely because they use the external yardstick while exploiting variation in the timing of the treatment, which lets the researcher separate the effect of accountability from the national trends that were already in motion. The findings that replicate are the findings that survive the external ruler. The findings that collapse are the ones that depended on the state ruler alone.
The stakes of getting the evidence right extend beyond the reputation of a single statute. The 2001 law was the largest federal experiment in school accountability ever attempted, and its record is the primary evidence base for every subsequent debate about testing, consequences, and the federal role in schooling. If the record is misread, through the state ruler alone, or through the attribution error, or through the confusion of take-up with effectiveness, the misreading propagates into the next design. An accountability system built on the belief that the consequence ladder worked will repeat a ladder that was never tested; a system built on the belief that nothing worked will discard the subgroup reporting that was the law’s most durable innovation. Scrupulousness about this evidence is therefore not an academic virtue. It is the precondition for learning anything from the largest policy experiment in the field’s modern history, and the article treats it as such.
The principle generalizes beyond education, and it is worth stating in its general form because it is the reason the two-ruler problem keeps recurring in public policy. When a measure becomes the basis for rewards or sanctions, the people being measured gain an incentive to improve the measure rather than the thing it measures, and any component of the measure they control will move first. Test scores can be raised by teaching or by redefining proficiency. Reported crime can fall because crime fell or because reporting standards changed. Hospital ratings can improve because care improved or because the sickest patients were turned away. The pattern is the same in every case: the measured party optimizes the controllable margin. The defense is also the same in every case, and it has two parts. First, separate the definition of the measure from the control of the measured, so that no party can move its own ruler. Second, check every self-reported improvement against an external yardstick that the measured party cannot touch. No Child Left Behind violated the first part by design, letting states define proficiency while facing a federal deadline, and the literature supplied the second part after the fact, through the national assessment and the mapping studies. A regime that built both parts in from the start would not need a decade of research to discover whether its ruler had moved.
The Instruments, and What Each One Can See
To apply the two-ruler discipline, a reader needs to know what each instrument actually measures, because the instruments differ in design as well as in control. The federal assessment operates two components. The long-term trend component tests representative samples of students at ages nine, thirteen, and seventeen, using assessment designs held constant across decades so that a score from the nineteen seventies can be compared with a score from the two thousands. It reports by age rather than by grade, a choice that keeps the series comparable even as retention and enrollment patterns shift, and it covers reading back to 1971 and mathematics back to 1973. Because it samples rather than testing every child, it cannot report results for individual schools or districts, which is why it never served as the statute’s accountability instrument. What it offers instead is the longest continuous external series in American education, and therefore the background against which any claim about the statute’s effects has to be judged.
The main assessment component tests grades four, eight, and twelve against frameworks that are updated periodically to reflect curricular change, which makes its series shorter but more instructionally detailed. It reports state-level results for grades four and eight, so researchers can compare states with one another on a common ruler, and it uses achievement levels, Basic, Proficient, and Advanced, set by panels of educators and citizens rather than by any legislature. The grade twelve results arrive less often and cover the nation as a whole. The distinction between the two components matters for this article: the long-term trend supplies the decades-long background, and the main assessment supplies the state-by-state comparisons that the staggered-adoption studies use.
The state instruments were built for a different purpose. The statute required annual testing in reading and mathematics in grades three through eight and once in high school, and each state chose or built its own tests, set its own proficiency cut scores through its own board or agency, and defined its own rules for which students counted. Adequate yearly progress was then computed from those state instruments: the share of tested students above the state’s cut score, reported for the school as a whole and for each qualifying subgroup. The design gave the federal government annual, universal data at the school level, which the federal assessment could never provide, and it gave each state control over the definition of success, which the federal assessment would never permit. The tradeoff is the two-ruler problem in miniature: the instrument with the finest grain is the instrument the measured parties control.
The scale on which these instruments report also deserves a reader’s attention, because the mapping studies express their findings in scale points. The federal assessment reports scores on a scale that runs from zero to five hundred, and the achievement levels mark regions of that scale: Basic denotes partial mastery of the tested domain, Proficient denotes solid academic performance, and Advanced denotes superior performance. When the mapping studies report that nearly all state proficiency standards sat at or below the federal Basic level, the translation is that the typical state definition of proficient corresponded, on the external ruler, to partial mastery or less. When they report that standards varied by at least sixty scale points, the translation is that the word proficient covered an enormous range of demonstrated knowledge, from well below Basic in the most lenient states to far higher in the most demanding ones. The numbers are not decorations. They are the reason the state ruler cannot be trusted on its own.
One more design detail matters for interpreting the federal numbers. The federal assessment tests samples, not censuses: a representative selection of students in each age group or grade, drawn so that the results describe the population without testing every child. Sampling keeps the burden on schools low and the conditions uniform, and it means the reported scores come with margins of error, which is why the article reports some changes as not significantly different rather than as zero. A change that falls within the margin of error is not evidence of stasis; it is evidence that the data cannot distinguish the change from no change. The statute’s state tests, by contrast, tested every child, which gave them precision at the school level at the cost of the comparability the federal samples preserved.
No instrument sees everything. The long-term trend sees decades but cannot distinguish states. The main assessment distinguishes states at grades four and eight but not schools. The state tests see every child every year but entangle learning with measurement choices. The research designs that survive are the ones that match the question to the instrument: national trends for the background, state comparisons on the external ruler for the causal estimates, and the mapping studies for the relationship between the two rulers. A claim that uses the wrong instrument for its question, state proficiency rates as evidence of national learning, or a single national trend as evidence of a law’s causal effect, fails before the evidence is even examined.
Finding One: What the National Assessments Show
The longest available external series, the long-term trend of the federal assessment, tells a story of slow improvement over decades, concentrated in particular subjects and ages, with the statute’s period sitting near the end of a much longer movement. The National Center for Education Statistics published the series covering reading from 1971 through 2008 and mathematics from 1973 through 2008 in April 2009 as NCES 2009-479. In reading, average scores for nine year olds rose twelve points over the full period, scores for thirteen year olds rose four points, and scores for seventeen year olds were not significantly different at the end of the period than at the beginning. In mathematics, average scores for nine year olds rose twenty four points and scores for thirteen year olds rose fifteen points, while seventeen year olds again showed no significant difference. The gains accumulated across nearly four decades, and the mathematics gains were substantially larger than the reading gains at every age where gains occurred.
The main assessment, which tests grades rather than ages, adds shorter series with a closer view of the statute’s period. The grade twelve results released in November 2010 showed mathematics rising from a scale score of one hundred fifty in 2005 to one hundred fifty three in 2009, a small gain at the end of high school. Reading at grade twelve rose from two hundred eighty six to two hundred eighty eight over the same years, but the 2009 reading score remained below the nineteen ninety two average of two hundred ninety two, which means the small gain recorded in those years did not even recover the ground lost across the preceding decade and a half. High school performance, on both components of the federal assessment, was essentially flat across the statute’s lifetime.
The attribution question is where care becomes indispensable. The statute’s accountability requirements began phasing in with the 2002 to 2003 school year, and the mathematics gains visible in the long-term trend had been accumulating since the early nineteen seventies. A reader who sees rising fourth grade mathematics scores in the two thousand and credits the statute has confused the tail of a decades-long trend with the effect of a law enacted near its end. This is the recurring error the brief warns against: attributing pre-existing trends to the statute. The long-term trend for nine year olds in mathematics was rising through the nineteen eighties and nineteen nineties, before any state was operating under the federal accountability regime, and the reading gains for nine year olds likewise accumulated over decades. Nothing about the shape of the national series identifies the statute as the cause of its slope.
That does not mean the statute caused nothing. It means the national series alone cannot say what it caused, which is why the staggered-adoption research, covered in the third finding, exists. The national trend is the background against which the accountability effect has to be isolated, and any study that treats the background as the effect is committing the oldest error in program evaluation, crediting a treatment with the work of time. The defensible reading of the first finding is therefore twofold. Scores on the federal assessment rose in mathematics for younger students across the period while reading gains were smaller and high school performance was largely flat, and the movements that began decades before enactment cannot be attributed to the statute without a research design that separates the law’s contribution from the trend. The rating for this finding in the evidence table is contested, not because the direction of the scores is in doubt, which it is not, but because the attribution of the direction to the statute is the contested question.
Why can the statute not take credit for the long rise in mathematics scores?
Because the rise began decades before the law existed. The long-term trend shows nine year old mathematics scores climbing through the nineteen eighties and nineteen nineties, before the 2002 to 2003 phase-in. Crediting the statute with the whole slope is the attribution error this literature exists to prevent.
No feature of the national trend identifies the statute as the cause of its slope, since the law arrived near the end of a forty-year movement. One more feature of the national record deserves emphasis before moving on, because it sharpens the contrast the next finding develops. The federal assessment and the state assessments were measuring during the same years, often the same students, and their verdicts on the period diverged. Mathematics for younger students improved on both instruments, but the magnitude of improvement reported by the states frequently exceeded what the external yardstick showed, and in reading, where the external yardstick showed small gains or none, many states reported substantial progress. The divergence is not a quirk of a few states. It is the pattern the mapping studies documented systematically, and it is the subject of the second finding.
The shape of the national trend deserves one more look, because the attribution error has a specific form that recurs in public discussion. The error is not simply crediting the statute with gains it did not cause. It is treating the level of scores in the statute’s period as the statute’s achievement, as if the decades of improvement before 2002 were a flat baseline from which the law’s work began. The long-term trend shows no such flat baseline. Nine year old mathematics scores were already climbing through the nineteen eighties, a period when no state operated anything resembling the federal accountability regime, and the climb continued through the nineteen nineties at roughly the same pace. A reader who draws a line at 2002 and measures everything after it is measuring the continuation of a slope, not the creation of one. The honest question is whether the slope changed, whether the rate of improvement accelerated after the accountability phase-in began, and the national series does not show a visible break. That absence does not prove the statute had no effect; national aggregates can conceal effects concentrated among particular students, which is exactly what the staggered-adoption studies find. But it does mean the national trend, taken alone, cannot carry the causal claim, and anyone who cites it as proof that the law worked has mistaken the background for the effect.
The grade twelve window sharpens the point. The main assessment’s grade twelve results for 2005 to 2009 cover the heart of the statute’s mature period, when the testing machinery was fully built and the consequence ladder was operating. Mathematics gained three scale points and reading gained two, with reading still below its nineteen ninety two level. If the accountability regime were transforming achievement, the end of high school is where the transformation would have to appear, because twelve years of schooling under the regime should compound. The flatness at the end of the pipeline does not refute the grade four mathematics findings; younger students and older students can genuinely move differently, and the staggered-adoption literature measures the grades where the treatment was most intense. But the flatness does discipline the rhetoric. A law described by its advocates in transformative terms produced, on the external ruler, modest early-grade mathematics gains, no reading gains, and a high school record indistinguishable from stasis. The evidence supports the modest claim and not the transformative one, which is why the authors of the strongest positive study framed their results against the moonshot rhetoric rather than beneath it.
The attribution problem is the finding, not a footnote to it. A line that rises from 1973 to 2008 was rising in 1973, in 1983, in 1993 and in 2001, and the segment of that line that falls inside the statute’s window has to be compared against the trajectory the line was already on. This is the first reason the two-ruler problem matters: even with an honest external yardstick, a before-and-after comparison confounds the statute with everything else that changed across those years, from demographics to funding to earlier state reforms.
The flat scores of seventeen-year-olds deserve a slower look, because they are the part of the trend that most resists a simple story. One reading is discouraging: whatever the elementary schools were doing right, its benefits did not survive to the end of high school. But several less discouraging readings are available, and the data cannot choose among them. The long-term trend is a repeated cross-section, not a longitudinal study: the nine-year-olds of 2008 are not the seventeen-year-olds of 2008, and the seventeen-year-olds of 2008 were nine in 2000, before the accountability regime began. Their elementary years fell mostly outside the statute’s window, so their flat scores say little about what the regime did to younger students. A second possibility is fadeout, the well-documented pattern in which early achievement gains diminish as students move through later grades without continued intervention. A third is institutional: high schools are departmentalized, students see each teacher for less time, and the levers available to elementary schools, above all the reallocation of the whole instructional day, are weaker when the day is divided among subjects and teachers. A fourth is selection: seventeen-year-olds still in school in any given year are a different population than the nine-year-olds, because the students most at risk have already left. None of these explanations is proven by the trend data, and this article does not choose among them. The point of listing them is to show how little a flat line at age seventeen settles. It rules out the claim that the era produced large, compounding, system-wide gains. It does not rule out real gains for younger students, and finding three will show that those gains appear in the research designs built to detect them.
None of this means the national assessment trends are useless. They establish the boundaries of what any honest claim must fit inside. A claim that the statute produced dramatic nationwide gains has to explain the flat high school trend. A claim that it produced nothing has to explain the mathematics gains for younger students, even allowing that some of those gains belong to the pre-statute trajectory. And a claim built on state test scores has to survive comparison with the federal assessment, which is where finding two begins. The national trends are the outer frame. The state-federal divergence is the inner complication.
One more caution about reading the point changes. A twelve-point or twenty-four-point move on the long-term trend scale is a real move, but the scale is not calibrated in grade levels and this article does not convert it into any. Translating scale points into years of learning requires assumptions the assessment’s designers do not endorse, and the translation is one of the quieter ways that advocacy literature inflates modest findings into dramatic ones.
Finding Two: The State Test Problem
If the first finding is about what an external ruler shows, the second finding is about what happens when the measured parties control the ruler. During the statute’s lifetime, state proficiency rates rose in many states at a pace that the federal assessment did not corroborate, and the divergence is the strongest available evidence that some of the apparent improvement reflected standard-setting rather than learning. This is the finding that most complicates any simple verdict on the law, because it means the numbers the statute itself produced, the adequate yearly progress determinations and the proficiency rates that filled state report cards, cannot be read as a record of learning without an external check.
The mapping studies, published by the National Center for Education Statistics, provide the systematic evidence. The first, “Mapping State Proficiency Standards Onto NAEP Scales: 2005 to 2007,” published as NCES 2010-456 in October 2009, placed each state’s proficiency standard on the federal assessment’s scale so that standards could be compared across states and across time. Two results matter. First, in between one third and one half of the states, the proficiency standard was lower in 2007 than it had been in 2005, which means the bar was being eased even as reported proficiency was rising. Second, the federal assessment agreed with the direction of the state’s reported progress in only two fifths to three fifths of the states, which means that in a substantial minority of states, the state test said progress while the external ruler said otherwise. A standard that falls while the reported pass rate rises is not evidence of learning. It is evidence of redefinition.
The second mapping study, “Variation and Change in State Standards for Reading and Mathematics, 2005 to 2009,” published as NCES 2011-458 and released on August 10, 2011, extended the analysis and deepened the concern. For the 2007 to 2009 window, the proficiency gains that states reported were not corroborated by the federal assessment in at least half of the comparison states, twenty two to twenty six of forty, and in most of those states, seventeen to twenty two, the state’s own results painted a more positive picture than the external measure. The standards themselves varied enormously, by at least sixty points on the federal assessment’s scale, and nearly all of them mapped at or below the federal assessment’s Basic achievement level, which means the typical state definition of proficient would not have counted as proficient on the external ruler. The two reports must not have their windows collapsed: the first covers 2005 to 2007 and the second covers 2007 to 2009, and the corroboration figures belong to their respective windows. Taken together, they show a system in which the reported progress depended heavily on where each state set its bar, and in which the bars were often set low and sometimes lowered further.
The mechanism behind the divergence is not mysterious, and the article states it plainly rather than insinuating bad faith. State officials faced a statutory trajectory that demanded universal proficiency by the 2013 to 2014 school year, a deadline that grew less plausible with each passing year as the adequate yearly progress targets climbed. The officials who set cut scores, chose tests, and defined which students counted were the same officials whose schools were being judged by the resulting numbers, and the political pressure from the consequence ladder, described in the fifth finding, fell on them as well as on the schools. A state that maintained a demanding standard watched its campuses fail in growing numbers; a state that eased its standard watched the failures recede. The system rewarded leniency and punished rigor, and the mapping studies show that a substantial number of states responded to the incentive. The deeper point is not that officials behaved cynically but that the statute gave them no reason to hold the line and every reason to move it.
There is a subtler version of the same point. Even in states where standards were not lowered, the state test and the national assessment measure overlapping but different things, and a state can improve on its own test by aligning instruction narrowly to that test’s content without improving the broader domain the national assessment samples. That is test-specific preparation rather than standard-setting, and the national assessment is less sensitive to it because its content is not the content any state teaches to. The distinction matters for finding four, where the curriculum evidence shows instructional time shifting toward tested subjects, which is a mechanism that would produce exactly this pattern: state scores up, national scores flatter. The two findings reinforce each other without either one proving the other. A state that wanted its proficiency rates to rise had two ways to get there: teach more effectively or define proficiency down.
The consequence for the evidence is direct. State proficiency rates cannot be cited as evidence of learning under this statute without an external check, and when the check is applied, a significant share of the reported improvement disappears. The recurring error the brief identifies, citing state proficiency rates as evidence of learning, is not a minor technicality. It is the error that sustains the most inflated claims about the statute’s achievement effects, and it is the error the two-ruler problem exists to prevent. The rating for this finding is replicates, because the mapping studies, across two reports and two windows, converge on the same conclusion: the state ruler and the external ruler told different stories, and the difference was driven in significant part by where the states set their bars.
What made the state and national results diverge?
The states controlled the instruments that produced their results. The mapping studies show proficiency bars lowered between 2005 and 2007 in up to half the states, and reported gains from 2007 to 2009 uncorroborated by the external assessment in at least half of the comparison states. When the measured party sets the bar, reported progress and learning can move independently.
The practical meaning of the mapping results becomes clearer with a concrete translation. When the second mapping study reports that state proficiency standards varied by at least sixty points on the federal assessment’s scale, it is describing a range so wide that the same student, with the same knowledge, could be labeled proficient in one state and far below proficient in another. Sixty points on that scale spans multiple years of typical learning growth; it is not a rounding difference or a technical artifact. And when the study reports that nearly all state standards mapped at or below the federal Basic level, it is saying that the typical state definition of proficient corresponded, on the external ruler, to partial mastery of the domain or less. A parent reading a state report card that announced rising proficiency would have had no way to know that the bar being cleared sat at partial mastery, or that the bar had been lowered since the previous report. The information the public needed, where the bar sat and whether it moved, was precisely the information the reporting system did not provide.
The states where the two rulers agreed deserve their due, because the divergence finding is sometimes read as saying that no state made real progress, which the studies do not say. In the first mapping study’s window, the national assessment agreed with the direction of state-reported progress in two fifths to three fifths of the states, which means that in a substantial share of the country the reported gains were corroborated by the external ruler. Those are the states where the accountability pressure most plausibly produced learning rather than redefinition, and they are disproportionately the states that held their standards steady. The divergence finding is therefore not a claim that all state progress was illusory. It is a claim that the illusory and the genuine were mixed together in the national totals, and that only the external ruler can separate them. A reader who wants to know whether a particular state’s gains were real should look at that state’s mapping, not at the national average of state reports.
The political economy behind the movement is worth stating carefully, because it explains why the pattern was so widespread without requiring any theory of misconduct. Cut scores in most states were set by state boards of education or by agencies acting under board authority, bodies whose members watched the same failure counts the public watched. As the adequate yearly progress targets climbed toward the 2013 to 2014 deadline, the number of schools missing their targets rose, and each newly failing school generated local news coverage, parental anger, and political pressure on the officials who had set the bar. Lowering the cut score, or holding it steady while the test was made more forgiving, relieved that pressure immediately and visibly. No official had to act cynically for this to happen; it was enough for officials to resolve genuine technical doubts, about where exactly the cut score should sit, in the direction that reduced the political pain. The statute supplied the pain and the relief mechanism together: it demanded universal proficiency on a fixed timetable while delegating the definition of proficiency to the officials who would be blamed for the failures. The mapping studies are the record of how that delegation was used.
There is also a federalism irony worth naming. The statute’s defenders described it as a civil rights measure that would end the soft bigotry of low expectations, and the subgroup reporting requirement was genuinely aimed at that goal. But the standard-setting power the statute left with the states created a different, quieter form of low expectations: a state could report rising proficiency for every subgroup while defining proficiency down for all of them. The subgroup gaps on the state’s own ruler could narrow on paper while the gaps on the national assessment stood still. This is not a hypothetical. It is what the combination of subgroup reporting and state-defined standards made possible, and it is the reason the two-ruler problem is not a technical footnote but the central moral of the evidence.
There is a final implication that reaches beyond the statute’s own period. The divergence between the rulers means that the entire public record of the law’s achievement, the proficiency rates published in state report cards, the adequate yearly progress determinations reported to the federal government, the trend lines shown in legislative hearings, was systematically more optimistic than the external evidence warranted. Policymakers who relied on the state numbers were steering by a compass that pointed partly toward the magnet of political convenience. The lesson is not that state assessments are useless; they remain the only instruments that test every child annually, and no accountability system can operate without them. The lesson is that any system which lets the measured party define success must build in an external check, or the reported progress will drift from the learning it claims to measure. The statute did not build in that check, and the mapping studies are the audit it never commissioned.
Finding Three: The Staggered-Adoption Research
The first two findings describe trends. The third finding describes causes, or more precisely, the closest the literature comes to causes. Because the federal statute imposed accountability on all states at once, its own introduction offers no comparison group: there is no America without the law against which to measure the America with it. But the states did not all arrive at accountability at the same time. In the years before the federal law, some states had already built consequential accountability systems, with public reporting, ratings, and sanctions attached to test results, while others had weak systems or none. Researchers exploited that staggered introduction, comparing the achievement trajectories of early-adopting and late-adopting states on the external yardstick, the federal assessment, before and after the federal law extended accountability everywhere. The design is the strongest in the literature for isolating the effect of accountability from the national trends, because it holds the background constant and varies only the timing of the treatment.
The earliest study in this line, by Martin Carnoy and Susanna Loeb, published in Educational Evaluation and Policy Analysis in 2002, volume 24, issue 4, pages 305 to 331, constructed a zero to five index of state accountability strength and examined gains on the federal assessment’s eighth grade mathematics test from 1996 to 2000, before the federal law. States with stronger accountability systems gained significantly more than states with weaker systems over that period. The study also tested two of the feared side effects of accountability pressure and found no evidence for either: stronger accountability did not increase student retention in grade or reduce high school completion. The Carnoy and Loeb result is a pre-federal-law finding, which makes it evidence about accountability systems in general rather than about the 2001 statute in particular, but it established the design and the basic pattern that later studies would refine.
Eric Hanushek and Margaret Raymond, writing in the Journal of Policy Analysis and Management in 2005, volume 24, issue 2, pages 297 to 327, extended the analysis to the accountability systems states operated through the nineteen nineties and reached a compatible conclusion: state accountability had a clear positive impact on the growth of federal assessment scores. Their study added an important qualification on the distribution of the gains. The positive effects did not narrow the gap between Black and White students, though they did narrow the gap between Hispanic and White students, a distinction the literature has preserved because it resists the simpler story that accountability helps all disadvantaged groups equally. Like Carnoy and Loeb, Hanushek and Raymond tested a feared side effect and found no evidence for it: there was no sign that accountability pressure increased the placement of students into special education as a way of removing low scorers from the tested population. The null finding on special education placement matters because it addresses one of the most common gaming hypotheses directly, and the evidence did not support it.
A different study in the same period reached a null result that the literature has had to accommodate. Jaekyung Lee and Kenneth Wong, in the American Educational Research Journal in 2004, volume 41, pages 797 to 832, examined whether state accountability systems narrowed achievement gaps and found essentially no effect on racial gaps and largely insignificant effects on socioeconomic gaps. The Lee and Wong null finding is important, and this article states it plainly rather than burying it, because it constrains what the positive findings can claim. Accountability systems of the nineteen nineties variety appear to have raised average achievement, particularly in mathematics, without reliably narrowing the gaps between groups. The memo for this article is explicit that Lee and Wong must not be cited for the positive concentrated-gains pattern; that pattern belongs to the study that follows, and the distinction is preserved here.
A reader who notices that Lee and Wong found null effects while Dee and Jacob found positive ones may wonder whether the literature contradicts itself. It does not, and the reason is that the two studies asked different questions about different periods with different designs. Lee and Wong studied the state accountability systems of the 1990s and asked whether those systems narrowed demographic gaps. Dee and Jacob studied the arrival of the federal statute on top of those state systems and asked whether it raised average achievement, finding that the gains concentrated among lower performers. A system can raise the floor without closing gaps, which is exactly what happens when lower-performing groups gain but higher-performing groups gain as well, or when the gains accrue within groups in ways that do not change the distance between group averages. The two results are compatible because gap-closing is a stricter test than floor-raising, and the literature passes the looser test while failing the stricter one. Treating Lee and Wong as a refutation of Dee and Jacob, or Dee and Jacob as a refutation of Lee and Wong, is the kind of selective reading this article exists to prevent.
The most influential study of the federal law itself, by Thomas Dee and Brian Jacob, published in the Journal of Policy Analysis and Management in 2011, volume 30, issue 3, pages 418 to 446, applied a comparative interrupted time series design to the introduction of the federal accountability regime, exploiting the same staggered-adoption logic: states that already had consequential accountability before the federal law served as the comparison group for states that received the treatment only when the federal law arrived. The results were the strongest positive findings in the literature for the statute’s achievement effects, and they were also the most qualified. By 2007, the study found large and statistically significant gains in fourth grade mathematics attributable to the law, with effect sizes around twenty two to twenty three hundredths of a standard deviation, and the gains at grade four were described as almost universal across the states newly subject to the regime. In eighth grade mathematics, the study found moderate gains targeted at low-performing groups rather than spread across the distribution. In reading, at both grades four and eight, the study found no significant gains. The authors themselves framed the results against what they called the moonshot rhetoric of the law’s advocates, the promise that accountability would transform achievement across the board, and the findings supported a smaller claim: the law moved mathematics achievement for younger and lower-performing students, and it did not move reading. The date matters because the study is sometimes cited to 2009, which was the year its working-paper version appeared as National Bureau of Economic Research working paper 15531, and the published version is the one this article cites.
Taken together, the staggered-adoption literature supports the brief’s summary: positive effects on mathematics achievement concentrated among lower-performing students and schools, with weaker effects in reading. The rating for this finding is replicates, because the pattern, mathematics gains concentrated at the bottom of the distribution, reading effects weak or null, appears across studies with different periods, different designs, and different comparison groups. Two qualifications keep the finding honest. First, the gains are modest in absolute terms; an effect size of roughly one fifth of a standard deviation is real and policy-relevant, but it is not a transformation. Second, the gap-narrowing claim remains the weakest part of the achievement story: Hanushek and Raymond found no narrowing of the Black-White gap, Lee and Wong found null results on racial gaps, and the Dee and Jacob gains, while concentrated among lower performers, were not framed by the authors as a gap-closing triumph. The literature supports the claim that accountability raised mathematics achievement for the students who had been achieving least. It does not support the claim that it closed the gaps between groups.
What does staggered adoption let researchers do that a single national launch would not?
It creates a comparison group from timing alone, which separates the law’s effect from the national trend. Early-adopting states serve as the baseline for states that received the treatment only in 2002 to 2003. Comparing their trajectories on the external federal assessment isolates what accountability added.
A reader unfamiliar with the research designs may wonder what a comparative interrupted time series actually does, and the intuition is worth stating because the design carries most of the article’s causal weight. Imagine plotting each state’s mathematics trajectory on the federal assessment across the years, then marking the year the federal accountability regime began. In states that already had strong accountability, nothing changes at the mark; they were already treated. In states that had weak or no accountability, the mark is the moment the treatment arrives. If the newly treated states bend upward relative to their own prior trend, and relative to the already-treated states that show no bend, the difference in the bends is the estimated effect of the law. The design is comparative because it uses the early adopters as the comparison, interrupted because it looks for a break at the moment of introduction, and time series because it uses the full trajectory rather than a single before-and-after snapshot. Its strength is that it does not require the researcher to believe the national trend stood still; the trend is visible in the comparison states and subtracted out. Its weakness is the one all such designs share: if something else changed for the newly treated states at the same moment, the design cannot separate it from the law. The authors tested the obvious alternatives and found the pattern robust, which is why the study anchors the literature.
The size of the effect deserves a plain-language translation. An effect size of twenty two to twenty three hundredths of a standard deviation means that the average student in the treated states scored about a fifth of the typical spread of scores higher than the comparison implies they otherwise would have. In the distribution of test scores, that is a noticeable shift, visible in the data and unlikely to be chance, but it is not a transformation of the distribution. It does not move a typical student from the middle of the pack to the top, and it does not erase the distance between high-achieving and low-achieving groups. Policy researchers sometimes translate such effects into months of learning, but those translations depend on assumptions about typical annual growth that vary by grade and subject, so this article reports the effect in the units the authors used. The honest characterization is the one the authors offered: large enough to be statistically solid and policy-relevant, concentrated where the need was greatest, and far smaller than the rhetoric that accompanied the law’s passage.
The subject asymmetry, mathematics gains without reading gains, also deserves more than a passing note, because it recurs across every study in this finding and therefore demands an explanation even if the studies cannot supply a proven one. Three hypotheses fit the evidence. The first is curricular: mathematics content is sequential and its tested formats are well understood, so additional instructional time and test-aligned preparation transfer more directly to mathematics scores than to reading scores. The second is about the nature of the skill: reading comprehension draws on vocabulary and background knowledge accumulated over years, which makes it less responsive to the additional drill that accountability pressure tends to produce. The third is about where the time went: the survey evidence shows large reported increases in English language arts time, but if those minutes went to tested formats and test preparation rather than to wide reading and rich content, the national assessment, which samples broad comprehension, would not register them. None of the three is established by the research designs, and they may all operate at once. What is established is the asymmetry itself, and any account of the statute’s effects that does not name the subject is not describing this literature.
One final qualifier on the designs. Carnoy and Loeb’s accountability index measured systems with real stakes attached to performance, and Hanushek and Raymond’s positive results likewise came from the consequential systems of the 1990s wave. The literature’s positive findings are findings about accountability with consequences, not about testing alone. A testing regime without stakes, or with stakes as weakly implemented as finding five documents, would not be expected to reproduce these effects, and the distinction between the pressure half of the statute and the consequence half runs through the entire verdict.
The second is about the mechanism the designs cannot see. The staggered-adoption studies estimate the effect of accountability pressure on the external national assessment, which is exactly the right instrument for avoiding the two-ruler problem, but they cannot distinguish learning from test-specific preparation that happens to transfer to the national assessment’s content. The gains are real in the sense that they appear on an assessment no state controlled. Whether they represent deeper learning or better-aligned preparation is a question the designs cannot answer, and honest summaries say so.
Finding Four: The Narrowing of the Curriculum
The fourth finding is the cost side of the ledger, and the brief requires that it be reported as a documented cost rather than as a rhetorical point. The statute tested reading and mathematics annually, attached consequences to performance in those subjects, and left other subjects outside the accountability machinery. Schools responded as the incentive structure predicted: instructional time moved toward the tested subjects and away from the untested ones. The evidence for the shift comes from surveys of school districts, and the article reports what the surveys found while stating their limitation, which is that surveys record what administrators report, not what time-use diaries would measure.
The principal source is the Center on Education Policy’s report “Choices, Changes, and Challenges: Curriculum and Instruction in the NCLB Era,” written by Jennifer McMurrer and published in July 2007. The Center surveyed three hundred forty nine school districts and asked about changes in instructional time at the elementary level since the 2001 to 2002 school year. Sixty two percent of districts reported increasing instructional time in English language arts and mathematics. Among the districts that increased time, the average increase was forty six percent in English language arts and thirty seven percent in mathematics, a combined average increase of forty two percent for the two subjects together. Forty four percent of districts reported cutting time from at least one other subject, by an average of thirty one percent. The subjects losing time were the ones the statute did not test: science, social studies, art, music, physical education, and in some districts recess. A follow-up report from the Center in February 2008 extended the analysis. The figures are self-reports by district administrators, not measurements from classroom observation, and the article states that limitation wherever the figures appear. What administrators say they did is not the same as what happened minute by minute in classrooms, but when nearly two thirds of districts report the same directional shift, the pattern is not a reporting artifact.
The narrowing finding is the one the statute’s supporters have generally acknowledged as a real cost, and the acknowledgment matters for the article’s neutrality. There is no need to inflate the finding into a claim that the arts were destroyed or that science instruction vanished; the survey documents a reallocation of time, substantial in magnitude, concentrated at the elementary level, and directionally consistent with the incentives the statute created. Nor is there any need to minimize it into a claim that schools were merely focusing on the basics; a forty two percent average increase in tested-subject time and a thirty one percent average cut elsewhere is a large reallocation, and it fell on the subjects that give schooling its breadth. The honest summary is that the statute bought its mathematics gains in part with time taken from everything it did not test, and that the purchase was visible in the survey data within a few years of the law’s implementation.
The curriculum finding interacts with the first three in ways that complicate every verdict. The instructional shift toward tested subjects is a plausible mechanism for the mathematics gains in finding three: more minutes in mathematics, particularly for lower-performing students who received the most additional instruction, would be expected to raise mathematics scores. It is also a plausible mechanism for part of the state-national divergence in finding two: instruction aligned tightly to a state test’s content raises state scores more than it raises national assessment scores. And it is the reason the reading null results in finding three deserve a second look. If districts increased English language arts time by a reported forty-six percent on average and reading scores on the national assessment still did not move, that is evidence about the limits of additional instructional time as a lever, or about the difference between the reading the tests reward and the reading the national assessment measures. The findings do not sit in separate boxes. They are one system, and the curriculum evidence is the connective tissue.
Two further points keep the finding in proportion. First, the narrowing was concentrated at the elementary level, where the testing grades clustered and where the school day offered the most discretionary time to reallocate. High schools, with their departmental structures and graduation requirements, shifted less. Second, the survey evidence cannot say whether the reallocated time was well used; more minutes in mathematics is a cost to the untested subjects regardless of whether the additional mathematics instruction was effective, but the achievement findings suggest at least some of it translated into learning. The cost is real whether or not the purchase was worthwhile, and the article treats it as a cost, documented and quantified, rather than as an argument.
Was the time shift in elementary classrooms measured or merely reported?
Reported, and the distinction matters for how much weight the figures carry. The Center on Education Policy surveyed three hundred forty nine districts and recorded what administrators said about instructional time, not what observers timed in classrooms. When sixty two percent of districts describe the same reallocation, the direction of the shift is established even if the minutes are approximate.
The distribution of the losses matters as much as their size. The subjects that lost time were not random; they were the subjects furthest from the accountability machinery. Science and social studies, the two subjects with the strongest claims to tested-adjacent status, lost time at the elementary level in many districts, which is notable because both subjects would later gain their own testing constituencies. Art, music, and physical education lost time in districts where the schedule had the least slack, and in some districts recess itself was shortened or eliminated to make room for additional tested-subject blocks. The pattern reveals the logic of the reallocation with unusual clarity: time flowed toward whatever the statute measured and away from whatever it did not, regardless of the subject’s educational value. A defender of the law might reply that the reallocation was rational, that struggling readers need reading time more than they need art, and that the mathematics gains vindicate the tradeoff. A detractor might reply that the tradeoff sacrificed the breadth that makes schooling worthwhile for gains the external ruler shows to be modest. The survey does not adjudicate that dispute. It establishes that the tradeoff happened, at scale, within a few years.
Time taken from science and social studies is time taken from the content knowledge that supports reading comprehension itself, which creates a paradox inside the statute’s own theory: the pressure to raise reading scores may have reduced the very instruction, in history and science, that builds the background knowledge struggling readers need.
The concentration at the elementary level also deserves emphasis, because it explains why the narrowing debate played out differently across the grades. Elementary schools typically assign one teacher to a class for most of the day, which gives the school enormous discretion over how the hours are divided; moving forty minutes from science to reading is an administrative decision, not a staffing one. High schools, organized by departments with subject-certified teachers and graduation requirements set by the state, could not reallocate as freely, and the testing footprint in high school was lighter in any case. The result is that the curriculum cost fell most heavily on the youngest students, the same students for whom the mathematics gains were largest. Whether that coincidence strengthens the case for the tradeoff or deepens the concern about it depends on the reader’s values, and the article does not resolve it. It records the coincidence because any honest accounting of the statute has to hold the gain and the cost in the same frame.
The survey’s design centered on the elementary level, and that focus limits what can be said about middle and high schools. The elementary grades are where the instructional day is most flexible, which is also why the reallocation was largest there and why the achievement effects in finding three concentrate in the early grades. Whether secondary schools narrowed their curricula in the same way is less well documented in the corpus this article draws on, and the flat high school trends in finding one are consistent with a regime whose pressure and whose time reallocation both bit hardest in the early grades. An honest evidence base says what it knows about elementary schools and marks the secondary picture as thinner, rather than projecting the elementary finding across all grades.
The February 2008 follow-up from the Center on Education Policy extended the survey analysis and confirmed the direction of the findings, adding detail on how districts described their instructional changes. The follow-up matters for the rating because replication across reports, even from the same research organization, strengthens the claim that the pattern was not an artifact of a single questionnaire. The limitation remains what it was: administrators reporting on their districts, not observers timing classrooms. But the convergence of the two reports, the size of the reported shifts, and the consistency of the direction across nearly two thirds of districts make the narrowing the most solidly documented cost in the literature on this statute.
Finding Five: The Machinery That Was Never Really Tested
The fifth finding concerns the consequence side of the statute: the ladder of interventions that was supposed to follow when schools missed their targets. The brief’s summary is that the choice and supplemental service provisions had very low take-up rates and the restructuring interventions were implemented inconsistently, so the consequence side of the statute is largely untested rather than proven ineffective. The distinction between untested and ineffective is the most important conceptual point in this section, because public discussion routinely treats the two as the same, and the evidence supports only the first.
The choice provision allowed students in schools identified as needing improvement to transfer to a higher-performing public school in the district, with transportation provided. In the 2003 to 2004 school year, fewer than one percent of eligible students transferred, according to the Government Accountability Office’s December 2004 report GAO-05-7. The Congressional Research Service, in report RL31487, put the figures for the same school year at one percent for choice and eleven point three percent for supplemental services. The Department of Education’s own figures for the following year showed choice participation at about one percent of eligible students in 2004 to 2005. The reasons for the low take-up were structural rather than mysterious: receiving schools had limited capacity, districts were slow to notify parents of their options, transportation logistics were difficult, and many parents preferred to keep their children in their neighborhood schools. A provision that almost nobody used cannot tell the researcher whether school choice, as a mechanism, improves achievement. It can only tell the researcher that this particular provision, as implemented, did not operate at a scale where its effects could be measured.
Supplemental educational services, the free tutoring the statute offered to low-income students in schools missing their targets, reached more students than choice but still a small fraction of those eligible. The Government Accountability Office’s August 2006 report GAO-06-758 found that participation rose from twelve percent of eligible students in 2003 to 2004 to nineteen percent in 2004 to 2005, and that the number of recipients grew from about one hundred seventeen thousand in 2002 to 2003 to about four hundred thirty thousand in 2004 to 2005. About twenty percent of the districts required to offer the services had zero recipients, which means the provision existed on paper in those districts but reached no child. The memo for this article cautions against mixing the GAO’s twelve percent figure for 2003 to 2004 with the Department’s seventeen percent figure for the same year; the article keeps the same-year pairs together, GAO with GAO, and does not blend sources. Even on the more generous trajectory, four out of five eligible students were not receiving the services, and the research literature on the tutoring itself, as distinct from the take-up question, remained thin within the article’s horizon. The provision was underused, unevenly implemented, and consequently uninformative about whether tutoring at scale would have worked.
Restructuring, the most severe rung of the ladder, required fundamental changes in schools that missed their targets for five consecutive years, chosen from a statutory menu of five options. The Government Accountability Office’s September 2007 report GAO-07-1036 examined the 2005 to 2006 school year, when two thousand seven hundred ninety Title I schools were in corrective action or restructuring, and found implementation that could fairly be called nominal. About six percent of the schools in corrective action had taken none of the required corrective actions. About forty percent of the schools in restructuring had taken none of the five statutory restructuring options, which means they carried the label without adopting any of the interventions the label was supposed to trigger. Forty two percent of the schools had not received all of the technical assistance the statute required the states and districts to provide. A sanction that two fifths of its targets ignore, and for which the required support fails to arrive in more than two fifths of cases, is not a tested intervention. It is a designation.
The adequate yearly progress trajectory that drove schools onto the ladder deserves its own statement, because the scale of the failure determinations shaped everything else. The statute required all students to reach proficiency not later than twelve years after the end of the 2001 to 2002 school year, which set the deadline in the 2013 to 2014 school year. As the targets climbed toward that deadline, the share of schools missing them rose. The Center on Education Policy’s report “AYP Results for 2010 to 2011,” published December 15, 2011, estimated that forty eight percent of the nation’s public schools, more than forty three thousand schools, missed adequate yearly progress in the 2010 to 2011 school year, up from thirty nine percent the year before. The forty eight percent figure was a six-year high in thirty five states; in twenty four states plus the District of Columbia, at least half of schools missed; in five states plus the District, at least three quarters missed. The article labels these as the Center’s estimates, which the Center itself noted could be revised by one or two percentage points, and it does not use the higher figure that circulated in public discussion, the prediction of eighty two percent attributed to the Secretary of Education, because that figure was a pre-data projection that the Center’s actual data contradicted. The trajectory was doing exactly what its critics predicted: as the deadline approached, the system labeled nearly half the nation’s schools as failing, including many schools whose students were learning at respectable levels but not at the pace the fixed trajectory demanded.
The conceptual payoff of this finding is the one the brief demands. Low take-up is not proof that consequences failed; it is proof that consequences, as designed, were barely tried. A choice provision used by fewer than one percent of eligible students, a tutoring provision reaching fewer than one in five, and a restructuring regime ignored by two fifths of its targets cannot support any conclusion about whether strong consequences would have improved achievement, because the strong version was never implemented. The consequence side of the statute is untested. That verdict cuts against both camps: it denies the defenders the claim that the sanctions drove the gains, since the sanctions barely operated, and it denies the detractors the claim that the sanctions were proven harmful, since what was barely implemented cannot have been proven anything. The rating in the evidence table is untested, and the article means the word literally.
Does low take-up prove that choice and tutoring failed?
No. Fewer than one percent of eligible students used the transfer option in 2003 to 2004, and supplemental services reached nineteen percent of eligible students in 2004 to 2005, which means neither provision operated at a scale where its effects could be measured. A mechanism that almost nobody used was not proven ineffective. The consequence ladder is untested, not refuted.
The reasons for the low take-up were structural, and they are worth enumerating because they explain why the provisions failed as mechanisms rather than merely as statistics. The choice provision required districts to notify parents of their transfer rights, to identify receiving schools with capacity, and to provide transportation, all on timelines set by the statute. In practice, notification often arrived late or in language parents found impenetrable; receiving schools, frequently the highest-performing schools in the district, had limited seats and little incentive to accept large numbers of transferring students; and transportation across a district’s attendance zones proved logistically difficult and expensive. Parents, for their part, often preferred the neighborhood school, with its familiar teachers and shorter commute, to a distant school whose higher test scores might reflect its students’ advantages rather than its teaching. None of these obstacles implies that school choice as an idea cannot work. They imply that this particular choice mechanism, layered onto existing attendance zones and capacity constraints, could not operate at the scale its drafters imagined.
The tutoring provision faced a parallel set of constraints. States had to approve providers, districts had to contract with them, and sessions had to be scheduled outside the regular school day, which meant evenings and weekends for the students most in need of rest and the families least able to arrange transportation. Provider capacity was thinnest in the rural districts and the most distressed urban districts, the places where eligible students were concentrated. The twenty percent of required districts with zero recipients were not necessarily defying the law; many were small districts with no approved provider willing to operate within their boundaries. The growth from one hundred seventeen thousand to four hundred thirty thousand recipients shows the machinery expanding, but the nineteen percent participation rate in 2004 to 2005 shows it expanding far more slowly than eligibility. A tutoring program that reaches one student in five is not a test of tutoring. It is a test of outreach, logistics, and market capacity, and it is those that failed.
And there is a structural conflict that runs beneath all of it. The districts responsible for informing parents about choice and tutoring, arranging the transfers, contracting with the providers and paying for the transportation were the same districts whose schools had been identified as failing. The statute assigned the implementation of the sanctions to the sanctioned. That arrangement did not require bad faith to fail. It required only the ordinary friction of institutions asked to administer their own punishment, and the take-up figures show how much friction there was.
Restructuring failed differently, as a problem of local capacity rather than of parent behavior. The five statutory options, which included reopening as a charter school, replacing staff, contracting with an outside manager, state takeover, and fundamental governance restructuring, each demanded actions that districts found difficult, expensive, or politically impossible. Replacing a school’s staff meant confronting labor contracts and community opposition; chartering or outside management meant surrendering local control; state takeover meant admitting that the district could not fix its own schools. The forty percent of restructuring schools that adopted none of the options were not necessarily ignoring the law out of defiance; many were going through the motions of planning while the political costs of any real option blocked action. The forty two percent that did not receive all required assistance point to the other side of the failure: states and districts lacked the personnel, the expertise, and in some cases the will to support several thousand simultaneous restructurings. A sanction is only as real as the system that implements it, and the system that was supposed to implement restructuring was the same system being sanctioned.
What the Evidence Does Not Cover
Five findings is a lot, and it is worth saying what they leave out, because the boundaries of an evidence base are themselves evidence about what the statute was designed to notice. The literature this article draws on is, with few exceptions, a literature about test scores. That is not an accident. The statute made test scores the measure of everything, so researchers studied what the statute measured, using the instruments the statute’s own data systems made available. The questions the statute did not ask are the questions the literature largely did not answer.
Graduation and completion are the most important omission. Carnoy and Loeb’s 2002 study found no effect of 1990s accountability systems on student retention or high school completion, which is a useful null result for the gaming question, but it is not a finding about whether the federal regime changed graduation rates, and the corpus this article draws on contains no such finding for the 2002 to 2011 period. A reader who wants to know whether accountability kept more students in school, or pushed more out, will not find the answer here, because the studies were not built to give it. The same holds for longer-run outcomes: college enrollment, earnings, any measure of what happened to the tested cohorts after they left school. The statute’s theory ended at proficiency, and so did most of the research.
School-level variation is a second omission. Every finding above is stated as an average or a national pattern, and averages hide the distribution. Some schools and districts used the accountability pressure well, focusing intervention time on struggling students and improving instruction in ways that the national assessment registered. Others gamed the system, narrowed the curriculum to test preparation, or ignored the requirements entirely. The mapping studies show states making different choices about their standards, and the implementation studies show districts making different choices about their obligations. A national verdict is therefore a summary of a varied reality, and any policymaker or educator reading it should remember that the statute’s effects were not uniform. The evidence table’s ratings describe what replicates across the studies, not what happened in any particular school.
The third omission is the waiver period itself. The accountability regime this article evaluates was effectively unwound in most states beginning with the waiver offer of September 23, 2011 and the first approvals of February 9, 2012. What replaced it, a patchwork of state-designed systems operating under federal flexibility, is outside this article’s scope, and the research literature on those replacement systems was still thin within this article’s horizon. The unwinding is therefore a boundary rather than a finding: the evidence describes a regime that ran from the 2002 to 2003 school year through the waiver years, and what came after belongs to the series’ account of the statute’s successor.
Naming these omissions is part of the scrupulousness the brief for this article demanded. An evidence article that pretended its five findings exhausted the subject would be committing the same selective-vision error it criticizes in political argument. The findings cover what the research measured well: national trends, the state-national divergence, the staggered-adoption effects, the curriculum reallocation and the unused consequences. What they do not cover, graduation, long-run outcomes, local variation and the post-waiver years, remains open, and any future evaluation of the era will need student-level longitudinal data and implementation measures that the studies of this period largely lacked.
The Five-Finding Evidence Table
The table gathers the article’s findings into the findable artifact the brief requires. Each row carries the outcome, the direction of the evidence, the period the evidence covers, the principal research with its date, and the rating: replicates, contested, or untested. The ratings apply to the finding as stated, not to the broader claims sometimes built on it.
| Outcome | Direction | Period | Principal research | Rating |
|---|---|---|---|---|
| National assessment scores in mathematics, ages nine and thirteen | Up | 1971 to 2008, long-term trend; 2005 to 2009, main assessment grade twelve | NCES 2009-479, April 2009; grade twelve release, November 2010 | Contested |
| National assessment scores in reading, ages nine and thirteen | Small gains at nine, minimal at thirteen | 1971 to 2008, long-term trend; 2005 to 2009, main assessment grade twelve | NCES 2009-479, April 2009; grade twelve release, November 2010 | Contested |
| High school performance on the federal assessment | Flat | 1971 to 2008, age seventeen; 2005 to 2009, grade twelve | NCES 2009-479, April 2009; grade twelve release, November 2010 | Replicates |
| State proficiency rates versus the federal assessment | State rates rose faster; divergence documented | 2005 to 2007; 2007 to 2009 | NCES 2010-456, October 2009; NCES 2011-458, August 2011 | Replicates |
| Mathematics achievement under accountability, concentrated among lower performers | Positive, concentrated at the bottom of the distribution | 1996 to 2000; nineteen nineties state systems; through 2007 | Carnoy and Loeb 2002; Hanushek and Raymond 2005; Dee and Jacob 2011 | Replicates |
| Reading achievement under accountability | Weak to null | 1996 to 2000; through 2007 | Carnoy and Loeb 2002; Dee and Jacob 2011 | Replicates |
| Achievement gaps between racial and socioeconomic groups under accountability | No reliable narrowing | Nineteen nineties state systems; 2004 study window | Hanushek and Raymond 2005; Lee and Wong 2004 | Contested |
| Instructional time in tested versus untested subjects | Tested subjects up, untested subjects down | 2001 to 2002 through 2007 | Center on Education Policy, McMurrer, July 2007; follow-up February 2008 | Replicates |
| Public school choice take-up | Below one percent of eligible students | 2003 to 2004; 2004 to 2005 | GAO-05-7, December 2004; CRS RL31487; Department of Education | Untested |
| Supplemental educational services take-up | Twelve to nineteen percent of eligible students | 2002 to 2003 through 2004 to 2005 | GAO-06-758, August 2006 | Untested |
| Restructuring implementation | Inconsistent; two fifths took none of the statutory options | 2005 to 2006 | GAO-07-1036, September 2007 | Untested |
A note on how to read the ratings, because the table is the article’s instrument for the One Test and it should not be misread. Replicates means the finding appears across independent studies, periods, or instruments, and would be expected to appear again if the measurement were repeated. Contested means the underlying numbers are not in doubt but their interpretation is, which is the case for the national trend attribution and for the gap findings, where honest researchers disagree about what the patterns license. Untested means the intervention did not operate at a scale or with a fidelity that would let the evidence speak, which is the case for the entire consequence ladder. The table does not rate the statute. It rates the evidence, which is the distinction the whole article is built to maintain.
What Both Verdicts Get Wrong
The complication the brief requires this article to address is the pair of verdicts that dominate public discussion, and the article’s answer to both is that they rest on selective evidence. The first verdict holds that accountability raised achievement, and it typically proceeds by citing state proficiency rates, which rose impressively in many states during the statute’s lifetime. The second verdict holds that the statute accomplished nothing but narrowing, and it typically proceeds by citing the curriculum surveys and the mapping studies while ignoring the staggered-adoption research. Each verdict is built by selecting the ruler that flatters it: the first uses the state ruler for achievement and ignores the external check, the second uses the external ruler for achievement and the survey evidence for costs while declining to ask whether the modest mathematics gains were real.
The defensible summary, stated in the brief and supported by the findings above, is this. The statute’s accountability regime produced modest positive effects on mathematics achievement concentrated among lower-performing students and schools, a finding that replicates across the staggered-adoption literature and that survives the external yardstick. It produced real curriculum costs, documented in the district surveys, with instructional time moving toward tested subjects and away from untested ones at the elementary level. It produced illusory gains in some state data, where reported proficiency rose because standards fell or were set low, a finding the mapping studies document across two windows. And it produced an untested consequence apparatus, because the choice, tutoring, and restructuring provisions were implemented at too small a scale and with too little fidelity to support conclusions about whether consequences work. That summary is less satisfying than either verdict, which is why neither camp prefers it, and it has the advantage the verdicts lack, which is that every clause carries named research with its period and its method.
The modesty of that summary is its strength, and it is worth defending against the impatience it will provoke in both camps. The success camp will ask why real mathematics gains for lower-performing students, documented on an external assessment by the best research designs available, do not add up to vindication. The answer is that vindication would require more than the evidence gives: gains in reading as well as mathematics, effects that reach high school, gap-closing rather than floor-raising, and state-reported progress that survives comparison with the national ruler. The evidence gives mathematics, the early grades, the floor and a divergence. That is a real achievement purchased at a real cost, not a vindication. The failure camp will ask why documented curriculum narrowing, illusory state gains and an unused consequence ladder do not add up to indictment. The answer is symmetrical: indictment would require showing that the mathematics gains were artifacts, that the curriculum costs bought nothing, and that the consequences failed when tried. The evidence shows the gains on the external ruler, the costs alongside real improvement, and consequences that were never tried at scale. That is a costly, partial, badly measured reform, not a proven failure.
Two further cautions belong in this section because they are the places where even careful readers slip. The first is the gap question. The statute is often described, by both supporters and detractors, in terms of its effect on achievement gaps, and the evidence on gaps is the weakest part of the achievement story. Hanushek and Raymond found that the nineteen nineties accountability systems did not narrow the Black-White gap, though they did narrow the Hispanic-White gap; Lee and Wong found null results on racial gaps and largely insignificant results on socioeconomic gaps; Dee and Jacob found gains concentrated among lower performers without claiming a gap-closing triumph. A reader who wants to say the statute closed gaps is reaching beyond what the literature supports. A reader who wants to say it widened them is reaching further still. The honest statement is that accountability raised the floor in mathematics without reliably narrowing the distances between groups, and that the distances remain the central unsolved problem the evidence identifies.
What remains after both narratives are trimmed to what replicates is a smaller, more serious disagreement, and this article does not pretend to settle it. The first open question is the trade: were the mathematics gains for lower-performing students worth the curriculum costs? The gains are measured in fractions of a standard deviation on a federal assessment. The costs are measured in hundreds of thousands of instructional hours redirected from science, social studies and the arts, concentrated in the schools serving the poorest students. There is no common unit that converts one into the other, which means the trade is a value judgment, not an empirical finding, and honest people will weigh it differently. The second open question is the counterfactual about consequences: would the choice, tutoring and restructuring provisions have worked if they had been implemented at scale? The evidence says they were barely tried, which leaves the question open by definition. A reader who believes well-designed consequences would have amplified the gains is making a claim the data cannot check, and so is a reader who believes they would have failed. The third open question is whether the information half of the statute, the testing and subgroup reporting that both sides tend to keep, justifies the regime that housed it. That question belongs to political judgment, not to research.
The second caution concerns what the evidence changed. The waiver process that effectively unwound the regime in most states beginning in 2011 and 2012, and the legislative response that followed the article’s horizon, were both shaped by these findings: the implausibility of the universal proficiency deadline, the perverse incentives of state-controlled measurement, the documented costs to the untested curriculum, and the demonstrated value of subgroup reporting. The series examines what the evidence changed at nclb-vs-essa-accountability, and the statutory response itself is covered at every-student-succeeds-act-2015-guide. This article does not anticipate those developments. It records the evidence as it stood in November 2014, which is the condition under which any later judgment about the statute has to be evaluated.
The mechanics of cherry-picking deserve a final explicit statement, because recognizing them is the practical skill this article aims to teach. The pro-accountability cherry-pick works by citing state proficiency rates without the external check: the rates rose, the rise is real as a description of what the states reported, and the citation stops before the mapping studies. The anti-accountability cherry-pick works by citing the curriculum surveys and the mapping studies while omitting the staggered-adoption research: the costs are real, the standard-erosion is real, and the citation stops before Dee and Jacob. Both moves exploit the same feature of the literature, its size. With enough studies in circulation, any determined advocate can assemble a bibliography that supports a predetermined conclusion, and the abundance of citations creates an illusion of scholarly consensus where none exists. The defense against the maneuver is not more citations but better questions, and the questions are the ones this article has repeated: which ruler produced the number, what period does it cover, what comparison separates the law from the trend, and does the finding replicate on an instrument the measured party does not control. A claim that cannot survive those questions is not evidence, however many footnotes accompany it.
The role of priors is the uncomfortable corollary. Education policy attracts strong prior beliefs, about testing, about the federal role, about whether schools or society are responsible for unequal outcomes, and those priors shape which findings feel plausible before they are examined. A reader who believes testing corrupts instruction will find the narrowing evidence obvious and the mathematics gains suspicious; a reader who believes accountability disciplines institutions will find the mathematics gains obvious and the narrowing evidence overblown. The article cannot remove priors, and it does not try. It asks only that the reader hold them constant while the evidence is weighed, and that the verdict at the end be the one the surviving evidence supports rather than the one the reader held at the beginning. The modest summary, real but limited mathematics gains, real curriculum costs, illusory state-reported gains, untested consequences, is nobody’s prior. That is a point in its favor.
Presenting This Evidence in Class
The series thesis thread for this article is that assessing a statute against its own aims is difficult when the statute’s own measurement system is part of what needs assessing, and that difficulty is also what makes the article useful in a classroom. Students arrive with strong priors about testing, accountability, and the federal role in schooling, usually formed from personal experience rather than from evidence, and the five findings give an instructor a way to discipline those priors without dictating conclusions. The two-ruler problem is the pedagogical core: once a student understands that the measured party controlled the instrument, the student can evaluate every subsequent claim about the statute’s effects by asking which ruler produced it. The evidence table is designed to be projected, assigned, and argued over, because the ratings force the question the article wants students to ask, which is not whether the statute worked but what it would take to know.
Instructors looking for structured ways to teach this material, including discussion prompts on the measurement problem and exercises in distinguishing replicated findings from contested ones, will find companion material at teaching-federal-education-policy. Students who want to build their own study notes on the statute, organizing the five findings, the research citations, and the evidence table into a reviewable format, can use the VaultBook study notebook. The two-ruler problem also gives instructors a way to teach the difference between a technical dispute and a values dispute. Whether state proficiency rates overstated learning is a technical question, settled by the mapping studies. Whether the curriculum tradeoff was worth the mathematics gains is a values question, which no study can settle. Students who learn to sort the article’s findings into those two bins will read every subsequent education study more intelligently, because most public arguments about schooling blur the bins together. The evidence table’s ratings help with the sorting: replicates and contested mark the technical ground, while the costs and tradeoffs mark where values enter. The point of both resources is the same as the point of the article: the evidence on this statute rewards the reader who checks the ruler before trusting the reading.
How to Read the Next Claim You Hear About This Law
Readers will continue to encounter claims about this statute long after finishing this article, in journalism, in advocacy, in legislative debate, and in classrooms, and the article’s lasting value is the set of questions with which to meet them. The first question is always which ruler produced the number. A claim about rising proficiency that rests on state tests is a claim about what the state system reported, and it earns credence only when the federal assessment or another external measure corroborates it. A claim about falling scores that rests on a single state’s test after a cut-score change is a claim about redefinition until proven otherwise. The two-ruler discipline is not cynicism about state officials. It is the recognition that measurement choices move numbers, and that any serious claim about learning has to survive a ruler the claimant does not control.
The second question is what period the claim covers and whether the trend predates the law. Any achievement claim that begins its story in 2002 and ignores the decades before it is vulnerable to the attribution error, because the national series was moving before the statute existed. The relevant comparison is not the level of scores during the law’s lifetime but the change in the slope after its introduction, measured against a comparison group that did not receive the treatment at the same time. Claims that present a rising line as proof of the law’s success, without a comparison and without attention to the pre-trend, are describing the background and calling it the effect.
The third question is whether the claim confuses take-up with effectiveness. The consequence provisions failed as mechanisms, not as ideas: choice reached fewer than one percent of eligible students, tutoring reached fewer than one in five, and restructuring was implemented nominally. Any argument that the sanctions were proven to work, or proven to fail, on the basis of this record is arguing from an intervention that barely operated. Untested is not a synonym for ineffective, and it is not a synonym for effective either. It is a statement about the limits of what the record can teach, and honoring it is what separates the careful reader from the advocate.
The fourth question is whether the claim extends a finding beyond the population and subject where it was measured. The positive results live in elementary mathematics, concentrated among lower-performing students; they do not live in reading, and they do not live in high school. A claim that the law raised achievement in general, without the subject and grade qualifications, is larger than the evidence. A claim that the law narrowed the curriculum in general, without noting the elementary concentration and the survey basis of the evidence, is likewise larger than the evidence. Precision about scope is not pedantry. It is the difference between a finding and a slogan.
Held together, the four questions are a portable method, and they apply beyond this statute to any accountability regime a reader will encounter. Ask which ruler, ask about the pre-trend, ask whether the mechanism operated, ask about scope. The answers will rarely be as satisfying as a verdict, and they will always be more honest than one.
The last word belongs to the measurement problem, because it is the part of this story that outlives the statute. A future accountability regime that lets each measured party define its own ruler will reproduce this literature’s central frustration: years of argument over numbers that cannot settle the question they are cited to answer. The findings that replicated under No Child Left Behind are the ones that survived on a ruler nobody controlled. That is not a partisan conclusion. It is a methodological one, and it is the reason this article insisted on naming the assessment, the period, the authors and the method for every claim it allowed to stand.
Frequently Asked Questions
Q: Did No Child Left Behind improve test scores?
The answer depends on which test the question means, and that dependence is the entire lesson of the evidence. On the federal assessment, mathematics scores for nine and thirteen year olds rose substantially over the long term, twenty four and fifteen points respectively from 1973 to 2008, while reading gains were smaller and high school performance was largely flat. But those trends began decades before the statute, so they cannot be credited to it without a research design separating the law from the trend. The staggered-adoption research provides that design and finds modest positive effects on mathematics, concentrated among lower-performing students, with no significant reading gains. State proficiency rates rose faster than the federal assessment corroborated, and the mapping studies show that much of that reported improvement reflected standard-setting rather than learning. The defensible answer is therefore qualified: modest mathematics gains for weaker students, replicated across studies; flat reading; and state-reported gains that outran the external evidence.
Q: Did No Child Left Behind close achievement gaps?
The literature does not support that claim. The staggered-adoption studies find that accountability raised mathematics achievement concentrated among lower-performing students and schools, which raised the floor without reliably narrowing the distances between groups. Hanushek and Raymond, studying the nineteen nineties state accountability systems in 2005, found clear positive effects on score growth but no narrowing of the Black-White gap, though they did find narrowing of the Hispanic-White gap. Lee and Wong, in 2004, found null results on racial gaps and largely insignificant results on socioeconomic gaps. Dee and Jacob, in 2011, found mathematics gains concentrated among low-performing groups without framing the results as a gap-closing success. A reader who wants a simple verdict on gaps will not find one in the research. The honest summary is that the regime lifted lower achievers in mathematics while leaving the between-group distances largely where they were, which is a real accomplishment and a real limitation at once.
Q: Did states lower standards under No Child Left Behind?
In a substantial number of states, yes, and the mapping studies document it. The National Center for Education Statistics report NCES 2010-456, published in October 2009, placed state proficiency standards on the federal assessment’s scale for 2005 and 2007 and found that the standard was lower in 2007 than in 2005 in between one third and one half of the states. The follow-up report NCES 2011-458, released in August 2011, found that standards varied by at least sixty points on the federal scale and that nearly all of them mapped at or below the federal assessment’s Basic level. The mechanism was the incentive structure the statute created: the fixed trajectory toward universal proficiency by the 2013 to 2014 school year, combined with state control over the definition of proficiency, rewarded states that eased their bars and punished states that held them. The article does not attribute this to bad faith by individual officials. It attributes it to a design in which the measured parties controlled the ruler, which made the erosion of standards the predictable outcome rather than a surprising one.
Q: Did No Child Left Behind narrow the curriculum?
Yes, and the evidence is survey-based and substantial. The Center on Education Policy surveyed three hundred forty nine school districts for its July 2007 report by Jennifer McMurrer, “Choices, Changes, and Challenges,” and found that sixty two percent of districts had increased instructional time in English language arts and mathematics at the elementary level since the 2001 to 2002 school year, by average increases of forty six percent in English language arts and thirty seven percent in mathematics. Forty four percent of districts had cut time from at least one other subject, by an average of thirty one percent, with the losses falling on science, social studies, art, music, and physical education. A follow-up report in February 2008 extended the findings. These are self-reports by administrators rather than time-use measurements, so the exact minutes should be read as approximate, but the direction of the reallocation is unambiguous. The statute tested two subjects and attached consequences to them, and schools moved time accordingly. The finding is reported here as a documented cost, which is how the statute’s own supporters have generally described it.
Q: How many schools failed to meet No Child Left Behind targets?
The Center on Education Policy estimated that forty eight percent of the nation’s public schools, more than forty three thousand schools, missed adequate yearly progress in the 2010 to 2011 school year, up from thirty nine percent the year before, according to its report “AYP Results for 2010 to 2011,” published December 15, 2011. The forty eight percent figure was a six-year high in thirty five states; in twenty four states plus the District of Columbia, at least half of schools missed their targets; in five states plus the District, at least three quarters missed. These are the Center’s estimates, which it noted could be revised by one or two percentage points as final data arrived. The article does not use the eighty two percent figure that circulated in public discussion, because that number was a pre-data projection offered before the results were compiled, and the Center’s actual data contradicted it. The rising failure rate was the predictable product of the statute’s fixed trajectory: as the 2013 to 2014 deadline for universal proficiency approached, targets climbed beyond the reach of a growing share of schools, including many whose students were learning at respectable levels.
Q: What does research say about No Child Left Behind?
The research literature converges on four propositions and leaves a fifth open. First, the federal assessment shows mathematics gains for younger students across the period, smaller reading gains, and flat high school performance, but the trends predate the statute and cannot be attributed to it without a causal design. Second, the staggered-adoption studies find modest positive effects on mathematics concentrated among lower-performing students and schools, with weak or null effects in reading, a pattern that replicates across Carnoy and Loeb in 2002, Hanushek and Raymond in 2005, and Dee and Jacob in 2011. Third, the mapping studies show that state proficiency rates overstated learning, because standards were often set low and sometimes lowered, with state-reported gains uncorroborated by the federal assessment in at least half of the comparison states from 2007 to 2009. Fourth, the curriculum narrowed, with tested subjects gaining time at the expense of untested ones. The open question is the consequence ladder: choice, tutoring, and restructuring were implemented at too small a scale to evaluate, so the sanctions side of the statute is untested rather than refuted. That is the literature in one paragraph, and every clause of it carries a citation.
Q: Did No Child Left Behind help students with disabilities?
The evidence on this subgroup is thinner than the evidence on racial, ethnic, and economic subgroups, and the article states that limitation directly. The subgroup reporting requirement meant that the performance of students with disabilities was disaggregated and published for the first time on a national scale, which made their achievement visible in a way the prior regime had not. On the question of whether accountability pressure produced gaming through special education placement, Hanushek and Raymond tested the hypothesis in their 2005 study and found no evidence that stronger accountability increased special education placement rates. The Dee and Jacob 2011 study found its mathematics gains concentrated among lower-performing groups, a category that includes many students with disabilities, though the study did not isolate that subgroup’s gains separately. What the literature does not provide, within the article’s horizon, is a dedicated causal estimate of the statute’s effect on the achievement of students with disabilities as a distinct group. The visibility the reporting requirement created is well documented; the achievement effect is not separately established.
Q: Was subgroup reporting the lasting achievement of No Child Left Behind?
It is the strongest candidate for that title, and the claim survives the scrutiny this article applies to the statute’s other supposed accomplishments. Before the 2001 law, school-level results could be reported as averages that concealed the performance of minority students, students with disabilities, English learners, and economically disadvantaged students behind a respectable mean. The adequate yearly progress system required disaggregation by subgroup and required each qualifying subgroup to meet its targets independently, which meant a school could not offset one group’s failure with another group’s success. That design had a perverse statistical consequence, the diversity penalty, under which schools serving many subgroups faced more hurdles and failed at higher rates, a finding researchers documented and the statute’s framers never resolved. But the visibility itself changed the conversation: achievement gaps became a matter of published record rather than private knowledge. Whether later accountability designs preserved the disaggregation is a question for the comparison with the successor regime, but the innovation of measuring every group separately is the provision whose influence most clearly outlasted the statute’s own machinery.
Q: How did federal assessment scores change during the No Child Left Behind era?
The long-term trend of the federal assessment, published as NCES 2009-479 in April 2009, shows reading scores for nine year olds rising twelve points from 1971 to 2008 and for thirteen year olds rising four points, with seventeen year olds showing no significant change. In mathematics, nine year olds gained twenty four points and thirteen year olds gained fifteen points from 1973 to 2008, with seventeen year olds again unchanged. The main assessment’s grade twelve results, released in November 2010, showed mathematics moving from one hundred fifty in 2005 to one hundred fifty three in 2009, and reading from two hundred eighty six to two hundred eighty eight, still below the nineteen ninety two average of two hundred ninety two. The pattern is consistent: mathematics improved for younger students, reading improved less, and high school performance was flat. The crucial qualifier is that the long gains accumulated across decades, so the statute’s contribution cannot be read off the trend. The causal estimates come from the staggered-adoption studies, not from the national series.
Q: What is the two-ruler problem in No Child Left Behind research?
It is the measurement problem at the center of this article: the statute judged progress with instruments controlled by the parties being judged, so the central empirical question cannot be answered from state data alone. A state proficiency rate is the share of students above a cut score on a test the state chose, with the cut score set by the state and the tested population defined by state rules. Every one of those choices moves the reported number without moving student knowledge, which means the number is a joint product of learning and measurement decisions. The external ruler is the federal assessment, administered under uniform conditions with cut scores the states do not control. When the two rulers agree, the finding is strengthened; when they diverge, as the mapping studies show they often did, the divergence is itself the finding. Every credible result in this literature depends on the external yardstick, because the state ruler alone cannot distinguish improvement from redefinition.
Q: How did staggered state adoption of accountability help researchers study No Child Left Behind?
It supplied the comparison group that a single national launch could not. The federal statute imposed accountability on every state at once, which means its introduction alone offers no counterfactual: there is no version of the country without the law to measure against. But states had adopted consequential accountability systems at different times before the federal law, some with public reporting and sanctions in the nineteen nineties, others with weak systems or none. Researchers compared the achievement trajectories of early adopters and late adopters on the federal assessment, before and after the federal law arrived, which holds the national background trend constant and varies only the timing of the treatment. Carnoy and Loeb used this logic for 1996 to 2000 in 2002, Hanushek and Raymond for the nineteen nineties systems in 2005, and Dee and Jacob for the federal introduction itself in 2011. The design is what lets the literature claim modest causal effects on mathematics rather than merely describing trends.
Q: What was the diversity penalty under No Child Left Behind?
It was the statistical consequence of the subgroup-hurdle design: because a school had to meet its adequate yearly progress targets for every qualifying subgroup independently, each additional subgroup was an additional hurdle, so schools serving diverse populations failed at higher rates than schools with similar overall achievement but fewer subgroups. Researchers documented the pattern during the statute’s lifetime, and it embarrassed the law’s civil rights framing, since the schools most likely to be labeled as failing were the schools serving the most minority students. Defenders of the design replied that the hurdle structure was the point: a school should not be able to average away one group’s poor performance with another group’s strong results, and the penalty was the price of refusing to let disadvantage hide inside an average. The dispute was never resolved within the statute’s lifetime. The finding matters for the evidence because it shows how a measurement choice, disaggregation with independent hurdles, shaped which schools the system identified, independent of how much students were learning.
Q: Did accountability under No Child Left Behind help English learners?
English learners were one of the subgroups whose performance had to be reported separately and whose targets had to be met independently, which made their achievement visible in the same way the reporting requirement made other groups visible. The staggered-adoption literature’s positive mathematics findings, concentrated among lower-performing students and schools, plausibly include English learners, since they are disproportionately represented among lower achievers, but the major studies did not isolate English learners as a distinct group for causal estimation. Dee and Jacob’s 2011 finding of moderate eighth grade mathematics gains targeted at low-performing groups is the closest the causal literature comes, and it is a statement about the bottom of the distribution rather than about language status. A separate literature on the testing of English learners during this period raised concerns about the validity of assessing students in a language they were still acquiring, but validity concerns are distinct from achievement effects. Within the article’s horizon, the documented contribution is visibility through disaggregated reporting; a dedicated causal estimate of the achievement effect for English learners is not established.
Q: Why did rising test scores coincide with rising school failure rates?
Because the statute’s annual targets rose every year toward the 2013 to 2014 deadline for universal proficiency, so the bar moved up faster than scores did. A school could improve its students’ achievement and still miss its target if the target rose more than the scores, which is exactly what happened in many places as the deadline approached. The Center on Education Policy estimated that forty-eight percent of public schools missed their targets in 2010 to 2011 even though national assessment trends showed mathematics gains for younger students. The two facts are compatible because they measure different things: the national assessment measures absolute achievement on a fixed scale, while adequate yearly progress measured movement against a rising bar. This is why the failure rates of the late accountability years are evidence about the deadline’s arithmetic rather than evidence that learning stopped. Confusing the two is one of the most common errors in public discussion of the late accountability years.
Q: How did the one hundred percent proficiency deadline shape state behavior?
Section 1111(b)(2)(F) required states to set a timeline ensuring all students would reach proficiency not later than twelve years after the end of the 2001 to 2002 school year, fixing the deadline at 2013 to 2014. As the deadline neared and annual targets climbed, states faced a choice between the politically impossible admission that universal proficiency would not happen and the quieter path of adjusting their own definitions of proficiency. The federal mapping studies show that between one third and one half of states lowered their standards between 2005 and 2007. The waiver offer of September 23, 2011 gave states relief from the 2014 deadline in exchange for new accountability designs, which was the executive branch’s answer to a deadline Congress would not revise. The deadline’s main measurable effect may therefore have been on the ruler rather than on learning.
Q: Did schools focus extra attention on students near the proficiency cutoff?
The research literature discusses this triage behavior, sometimes called bubble-student focus or educational triage, as a predicted response to a proficiency-threshold system: when a school’s rating depends on the share of students above a cut score, the highest return on effort comes from students just below the line. Dee and Jacob’s finding that mathematics gains were concentrated among lower-performing students is consistent with schools directing additional instruction toward struggling students, though their design cannot distinguish broad-based improvement among low performers from narrow targeting of those nearest the cutoff. The distinction matters for interpreting the gains: improvement concentrated at the bottom of the distribution is still improvement, but a system that rewards moving students just across a line creates weaker incentives for the lowest performers far below it and for high performers above it. The national assessment data cannot resolve which pattern dominated. Either pattern is consistent with the concentrated gains among lower performers, which limits how much the triage concern can explain away the results.
Q: Did schools push low-performing students out to raise their scores?
The best available evidence says this particular gaming mechanism did not operate at scale. Carnoy and Loeb’s 2002 study of 1990s state accountability systems found no effect on student retention or high school completion alongside the mathematics gains, which is inconsistent with the theory that the gains were manufactured by holding students back or pushing them out. Hanushek and Raymond’s 2005 study found no evidence of increased special-education placement under accountability pressure, rebutting the related theory that schools reclassified low performers to remove them from the tested pool. These are null findings and should be read as such: they show no support for specific predicted forms of gaming in the data and periods studied, not proof that no gaming ever occurred anywhere. Other mechanisms, such as teaching narrowly to the test, are harder to detect in the available data and remain plausible contributors to the score patterns. The distinction matters: the gaming mechanisms the data can test show no support, while the ones it cannot test remain open questions rather than established facts.
Q: How did accountability change what teachers did in the classroom?
The district-level evidence shows reallocation of instructional time rather than a single national change in teaching method. The Center on Education Policy’s 2007 survey found elementary instructional time shifting substantially toward English language arts and mathematics, with reported average increases of forty-six and thirty-seven percent respectively, and away from other subjects. Within the tested subjects, the statute’s incentives rewarded tested content and tested formats: more practice with the kinds of items the state test used, more test preparation, and more intervention time for students at risk of missing proficiency. The national assessment’s weaker reading results despite large reported increases in English language arts time suggest that additional minutes did not reliably translate into the broader reading proficiency the federal assessment measures. What teachers did differently varied by district, but the direction of the reallocation was set by what the accountability system counted.
Q: Why did reading results disappoint despite more reading instruction?
Districts reported increasing elementary English language arts instructional time by an average of forty-six percent since 2001 to 2002, yet the national assessment showed only small reading gains for younger students and flat scores for seventeen-year-olds, and Dee and Jacob found no significant reading effects from accountability at grades four or eight. Several explanations are consistent with the evidence. Additional minutes may have gone to test preparation and tested formats rather than to the broad comprehension the national assessment measures. Reading may respond less to additional drill than mathematics does, since mathematics content is more sequential and more directly teachable in the tested formats. And the students receiving the most additional reading time were the lowest performers, whose gains may not have been large enough to move group averages. The reading null is a finding about the limits of the statute’s theory of action, not proof that reading instruction cannot work.
Q: What did the waiver era reveal about the original deadline?
It revealed that the 2013 to 2014 deadline for universal proficiency was not credible as written. By 2011, with annual targets climbing toward one hundred percent, the Center on Education Policy estimated that nearly half of public schools were missing their targets, and the trajectory pointed toward the great majority missing as the deadline arrived. Rather than let the deadline trigger consequences for most of the nation’s schools, the Department of Education offered waivers from the core requirements on September 23, 2011, with the first approvals on February 9, 2012, in exchange for states adopting new standards, accountability designs and teacher evaluation systems. The waivers effectively unwound the federal accountability regime in most states while the statute remained on the books. A deadline that the executive branch sets aside rather than enforces was a deadline the design could not sustain. The episode is the clearest demonstration that the uniform target, the element of the statute with the most ambitious rhetoric, was also the element with the least credible arithmetic.