Two federal laws measured American schools with almost identical instruments and then answered the only question that mattered in opposite ways. The first, Public Law 107-110, signed on January 8, 2002, read the measurement results and prescribed the response from Washington: a fixed national deadline, a ladder of federally defined interventions, and a Secretary of Education with broad discretion to attach conditions to relief. The second, Public Law 114-95, signed on December 10, 2015, kept the instruments, discarded the federal script for what happens next, and wrote explicit prohibitions on the Secretary into the statute itself. Anyone who wants to understand American education accountability must be able to state exactly which requirements survived the 2015 rewrite and which did not, explain that the shift concerned who designs consequences rather than whether measurement is required, and defend a verdict about which model surfaces inequity and which produces useful responses.

The One Test this comparison must pass
Every comparison in this series is built around a single test of the reader, and this one has three parts. First, the reader can state exactly which requirements survived the 2015 rewrite and which did not: annual testing in reading and mathematics in grades 3 through 8 and once in high school survived intact; disaggregated subgroup reporting survived intact; the 100 percent proficiency deadline did not; the federally prescribed ladder of interventions did not; the Secretary’s discretion to condition relief on favored policies did not. Second, the reader can explain that the shift was about who designs consequences rather than about whether measurement is required. The 2015 law did not reduce testing, did not weaken reporting, and did not end federal identification of struggling schools; it moved the design of the response from Washington to state governments and districts. Third, the reader can reach a defended verdict on which model does more to surface inequity and which does more to produce useful responses, with the deciding factor named explicitly: the earlier regime was better at forcing uncomfortable information into public view and worse at producing responses that helped, so a reader who values comparability across states should prefer elements of the earlier model while a reader who values locally designed intervention should prefer the later one.
That verdict is deliberately split, and the split is the point. The two statutes are usually reduced to a slogan about federal overreach, as if the entire difference between them were the volume of Washington’s voice. The statutes themselves tell a more precise story, and the precise story is the one this article follows: same measurement, different response. The two laws measure schools almost identically and differ almost entirely in who decides what happens next, so debates framed as testing versus no testing are describing a difference that does not exist in the statutes. This article takes no position on standardized testing as a practice; its judgments concern institutional design and the evidence about how each design behaved, with every performance claim attributed to named research and no reference to any state plan or dispute in force at any time.
Two statutes, one framework
Public Law 107-110 began as H.R. 1 in the 107th Congress, passed the House of Representatives on December 13, 2001 by a vote of 381 to 41 on Roll Call Number 497, passed the Senate on December 18, 2001 by a vote of 87 to 10 on Record Vote Number 371, and was signed into law on January 8, 2002. Public Law 114-95 began as S. 1177 in the 114th Congress, passed the House on December 2, 2015 by a vote of 359 to 64 on Roll Call Number 665, passed the Senate on December 9, 2015 by a vote of 85 to 12 on Record Vote Number 334, and was signed into law on December 10, 2015. Both statutes are reauthorizations of Public Law 89-10, the Elementary and Secondary Education Act of 1965, which first tied federal education aid to conditions on how states serve disadvantaged students. The lineage matters because it fixes the terms of the comparison: the 2015 act did not repeal its predecessor in the way one law repeals an unrelated one; it amended the same underlying framework, replacing the accountability title of the 2002 reauthorization with a new one while leaving the architecture of the 1965 act standing. Readers who want the full story of that underlying framework can begin with the Elementary and Secondary Education Act of 1965, which explains why federal money has carried federal conditions since the Johnson administration.
The comparison that follows is organized along four axes, because the differences between the two regimes concentrate in four places and nowhere else. The first axis is measurement: what tests are given, in which grades, and how the results are reported. The second axis is the target: what level of performance the law demands and who sets it. The third axis is the consequence: what happens, by law, when the results fall short. The fourth axis is executive discretion: what the Secretary of Education may and may not do with the authority the statute grants. On the first axis the two laws are nearly indistinguishable. On the other three they diverge sharply, and the divergence runs in a consistent direction: from federal prescription toward state design, with federal guardrails retained around measurement, identification, and evidence. Holding those four axes steady is what keeps the comparison honest, because the public debate about the two laws tends to collapse all four into one, usually under the banner of testing, which is precisely the axis where nothing changed.
A note on dates is necessary before the substance begins. This article is dated November 15, 2014, and the 2015 act was signed on December 10, 2015, so the later statute is presented throughout as the later development it is: a subsequent rewrite whose provisions are described in the past tense with explicit dates, in the same manner as the other retrospective comparisons in this series. The 2002 law is described as the earlier regime and the 2015 law as the later one, and where the article refers to “the earlier statute” it means Public Law 107-110, while “the later statute” means Public Law 114-95. The comparison is a decision exercise as well as a description: it resolves with a verdict and a named deciding factor, in keeping with the series thesis that a comparison is only finished when the reader can say which design they would choose and why.
Axis one: measurement stayed the same
The most consequential misunderstanding about the 2015 rewrite is also the most widespread: the belief that it reduced testing. It did not. The testing requirements of the two statutes are identical in substance, and the identity is verifiable section by section. Under Public Law 107-110, states were required to implement annual assessments in reading and mathematics in each of grades 3 through 8, with the full schedule in place by the end of the 2005-2006 school year, plus assessments at least once in the grades 10 through 12 band. Under Public Law 114-95, Section 1111(b)(2)(B)(v)(I), codified at 20 U.S.C. Section 6311, states must administer annual assessments in reading and mathematics in each of grades 3 through 8 and at least once in grades 9 through 12. The grades are the same, the subjects are the same, the annual cadence is the same, and the high school requirement is the same in substance: at least one administration in the high school years. A journalist, a student, or an administrator who learns only one fact from this comparison should learn this one, because nearly every other error in the public debate descends from getting it wrong.
Did the 2015 law end annual testing?
No. Section 1111 of the 2015 reauthorization requires annual reading and mathematics assessments in grades 3 through 8 and at least once in grades 9 through 12. The schedule matches the testing calendar of the 2002 law, which demanded yearly exams in the same grades by the end of the 2005-2006 school year. The rewrite never reduced testing.
The persistence of the testing schedule is worth dwelling on because it explains why the slogan version of the story took hold. The 2002 law arrived with testing as its most visible feature: for the first time, every state had to test every child every year in the tested grades, and the results had to be broken out by subgroup, a combination that made school performance newly legible to parents, reporters, and researchers. When Congress replaced the law thirteen years later, the political energy behind the replacement was directed at the consequences of the tests rather than the tests themselves: the impossible deadline, the cascade of sanctions, and the federal prescription of interventions. But in public memory the grievance attached to the whole package, testing included, so the rewrite was remembered as the law that ended the testing era. The statute says otherwise. The annual assessments continued without interruption, the grades did not change, and the subjects did not change. What changed was everything that happened after the scores came back.
Reporting, the second half of measurement, followed the same pattern of continuity with additions. Both statutes require assessment results to be disaggregated by the same subgroups: race, ethnicity, economic disadvantage, disability status, and English learner status. Both require the disaggregated results to be published in public report cards so that the performance of each group is visible rather than buried inside a schoolwide average. The disaggregation requirement was the engine of the earlier law’s information effects: by forcing the performance of historically underserved groups into the open, it made it impossible for a school to look successful on average while failing the students who most needed help. The later statute kept that engine and added reporting categories rather than removing any. The additions include per-pupil expenditure data reported at the school level and information about educator qualifications, both of which extended the public’s view from outcomes alone to the resources behind them. The direction of change in reporting was expansion, not contraction, and the widespread assumption that subgroup reporting was weakened has no basis in the statutory text.
Did subgroup reporting survive the rewrite?
Yes. Both statutes require test results to be disaggregated by race, ethnicity, economic disadvantage, disability status, and English learner status, and both require the figures to appear in public report cards. The 2015 law added reporting categories rather than removing them, including per-pupil expenditure data and information on educator qualifications. Reporting grew, not shrank.
A third element of measurement, less discussed but important for interpreting results across the two eras, is the 95 percent participation rule. Under the earlier law, testing at least 95 percent of students in each subgroup was part of adequate yearly progress: a school that failed to test enough of its students failed to make progress regardless of how the tested students performed. Under the later statute, Section 1111(c)(4)(E) requires states to factor the 95 percent participation of all students and of each subgroup into the statewide accountability system, with the statute specifying how non-participation affects the school’s rating. The rule survived in both letter and function: in each regime, the law refuses to let a school improve its measured performance by excluding students from the measurement. The continuity matters because participation rules are where measurement systems are most vulnerable to gaming, and both Congresses understood that a testing requirement without a participation requirement is an invitation to select the test-takers.
The measurement story also includes the National Assessment of Educational Progress, the federal testing program that both laws used as an external check. The earlier statute required states receiving Title I funds to participate in the state-level National Assessment in reading and mathematics in grades 4 and 8, and the later statute retained that requirement. The long-term trend component of the assessment, which tests representative samples of students at ages 9, 13, and 17 rather than in grades, provides the cleanest available record of what happened to achievement across the two eras. According to the National Center for Education Statistics, long-term trend reading scores rose by 12 points for 9-year-olds and 4 points for 13-year-olds between 1971 and the most recent administration, with 17-year-olds essentially flat; mathematics scores rose by 24 points for 9-year-olds and 15 points for 13-year-olds since 1973, with 17-year-olds again flat. Those figures describe decades, not the tenure of either statute, and this article attributes no causal claim to them; they are included because any verdict about the two regimes must be read against the background of what the national trend data actually show.
One more measurement fact belongs in this axis because it disciplined how the earlier law’s results were read. The National Center for Education Statistics published two mapping studies, NCES 2010-456 covering 2005 through 2007 and NCES 2011-458 covering 2005 through 2009, that compared each state’s own proficiency results against that state’s National Assessment results. In most cases the states’ own results showed more positive changes than the National Assessment did: 12 of 14 comparisons in the first study and 8 of 10 in the second. The mapping studies are the documented basis for the claim, developed further in the targets section, that the earlier law’s deadline created an incentive for states to define proficiency downward on their own tests. The point for the measurement axis is narrower: the federal testing apparatus included, from early in the earlier law’s life, a mechanism for checking state-reported results against a common national yardstick, and the later law preserved that mechanism. Readers who want the full account of the earlier regime’s measurement machinery, including the adequate yearly progress calculations that turned test scores into school ratings, can consult the No Child Left Behind Act of 2001 guide, which walks through the earlier statute’s provisions in the order they operated.
Axis two: who sets the target
If measurement is the axis where the two statutes agree, the target is the axis where they part most visibly, and the parting defines the character of each regime. The earlier statute set a single national proficiency deadline: 100 percent of students, in every subgroup, proficient in reading and mathematics by the end of the 2013-2014 school year. The mechanism was adequate yearly progress, the annual calculation that determined whether a school, a district, and a state were on track toward the deadline. Each state defined its own starting point and its own trajectory of annual objectives, but the endpoint was fixed by Congress and it was the same for every state: universal proficiency, on a date certain, thirteen school years after enactment. There was no provision for missing the deadline and trying again; the statute’s consequence ladder, described in the next section, was the answer to the question of what missing it meant.
The deadline did two things at once, and the two effects pulled in opposite directions. On one side, it concentrated political and administrative attention on student achievement with an intensity no prior federal education law had achieved. State legislatures, district offices, and school faculties organized their work around the annual objectives; the phrase “adequate yearly progress” entered the vocabulary of school board meetings across the country; and the public learned to read a school’s trajectory the way investors read a company’s earnings. The deadline made achievement the organizing fact of school governance. On the other side, the deadline was unreachable, and its unreachability distorted the system it was meant to discipline. As the 2013-2014 school year approached, the number of schools failing to make adequate yearly progress rose year after year, not because schools were collapsing but because the trajectory required gains that no school system in history had sustained. By the early 2010s the law’s central promise had become a scheduled embarrassment: nearly every school in America was on course to be labeled as failing, a label that would have described the system rather than distinguished within it.
The distortion ran deeper than the labels. Because each state set its own definition of proficiency on its own assessments, the fixed federal deadline created a direct incentive to define proficiency downward. A state that set a demanding standard and an honest trajectory would watch its schools fail in growing numbers as the deadline neared; a state that set a lenient standard could report steady progress toward the same federal goal. The National Center for Education Statistics mapping studies, described in the measurement section, documented the result: in most comparisons, states’ own results showed more favorable trends than the National Assessment showed for the same states over the same period. The incentive did not require dishonesty by any individual; it required only that the officials who set cut scores understood which choice made their schools look better under federal law. The deadline thus produced the worst combination a target can produce: it was simultaneously too demanding to be credible and too easy to evade by redefinition. Dee and Jacob, in their 2011 analysis published in the Journal of Policy Analysis and Management, found significant gains in fourth-grade mathematics attributable to the earlier law, with an effect size of 0.22 by 2007 and improvements at both the lower and top percentiles, alongside eighth-grade mathematics gains concentrated among low-achieving groups, but they found no consistent gains in reading. The pattern is consistent with a system that pushed hard on tested mathematics while leaving reading largely unmoved, and it is the kind of finding that a comparison of institutional designs must register without turning it into a slogan.
Did the 2015 law remove every performance target?
No. It removed the federal deadline of 100 percent proficiency by the end of the 2013-2014 school year, but it requires each state to adopt long-term goals with interim measurements for all students and each subgroup. The targets became state-designed rather than federally fixed, which removed the pressure to define proficiency downward while removing any national benchmark.
The later statute’s answer to the target question reflects everything Congress learned from the deadline’s failure. Section 1111(c)(4)(A) requires each state to establish long-term goals, including measurements of interim progress, for all students and for each subgroup of students. The goals must address, at a minimum, improved academic achievement as measured by proficiency on the annual assessments, high school graduation rates, and progress in achieving English language proficiency for English learners. But the statute sets no federal deadline and no federal definition of the goal’s endpoint. Each state decides what its long-term goals are, what interim progress looks like, and how far in the future the goals extend. The federal role is reduced to requiring that goals exist, that they cover the named indicators, that they apply to all students and each subgroup, and that progress toward them is measured and reported.
The change removed the perverse incentive at a stroke. With no federal deadline to satisfy, no state gains anything under federal law by defining proficiency downward; the cut score no longer determines whether the state can claim compliance with a national timetable, because there is no national timetable. The honest measurement that the earlier law’s deadline discouraged became, under the later law, simply the state’s own choice, undistorted by federal reward. But the change also removed the common benchmark, and the loss is real. Under the earlier regime, every state was at least nominally climbing the same mountain, and the public could ask why one state’s trajectory looked so different from another’s. Under the later regime, each state climbs a mountain of its own choosing, and cross-state comparisons of ambition become exercises in reading fifty different plans with fifty different definitions of success. A reader who values comparability across states should prefer elements of the earlier model on exactly this ground: the federal deadline was a bad target, but it was a common one, and common targets make evasions visible in a way that fifty bespoke targets do not.
There is a further subtlety in the later law’s goal provisions that deserves attention because it is often misread. The statute requires goals for each subgroup, not only for the student population as a whole, and it requires the goals to be long-term with interim measurements. The subgroup requirement means that a state cannot satisfy the law with goals that track only overall averages; the performance of each historically underserved group must have its own trajectory. The interim measurement requirement means that a state cannot set a distant goal and ignore it until the deadline arrives; progress must be checkable along the way. These are genuine federal constraints on state-designed goals, and they belong in the accurate description of the later regime: state-designed does not mean unconstrained, and the subgroup and interim-progress requirements are the constraints that keep the design honest. The accurate contrast is therefore not between federal targets and no targets, but between a single federally fixed endpoint and federally required state-designed trajectories. That contrast is less dramatic than the slogans, and it is the one the statutes actually contain.
Proficiency, growth and the indicator question
Beneath the target axis lies a technical question that shaped both statutes: what should count as performance in the first place. The 2001 statute answered primarily with proficiency, the share of pupils scoring above the state’s cut score. Proficiency as a status measure has a well-documented bias. A school serving advantaged, well-prepared pupils can post high proficiency rates while adding little learning, and a school serving disadvantaged pupils can add enormous learning while posting low proficiency rates, because the measure records where pupils stand rather than how far they have come. Under the earlier law’s adequate yearly progress system, the second school failed and the first succeeded, regardless of their respective contributions to pupil growth. Educators in high-poverty schools experienced this as a deep unfairness built into the law’s mathematics, and analysts who studied the law’s incentives noted that status-based accountability gave schools reason to concentrate effort on pupils just below the cut score, whose movement across the line improved the school’s rating, rather than on the lowest-performing pupils, whose growth would not reach the line within a single year.
The unfairness had a behavioral consequence that researchers documented. Because only movement across the proficiency threshold affected the school’s verdict, instructional attention flowed toward the pupils nearest the threshold and away from those furthest below it, the very pupils the statute’s title suggested it existed to serve. Dee and Jacob’s 2011 finding, that eighth-grade mathematics gains concentrated among low-achieving groups, complicates any simple telling, since it shows the law did reach some of those pupils in mathematics. But the incentive structure of a status measure points the other way, and the tension between the incentive and the finding is one of the reasons the literature on the period resists summary. A system can create incentives to neglect the lowest performers and still, through other channels such as the resource increases Dee, Jacob and Schwartz documented in 2013, produce gains among them.
The 2015 statute addressed the indicator question by requiring multiple measures. Section 1111(c)(4)(B) obliges each state’s accountability system to include, for all pupils and disaggregated by subgroup: academic achievement as measured by proficiency on the annual assessments; a measure of student growth or another valid and reliable statewide academic indicator for elementary and middle schools; graduation rates for high schools; progress in achieving English language proficiency; and at least one measure of school quality or student success, such as chronic absenteeism, school climate or access to advanced coursework. The statute further requires that the academic indicators receive substantial weight individually and much greater weight in the aggregate than the school quality indicator, a provision meant to prevent states from drowning test results in softer measures.
The shift from a single status indicator to a weighted set changed what performance meant in law. A school with low proficiency but strong growth could distinguish itself from a school with low proficiency and weak growth, a distinction the earlier system blurred. A high school’s graduation rate, measured by the four-year adjusted cohort method, entered the system as a freestanding indicator rather than as an additional adequate-yearly-progress check. English learners’ progress toward proficiency became a required element rather than an afterthought. And the school quality indicator, though subordinated by the weighting rule, gave states a federally sanctioned place to measure things the tests did not capture, from attendance to course access.
Three cautions keep the indicator story honest. First, the weighting requirement was the federal backstop against evasion, and its effectiveness depended on Department enforcement of the deliberately vague term substantial. A state determined to minimize the role of test scores could comply with the letter while stretching the spirit, and the statute gave the Department limited tools to police the difference, particularly given the new restrictions on Secretarial discretion. Second, growth measures are statistically demanding. They require matched pupil-level data across years, models that account for measurement error, and decisions about how to attribute growth to schools rather than to pupil mobility or demographic change. States implemented them with varying sophistication, which meant the growth indicator itself varied in quality across the country. Third, adding indicators multiplies complexity for the public the system is meant to inform. The earlier law’s single verdict, made adequate yearly progress or did not, was crude but legible to any parent or reporter. The later law’s differentiated system is fairer in principle but harder to summarize, and legibility is itself an accountability value: a system the public cannot understand is a system the public cannot use to apply pressure. The indicator question thus replays the comparison’s central trade in miniature, with the later design measuring performance more richly and communicating it less simply.
Axis three: who designs the response
The consequence axis is where the comparison earns its subtitle of same measurement, different response, and it is the axis on which the two statutes are most cleanly opposed. The earlier law prescribed a federal ladder of interventions, triggered automatically when a school missed its adequate yearly progress targets for consecutive years, with each rung specified in the statute. The later law requires the identification of defined categories of schools and leaves the design of the intervention to states and districts, subject to federal evidence requirements. The earlier regime answered the question “what happens next” in Washington; the later regime requires that the question be answered, names the schools for which it must be answered, sets standards for the quality of the answer, and leaves the substance of the answer to the state and the district. Everything else in this section is elaboration of that sentence.
The earlier law’s ladder, set out in Section 1116, operated on consecutive years of missed targets. After two consecutive years of missing adequate yearly progress, a school was identified for improvement and had to develop an improvement plan, while the district had to offer students the option to transfer to a higher-performing school, with transportation provided, and had to reserve a portion of its Title I funds for that purpose. After a third year, the district had to offer supplemental educational services, essentially federally funded tutoring, to eligible students from low-income families, reserving 20 percent of its Title I Part A allocation for the combination of choice-related transportation and tutoring. After continued failure, the school entered corrective action, a menu of federally defined steps that included replacing staff, implementing a new curriculum, extending the school day or year, and changing the school’s internal organization. After further failure came restructuring: the statute required the district to implement one of a defined set of governance changes, including reopening the school as a charter school, replacing most or all of the staff, contracting with a private management organization, or turning operation of the school over to the state. Each rung was more intrusive than the last, each was defined in federal law rather than left to local judgment, and each was triggered by the measurement results without any intervening exercise of discretion about what the particular school needed.
The ladder’s design reflected a theory of action: that schools failing their students needed external pressure of increasing severity, applied uniformly, and that local officials left to their own devices would avoid the hard decisions. The theory was not without evidence in its favor, and the article will return to what the research found. But the ladder’s operation in practice revealed three weaknesses that the later statute was written to address. First, the triggers multiplied until the system lost its power to distinguish. As the 2013-2014 deadline approached and adequate yearly progress became harder to make, the number of identified schools grew so large that the interventions designed for a minority of struggling schools were being applied to a substantial share of all schools, diluting both the resources and the stigma that made the early rungs meaningful. Second, the lowest rungs, the ones that reached students directly, had strikingly low take-up. Federal evaluations reported that roughly 1 percent of eligible students used the school choice option and roughly 17 percent of eligible students received supplemental educational services. A remedy that 99 percent of its intended beneficiaries do not use is a remedy in name more than in fact, and the gap between the statute’s promise and the families’ experience became one of the most documented failures of the earlier regime. Third, the higher rungs demanded governance changes that many districts lacked the capacity or the political will to execute well; restructuring on paper often meant restructuring in name, with the same staff and the same practices under a new organizational label.
The later statute replaced the ladder with a system of identification plus state-designed intervention. Section 1111(c)(4)(D) and Section 1111(d) require states to identify three categories of schools. Comprehensive support and improvement schools comprise the lowest-performing 5 percent of Title I schools in the state, all public high schools with four-year graduation rates at or below 67 percent, and schools with additional targeted support needs that have not improved within a state-determined number of years not to exceed four. Targeted support and improvement schools are those in which one or more subgroups of students are consistently underperforming, as defined by the state. Additional targeted support and improvement applies where a subgroup on its own performs at or below the level of the lowest-performing 5 percent of Title I schools, triggering the comprehensive support requirements for that subgroup’s needs. The categories are federally defined: no state may decline to identify its lowest-performing 5 percent of Title I schools, no state may exempt a high school graduating two-thirds or fewer of its students in four years, and no state may ignore a subgroup performing at the level of the state’s worst schools. The identification requirements are, in an important sense, more prescriptive than anything in the earlier law’s identification stage, because they name exact thresholds, the 5 percent and the 67 percent, that every state must apply.
What happens after identification is where the later law departs from its predecessor. For comprehensive support schools, the district must develop and implement an improvement plan in partnership with stakeholders, and the plan must include evidence-based interventions. For targeted support schools, the school itself develops the plan with similar requirements. The state must establish statewide exit criteria for identified schools, monitor the plans, and take action where schools fail to improve, including more rigorous state-determined interventions for comprehensive support schools that do not meet exit criteria within four years. But the content of the interventions, the actual changes in staffing, curriculum, schedule, or governance, is designed by the state and the district, not prescribed by Congress. The federal statute requires that there be a plan, that it be evidence-based, that stakeholders participate, that the state monitor it, and that failure to improve triggers stronger state action; it does not say which curriculum to adopt, which staff to replace, or which governance model to impose. That silence is deliberate, and it is the heart of the reallocation of authority.
The evidence requirement deserves close attention because it is the later law’s substitute for the earlier law’s prescription, and it is frequently misunderstood as a mere paperwork hurdle. Section 8101(21)(A) defines four tiers of evidence: strong evidence from at least one well-designed and well-implemented experimental study; moderate evidence from at least one well-designed and well-implemented quasi-experimental study; promising evidence from at least one well-designed and well-implemented correlational study with statistical controls for selection bias; and a rationale based on high-quality research findings or positive evaluation suggesting that the intervention is likely to improve student outcomes, paired with ongoing efforts to examine its effects. Interventions funded under the school improvement provisions must meet one of these tiers, with the stronger tiers required for certain uses of funds. The tier system does not tell a district what to do; it tells the district what quality of reason it must have for doing it. A district that wants to extend the school day must be able to point to research meeting one of the tiers; a district that wants to replace the reading curriculum must do the same. The requirement is a genuine constraint, and in one respect it is more demanding than the earlier law’s approach: the earlier law prescribed interventions without requiring that they be evidence-based, so a district could implement a federally mandated restructuring with no research basis at all, while the later law permits any intervention but demands a research basis for each.
Which regime gave districts more freedom to choose interventions?
The later one, by design. The 2002 law named the interventions: choice, tutoring, corrective action, and restructuring options, each triggered by consecutive missed targets. The 2015 law names only the schools to identify and requires evidence-based plans designed by states and districts. Freedom expanded inside federal guardrails: fixed identification thresholds, stakeholder participation, state monitoring, and four evidence tiers.
The contrast on this axis can be stated as a pair of sentences that the reader should carry forward. Under the earlier law, Washington decided which schools were in trouble using state-defined targets and a federal deadline, and Washington prescribed what happened to them. Under the later law, Washington requires states to identify troubled schools using federally defined categories and thresholds, and the states and districts design what happens to them subject to federal evidence standards. The earlier model bet that uniformity of response would overcome local evasion; the later model bets that local knowledge will produce better responses than federal prescription, provided the identification is honest and the evidence standard is real. Which bet was wiser is the subject of the verdict section, but the structure of the bets is not a matter of opinion; it is written in the statutes. Readers seeking the later statute’s full architecture, including the state plan provisions that house these identification and intervention rules, can turn to the Every Student Succeeds Act of 2015 guide, which presents the 2015 reauthorization in the order its titles operate.
Axis four: what the Secretary may not do
The fourth axis is the one that most sharply distinguishes the two regimes, and it exists because of what happened between 2011 and 2015. Under the earlier statute, the Secretary of Education possessed broad authority to grant waivers of statutory and regulatory requirements, and the executive branch used that authority in a way that transformed the law’s operation without amending it. On September 23, 2011, the administration announced that it would offer waivers from the 2002 law’s central requirements, including the 2013-2014 proficiency deadline, to states that agreed to adopt a set of favored policies. The first waivers were approved on February 9, 2012, for ten states: Colorado, Florida, Georgia, Indiana, Kentucky, Massachusetts, Minnesota, New Jersey, Oklahoma, and Tennessee. Eventually 43 states plus the District of Columbia and Puerto Rico received waivers, according to the Department of Education’s own accounting. The waiver conditions required states to adopt college- and career-ready academic standards, to implement teacher and principal evaluation systems based in significant part on student growth, and to design their own differentiated accountability systems identifying priority and focus schools. In substance, the executive branch replaced the statute’s accountability regime with a new one of its own design, using the waiver authority as the instrument and the states’ desire for relief from the impossible deadline as the leverage.
The waiver era had defenders and detractors, and this article takes no position on the wisdom of the policies the waivers promoted. The institutional fact that matters for the comparison is different: the executive branch demonstrated that broad waiver authority, combined with an unworkable statutory deadline, could be used to condition relief on the adoption of specific standards, assessments, and personnel policies that Congress had never enacted. Whether one regards the resulting policies as good or bad, the mechanism concentrated in the Secretary a power that the constitutional design assigns to the legislature: the power to decide what the law requires in exchange for not enforcing what it says. Members of Congress from both parties objected to the mechanism even when they supported some of the policies, and the objection became one of the principal motors of the 2015 rewrite. The later statute is, among other things, a legislative reclamation of authority from the executive, and its prohibitions on the Secretary are written with the waiver era plainly in view.
Did waivers from the earlier law cause the rewrite?
The waiver era explains the 2015 law’s provision, though the rewrite had many causes. Between 2011 and 2015, the executive branch conditioned relief from the 2002 law’s targets on state adoption of favored policies, including specific standards and teacher evaluation systems. Congress responded with express prohibitions: the Secretary may not condition plan approval on particular standards, assessments, or evaluation designs.
The prohibitions are specific, numerous, and deliberately redundant, and they repay close reading because they are the provision that most sharply distinguishes the two regimes. Section 1111(e)(1)(B)(ii) bars the Secretary from conditioning approval of a state plan, or approval of any waiver, on a state’s adoption of specific academic standards or specific elements of standards. Section 1111(e)(1)(B)(iii), in subclauses (IX) and (X), bars the Secretary from conditioning approval on the state’s adoption of particular teacher, principal, or other school leader evaluation systems or on particular measures of their effectiveness. Section 1111(j) provides that the Secretary may not attempt to influence, incentivize, or coerce state adoption of the Common Core State Standards or any other standards, nor state membership in any particular assessment consortium or partnership; this provision addresses standards, assessments, and partnerships, and the bars that concern accountability components as such sit in Section 1111(e)(1)(B)(iii) rather than in Section 1111(j), a distinction worth preserving because the two provisions cover different objects. Beyond Title I, Section 8526A, codified at 20 U.S.C. Section 7906a, imposes an act-wide prohibition on the Secretary mandating, directing, or controlling a state’s or district’s instructional content, academic standards, assessments, curricula, or programs of instruction, expressly including the use of waiver conditions under Section 8401 to do so. Section 1604, codified at 20 U.S.C. Section 6575, imposes a parallel bar scoped to Title I. And Section 8401, codified at 20 U.S.C. Section 7861, retains the Secretary’s general waiver authority but fences it: waivers may not touch allocations to states and districts, civil rights protections, and other enumerated categories, and the exercise of the authority is subject to the prohibitions just described.
The cumulative effect is a statute that distrusts its own administrator. The earlier law gave the Secretary tools and trusted the office to use them well; the later law gives the Secretary tools and writes into the law, repeatedly and in overlapping terms, the things the office may not do with them. The contrast is not merely a matter of tone. Under the earlier regime, a Secretary who wished to steer state policy could do so through conditions on waivers and approvals, and did. Under the later regime, the same steering is unlawful on its face, and a state presented with such conditions has statutory text to quote in refusal. The prohibitions are the later law’s answer to the constitutional anxiety the waiver era produced, and they are the single clearest textual evidence for the claim that the 2015 rewrite was about the allocation of authority rather than about the volume of federal requirements. A law that multiplies prohibitions on the federal executive is not deregulating; it is re-regulating the regulator.
There is a coda to the discretion story that belongs here because it completes the picture of how each regime handled the gap between statutory text and administrative reality. Under the earlier law, when the deadline became unworkable, the executive branch improvised a new regime through waivers because the statute offered no lawful path to relief; the improvisation was creative, effective in the short term, and corrosive to the statute’s authority in the long term. Under the later law, Congress wrote the relief into the statute itself: state-designed goals replaced the federal deadline, state-designed interventions replaced the federal ladder, and the Secretary’s discretion was fenced by express prohibitions. The later law can be read as Congress’s admission that the earlier law’s design had depended on an administrator’s willingness to bend it, and as Congress’s decision to write a law that did not require bending. Whether the new design works better is a separate question, addressed in the verdict; that it was designed to be operable as written is a fact about its architecture, and it distinguishes the two regimes as clearly as any provision in either statute.
What the waivers actually required
The waiver era is often summarized as the executive branch letting states off the hook, but the reality was a bargain with detailed terms, and the terms explain why Congress wrote the prohibitions it did. The September 2011 offer relieved states from a specific list of the 2002 law’s requirements: the 2013-2014 deadline for 100 percent proficiency, the adequate yearly progress determinations built on it, and the prescribed sanctions of improvement, corrective action, and restructuring, including the choice and tutoring mandates. In exchange, each applicant state had to commit to three principles, and the Department of Education reviewed the applications against detailed criteria for each.
The first principle was college- and career-ready expectations for all students. States had to adopt academic standards deemed to prepare students for postsecondary education and the workforce, or demonstrate that their existing standards met the bar, and had to administer aligned assessments, with most states joining one of two federally funded assessment consortia. The second principle was state-developed differentiated recognition, accountability, and support: each state had to design its own system for identifying reward schools, priority schools (the lowest-performing 5 percent of Title I schools, a concept the 2015 law later adopted in statutory form), and focus schools (those with the largest achievement gaps or low subgroup performance), and had to implement interventions in the priority and focus schools. The third principle was supporting effective instruction and leadership: states had to develop and implement teacher and principal evaluation and support systems that used multiple measures, including student growth as a significant factor, to rate educators and to inform personnel decisions.
Each principle carried extensive sub-requirements, and the application process gave the Department continuing leverage: states submitted detailed plans, received feedback, revised, and operated under monitoring, with the possibility of losing the waiver if they failed to implement their commitments. The arrangement was, in effect, a second federal accountability law, designed by the executive branch, applicable to 43 states plus the District of Columbia and Puerto Rico, and resting entirely on the Secretary’s waiver authority rather than on any act of Congress. Its defenders argued that it rescued states from an impossible deadline while advancing worthy policies; its detractors argued that it replaced legislation with administration. For the comparison, the important point is institutional rather than ideological: the waiver system proved that the 2002 law’s combination of an unworkable deadline and broad waiver authority handed the executive a standing invitation to write education policy by condition, and the invitation would remain open for every future Secretary unless Congress closed it. The 2015 law’s overlapping prohibitions are the closing of that invitation, written provision by provision against the three principles: no conditions on standards, no conditions on assessments or consortia, no conditions on evaluation systems. A reader who grasps the waiver bargain understands the prohibitions not as boilerplate but as a point-by-point reversal.
How adequate yearly progress actually worked
To understand why the earlier law broke down, a reader must understand the machinery of adequate yearly progress beyond the slogan of the deadline. Each state set its own starting point, generally based on the performance of its lowest-achieving demographic group or its lowest-achieving schools in the baseline year, and then drew a trajectory of annual measurable objectives rising to 100 percent proficiency by the end of the 2013-2014 school year. The trajectory had to include intermediate goals, typically with the steepest gains scheduled for the later years, which meant the required annual improvement grew larger exactly as the deadline neared and the remaining non-proficient students became harder to move. A school made adequate yearly progress only if every subgroup met the annual objective in both reading and mathematics, tested at least 95 percent of each subgroup, and met the state’s other academic indicator, usually graduation rate for high schools and attendance for elementary and middle schools. A single subgroup missing a single subject in a single year was enough to fail the school.
The statute and the regulations softened the edges of that unforgiving structure with three devices. Safe harbor allowed a school to make adequate yearly progress even while missing the annual objective, provided it reduced the percentage of students below proficient by at least 10 percent from the prior year and met the other indicators; the provision recognized improvement without requiring the full scheduled gain. Confidence intervals allowed states to treat small misses as statistical noise, acknowledging that a school’s measured proficiency rate is an estimate with a margin of error. Minimum subgroup sizes, the n-size, allowed states to exclude very small groups from the subgroup calculations, on the ground that the performance of a handful of students should not determine a school’s rating. Each of these devices was defensible in isolation, and each created an avenue for softening the law’s bite: states set their own n-sizes, some as high as 50 students, which excluded large numbers of minority and disabled students from subgroup accountability in small schools; confidence intervals varied by state; and safe harbor, while pedagogically sensible, meant the trajectory’s discipline depended on how the state implemented the exception.
The combined effect was a system that was simultaneously rigid and manipulable. The rigidity came from the all-subgroups rule and the fixed endpoint: as the deadline approached, the mathematics of the trajectory required gains that grew each year, and schools serving the most disadvantaged populations faced the steepest required climbs because they started furthest from the endpoint. The manipulability came from the state-set parameters: the definition of proficiency, the n-size, the confidence interval, and the shape of the trajectory were all state choices, and each choice could make adequate yearly progress easier or harder to achieve without changing anything in a classroom. The National Center for Education Statistics mapping studies caught the most visible form of the manipulation, the divergence between state-reported trends and National Assessment trends, but the subtler forms, a high n-size here, a generous confidence interval there, operated quietly in every state. The machinery did exactly what its designers feared local control would do and exactly what its defenders hoped federal pressure would prevent, at the same time, because the designers had split the difference: federal endpoints with state-defined measurement. The 2015 law’s answer was to move the endpoint to the states as well, accepting the loss of comparability in exchange for removing the incentive to game the federal target.
From deadline crisis to rewrite
The 2015 rewrite did not emerge from a single moment of decision; it accumulated over nearly a decade of failed reauthorization attempts, executive improvisation, and shifting coalitions. The earlier law was due for reauthorization in 2007, and the House education committee produced discussion drafts under Chairman George Miller and Ranking Member Buck McKeon, but no bill reached the floor in either chamber. The deadline was still years away, the law’s supporters could still point to rising scores in the early grades, and the political system found it easier to leave the statute unexamined than to renegotiate its bargains. By the time the next Congress took up the question, the deadline’s approach had changed the politics: the scheduled failure of nearly every school in America concentrated minds, and the administration of Barack Obama published A Blueprint for Reform in March 2010 proposing its own replacement framework, centered on college and career readiness, competitive grant programs, and a revised accountability structure.
Congress did not take up the Blueprint as a bill, and the reauthorization stalled again. What moved instead was the executive branch. The September 2011 waiver announcement and the approvals that followed, ten states in February 2012 and eventually 43 states plus the District of Columbia and Puerto Rico, created a de facto new regime without legislation. For the states operating under waivers, the 2002 law’s accountability provisions had been replaced by executive-designed substitutes: differentiated recognition systems, priority and focus schools, and the required policies on standards and evaluations. The waivers relieved the immediate crisis of the deadline, which reduced the pressure on Congress to act, while simultaneously demonstrating that the statute as written could not function, which increased the pressure from a different direction. The result was several years of stalemate in which the law on the books described a regime nobody operated and the regime everybody operated had never been enacted.
The stalemate broke in 2015. A conference committee reconciled the House and Senate versions into a single text, and the conference report returned to each chamber in December: the House adopted it on December 2, 2015 by 359 to 64 on Roll Call Number 665, the Senate on December 9, 2015 by 85 to 12 on Record Vote Number 334, and the President signed it on December 10, 2015 as Public Law 114-95. The bipartisan margins of the final votes reflected the breadth of the coalition against the earlier regime’s design: civil rights organizations that valued the disaggregated data joined with governors and state chiefs who wanted control of interventions, and with legislators of both parties who objected to the waiver-era concentration of power in the Secretary. The coalition agreed on what to dismantle more completely than on what to build, which is why the 2015 law is most specific about what the federal government may not do and leaves the affirmative design of improvement to the states.
The subgroup bargain: the earlier law’s most durable achievement
Beneath the fights over deadlines and sanctions, the earlier law contained a civil rights bargain that outlasted every other part of its design, and the 2015 law’s most important act of preservation was keeping it. The bargain was simple: in exchange for federal money, states would test every child every year and report the results broken out by race, ethnicity, poverty, disability, and English learner status, so that no school could hide the performance of its most vulnerable students inside a respectable average. Before that bargain, the hiding was routine. District and state reports typically published schoolwide averages, and a school could celebrate its mean score while failing its Black students, its Latino students, its students with disabilities, and its English learners at rates that would have scandalized the community had they been printed. The disaggregation requirement made the printing mandatory, and the report card requirement made the results public, which together created a form of accountability that operated through information rather than through sanction: parents could see, reporters could compare, advocates could organize, and researchers could study.
The bargain held together an unlikely coalition, and the coalition’s shape explains why the data provisions survived the rewrite untouched. Civil rights organizations defended the testing and reporting requirements even as they criticized the sanctions, because the data gave their advocacy its factual foundation; without disaggregated results, claims about inequity remained anecdotal and deniable. State chiefs and governors, who chafed at the deadline and the ladder, nevertheless valued the diagnostic power of the data for their own management. Researchers across the ideological spectrum relied on the annual, subgroup-level results as the empirical base for studying American schooling. When the 2015 rewrite was negotiated, the coalition’s message to Congress was consistent: change what happens after the scores arrive, but do not touch the scores’ production or their publication. Congress obliged, retaining the testing schedule, the subgroup definitions, and the report cards, and adding expenditure and educator data on top.
The durability of the bargain carries a lesson for the verdict. The earlier regime’s information effects, the documented movements in spending, staffing, and instructional time that Dee, Jacob, and Schwartz and Reback, Rockoff, and Schwartz recorded, all flowed from measurement pressure that the disaggregation made inescapable. A school could not evade the pressure by averaging its way out of it, because the law forbade the averaging. The later law preserved that inescapability while changing the use to which the information is put: the identified categories, the 5 percent and the 67 percent, are built directly on the disaggregated data, and the improvement plans must address the subgroup needs the data reveal. The through line from the 2002 bargain to the 2015 categories is one of the clearest continuities in the comparison, and it is why the claim that the rewrite weakened accountability for underserved students fails on the text: the instrument that made underserved students visible was never weakened, and the later law built its identification system on top of it.
The money behind the mandates
Accountability systems run on money as well as on rules, and the two laws financed their designs differently in ways that reveal their theories of action. Under the earlier law, the finance of intervention ran through set-asides from districts’ existing Title I allocations: a district with schools in improvement status had to reserve funds for choice-related transportation, and once supplemental educational services were triggered, the combined reservation reached 20 percent of the district’s Title I Part A allocation. The design meant the interventions were funded by redirecting money the district already received, which made the mandates budget-neutral for the federal government and budget-painful for the district. The pain was part of the theory: a district that lost a fifth of its compensatory education funding to transportation and tutoring had a financial incentive to get its schools out of improvement status. In practice, the low take-up of choice and tutoring, roughly 1 percent and 17 percent of eligible students respectively in federal evaluations, meant much of the reserved money went unspent on its intended purposes or was spent on services few families used, converting the financial incentive into deadweight.
The later law changed the finance along with the design. Instead of district set-asides for prescribed services, the statute requires states to reserve 7 percent of their Title I Part A allocation for school improvement activities, with the funds flowing to districts serving identified schools to support the development and implementation of improvement plans. The money follows the identification: the schools the federal categories name are the schools the improvement funds serve. The design reflects the later law’s theory that the binding constraint on struggling schools is capacity rather than will, and that resources directed to planning, evidence-based intervention, and implementation support will do more than resources reserved for transportation to schools families did not choose. Whether the theory is correct is an empirical question the research available through 2014 had not settled, but the structure of the bet is visible in the finance: the earlier law taxed districts for failure and spent the proceeds on exits, while the later law invests in the identified schools themselves.
A third financial provision belongs in the comparison because it extends the information function into dollars. The later law’s requirement that states report per-pupil expenditures at the school level, including personnel and nonpersonnel costs disaggregated by source, made resource distribution visible in the same way the earlier law had made achievement distribution visible. A district that spends substantially more per pupil in its affluent schools than in its high-need schools must publish the difference under the later law, and the publication creates the same kind of pressure the disaggregated test scores created: not a sanction, but a fact in public view that advocates, reporters, and parents can use. The expenditure reporting is the subgroup bargain applied to inputs rather than outcomes, and it is arguably the later law’s most significant addition to the measurement apparatus. A reader comparing the two regimes’ theories of change should weigh this provision carefully: the earlier law bet that publishing outcomes would move resources, and Dee, Jacob, and Schwartz documented that it did, with roughly 600 dollars in additional per-pupil spending; the later law bets that publishing resources directly will move them more efficiently.
Comparability lost: the cost of fifty different goals
Every design choice in the comparison has a price, and the price of state-designed goals is comparability: the ability to say, with a straight face, how one state’s ambitions compare with another’s. Under the earlier law, the single federal deadline supplied that ability, however crudely. Every state was climbing toward 100 percent proficiency by the end of the 2013-2014 school year, so a reader could compare trajectories, ask why one state’s adequate yearly progress record looked different from another’s, and treat the differences as information about policy rather than about definitions. The crudeness of the comparison was real, the deadline was unreachable and the definitions were gamed, but the commonness of the target made the gaming itself visible: the National Center for Education Statistics mapping studies could show that state-reported trends outpaced National Assessment trends precisely because there was a federal expectation against which to check the states.
Under the later law, each state sets its own long-term goals, its own interim measures, its own timeline, and, within federal parameters, its own definition of consistent underperformance for targeted support. A reader who wants to compare state ambition must read fifty plans with fifty vocabularies, fifty timelines, and fifty theories of what counts as progress. The federal requirements that goals cover all students and each subgroup, address achievement, graduation, and English learner progress, and include interim measurements keep the plans honest in a formal sense, but they do not make them comparable in substance. A state with ambitious goals and a state with modest ones both comply, and the statute provides no mechanism for ranking their ambition or for pressing the modest state to aim higher. The loss is not merely academic. Comparability is the mechanism by which public pressure crosses state lines: the reform that embarrasses one legislature into action is usually the neighboring state’s better record, and better records require common measures.
The later law’s defenders have two answers, and the honest comparison registers both. The first is that the National Assessment remains, with mandatory state participation in grades 4 and 8, so a common yardstick still exists for comparing outcomes even if the goals differ; the mapping function the assessment served under the earlier law, checking state-reported results against an independent measure, continues to operate. The second is that the earlier law’s comparability was purchased with perverse incentives, the downward redefinition of proficiency that the mapping studies documented, so the common measure was corrupting the very thing it measured; comparability bought at the price of honesty is a bad bargain. Both answers have force, and the verdict’s deciding factor is built to hold them: the reader who values comparability across states should prefer elements of the earlier model, and the reader who values locally designed, honestly measured improvement should prefer the later one. The price is real either way; the comparison’s job is to name it, not to pretend it does not exist.
The graduation rate: from side indicator to central goal
The treatment of high school graduation illustrates, in miniature, the whole movement from the earlier regime to the later one. Under the 2002 law, graduation rate served as the “other academic indicator” for high schools, the non-test measure that supplemented assessment results in the adequate yearly progress calculation. Its role was subordinate: a high school could meet its test-score targets and still fail adequate yearly progress on graduation, but the rate itself was not the law’s organizing goal, and for years states calculated it in inconsistent ways that made comparisons unreliable. Federal rules later standardized the four-year adjusted cohort method, counting only students who entered ninth grade together and graduated within four years, which made the rates comparable across states and exposed how many schools had been flattering their holding power with looser definitions.
The 2015 law promoted graduation from supporting indicator to required long-term goal. Section 1111(c)(4)(A) requires states to set long-term goals for the four-year adjusted cohort graduation rate, with interim measurements of progress, and the statute also calls for reporting an extended-year rate that credits schools for students who need more than four years. More consequentially, the 67 percent threshold made graduation an identification trigger: any public high school with a four-year rate at or below 67 percent enters comprehensive support and improvement by federal command, regardless of its test scores. A high school can no longer offset catastrophic dropout rates with acceptable assessment results; the statute treats a school that graduates two-thirds or fewer of its students as failing on that fact alone. The shift reflects a judgment about what accountability is for: the earlier law treated the diploma as context for the test scores, while the later law treats keeping students through graduation as an outcome coequal with achievement. For the comparison’s verdict, the graduation provisions strengthen the later law’s claim to surface inequity, because dropout is concentrated among the same disadvantaged groups the disaggregated test data track, and a school that loses a third of its students before graduation is failing precisely those students.
English learners: from separate title to core indicator
Students learning English occupied a different statutory location under each regime, and the move from one location to the other is among the 2015 law’s most consequential quiet changes. Under the 2002 law, English learner progress was tracked primarily under Title III, the language instruction title, through annual measurable achievement objectives that operated separately from the Title I accountability system. The separation meant that a school’s obligations to its English learners lived in a different part of the statute from its obligations to its other subgroups, with different metrics, different timelines, and less public visibility. The Title I report cards disaggregated results for English learners, so the subgroup was visible in the data, but the accountability consequences for the subgroup’s progress ran through the Title III channel.
The 2015 law moved English language proficiency into the core. Progress in achieving English language proficiency is one of the required indicators for the state’s long-term goals under Section 1111(c)(4)(A), standing alongside academic achievement and graduation rates, which means every state’s accountability system must set goals for how quickly its English learners gain proficiency and must measure interim progress toward them. The statute also requires states to establish standardized statewide procedures for identifying English learners and for exiting them from language programs, replacing the patchwork of district-level criteria that had made the population itself inconsistently defined. The standardization matters because an accountability indicator is only as honest as the group it measures: when districts define entry and exit differently, the measured population shifts with the definitions, and apparent progress can reflect reclassification rather than learning. By fixing the definitions at the state level and placing the progress indicator in Title I, the later law gave English learners the same structural position as every other subgroup: the same goals, the same interim measurements, the same public report cards, and the same role in triggering targeted support when the subgroup consistently underperforms. The change is easy to overlook beside the headline fights over testing and the Secretary’s powers, but for the students it concerns, the move from a separate title to the core accountability system is the difference between being tracked and being counted.
Teachers and leaders: the evaluation fight
Few provisions better illustrate the reallocation of authority than the two laws’ treatment of teacher quality, because the subject traveled from federal prescription to prohibited federal condition in a single rewrite. The 2002 law required that all teachers of core academic subjects be “highly qualified” by the end of the 2005-2006 school year, a federal definition demanding a bachelor’s degree, full state certification, and demonstrated subject-matter competence. The requirement was among the law’s most criticized: states complied formalistically, the definition said little about classroom effectiveness, and the deadline passed with widespread paper compliance. The waiver era then pushed further, conditioning relief on state adoption of teacher and principal evaluation systems that rated educators using multiple measures with student growth as a significant factor, which made evaluation design a condition of federal favor on a national scale.
The 2015 law dismantled both the prescription and the condition. It repealed the highly qualified teacher definition, returning the determination of teacher qualifications to the states, while retaining the requirement that states report on educator qualifications, including the distribution of inexperienced, out-of-field, and ineffective teachers across high- and low-poverty schools. And it wrote the prohibitions discussed in the discretion section: the Secretary may not condition plan or waiver approval on any particular teacher, principal, or school leader evaluation system or on any particular measure of effectiveness, under Section 1111(e)(1)(B)(iii)(IX) and (X). The combination is characteristic of the later law’s design philosophy. The federal government kept the information function, requiring public reporting on where the least experienced and least credentialed teachers work, and surrendered the design function, forbidding the Secretary from steering how states evaluate their educators. A reader comparing the two regimes should see the symmetry with the testing story: measurement retained, response released, and the executive fenced. The teacher provisions are also the clearest evidence that the rewrite was not a simple victory for any single interest: the same law that freed states from the federal qualification definition also forced them to publish data on teacher distribution that many would have preferred to keep private.
Assessment quality: who vouches for the tests
A comparison of accountability regimes that ignored the quality of the assessments themselves would miss the foundation on which everything else rests, because targets, identifications, and interventions are all computed from test scores. Under both laws, the federal government’s main instrument for assuring assessment quality was peer review: states submitted evidence that their assessments met statutory requirements for alignment to academic standards, technical quality, and inclusion of all students, and Department review teams evaluated the submissions. The process was bureaucratic and often slow, but it was the mechanism by which Washington checked the instruments without writing them, a division of labor both Congresses preserved.
The assessment landscape changed between the two enactments in ways the statutes registered differently. In the years after 2009, two multistate consortia built common assessments with federal grant support, and many states adopted them, which briefly promised a degree of cross-state comparability the earlier law’s fifty separate testing programs had never achieved. States then withdrew from the consortia in large numbers, returning the country to a patchwork of state-specific tests. The 2015 law responded to that history with the prohibitions already described: the Secretary may not coerce membership in any assessment partnership or consortium under Section 1111(j), and may not condition approval on particular assessments under Section 1111(e)(1)(B). The later law thus guarantees state control of assessment choice while retaining the federal peer review of assessment quality, the same pattern that governs every other axis: the state chooses, the federal government checks the choice against statutory criteria, and the Secretary may not use the checking process as leverage for favored policies. For the verdict, the assessment story reinforces the deciding factor. The earlier regime’s federal leverage over assessments, exercised most aggressively through the waiver conditions, produced common instruments at the cost of executive overreach; the later regime’s prohibitions protect state choice at the cost of the comparability the common instruments briefly offered.
The report card as a policy instrument
Both laws treat the public report card as a policy instrument in its own right, not merely as a compliance document, and the growth of the report card from one regime to the next is the clearest evidence that the later law took information more seriously than its deregulation caricature allows. Under the 2002 law, the report card carried the disaggregated assessment results, the adequate yearly progress determinations, and the teacher qualification data, and its publication was the mechanism by which the law’s pressure reached beyond administrators to parents, journalists, and advocates. The report card was the part of the statute ordinary citizens actually encountered: the document that told a parent whether the school’s averages concealed subgroup failure, and the document reporters used to rank schools and districts.
The 2015 law kept every element of the earlier report card and added categories that extended public view from outcomes to inputs and from single-year snapshots to trajectories. Per-pupil expenditure reporting, broken out by source and published at the school level, exposed resource distribution within districts. Educator qualification reporting, including the distribution of inexperienced, out-of-field, and ineffective teachers between high- and low-poverty schools, exposed staffing inequities. The required reporting of English learner progress toward proficiency, of four-year and extended-year graduation rates, and of assessment participation rates filled gaps the earlier card had left. The cumulative effect is a document that answers more questions than its predecessor: not only how each group performed, but what resources stood behind the performance, who taught the students, and how many students were actually tested.
The theory behind the instrument deserves explicit statement because it underwrites the verdict. Informational accountability assumes that facts in public view change behavior without any sanction being imposed: that a superintendent who must publish subgroup gaps will act on them, that a school board that must publish expenditure gaps will defend or close them, and that the publication itself disciplines the system by making evasion visible. The earlier law tested that theory on outcomes, and the mechanism studies found that it moved spending, staffing, and time. The later law extended the theory to inputs, betting that publishing resource data will move resources the way publishing outcome data moved them. The bet was unproven in the research available through 2014, but the direction of change is unambiguous: a Congress engaged in deregulation does not expand the public’s right to know, and the report card’s growth from one law to the next is incompatible with the claim that the federal role shrank.
What the two laws assume about why schools fail
Beneath the provisions about testing, targets, and interventions, the two statutes embody different theories of why schools fail, and naming the theories clarifies what is truly at stake in the comparison. The earlier law assumed evasion. Its designers believed that schools and districts knew which students were failing and chose, through inattention or misaligned incentives, not to act; the remedy was therefore pressure, applied uniformly and escalated automatically, that made inaction more costly than action. The deadline, the all-subgroups rule, and the federally prescribed ladder all follow from that assumption: if the problem is will, the solution is consequences the local actors cannot dodge. The theory had a distinguished pedigree in the economics of regulation, and the mechanism studies partly vindicated it: when the pressure arrived, systems moved money, personnel, and time toward the measured goals.
The later law assumes incapacity compounded by misdirected effort. Its designers believed that struggling schools generally wanted to improve but lacked the knowledge, the resources, or the organizational capacity to do so, and that federally prescribed interventions had compounded the problem by forcing districts to implement remedies disconnected from their circumstances. The remedy was therefore diagnosis plus supported design: federally mandated identification to ensure the struggling schools are named, evidence tiers to ensure the chosen remedies have a research basis, stakeholder participation to ensure local knowledge enters the plan, state monitoring to ensure follow-through, and improvement funds directed to the identified schools. The assumption is visible in every structural choice: the state designs the goals because the state knows its schools, the district designs the intervention because the district knows its constraints, and the federal government checks the evidence because enthusiasm is not a substitute for research.
Neither theory is complete, and the honest verdict says so. Some schools fail from evasion, and for those schools the earlier law’s automatic pressure was the right instrument; the later law’s reliance on state monitoring and plan quality offers such schools more room to drift. Some schools fail from incapacity, and for those schools the earlier law’s prescribed ladder was the wrong instrument; a restructuring mandate does not teach a district how to teach reading. The deciding factor, that the earlier regime surfaced inequity while the later regime bets on better responses, is another way of stating that each law is built for a different diagnosis of failure. A reader choosing between the models is choosing between theories of the problem, and the comparison is finished when the reader can state both theories, match each to its statute, and defend the choice with the evidence rather than with a slogan about federal overreach.
The long view: what survived, what did not, and what it means
With the four axes, the machinery, the history, and the evidence all on the table, the comparison can be restated as a ledger of survival, and the ledger is the most compact form of the One Test. What survived the 2015 rewrite: annual testing in reading and mathematics in grades 3 through 8 and at least once in high school, identical in grades, subjects, and cadence; disaggregated subgroup reporting by race, ethnicity, economic disadvantage, disability status, and English learner status, expanded with expenditure and educator data; the 95 percent participation rule for all students and each subgroup; mandatory state participation in the National Assessment in grades 4 and 8; the public report card as the instrument of informational accountability; and the federal insistence that struggling schools be identified and named. What did not survive: the single federal deadline of 100 percent proficiency by the end of the 2013-2014 school year; adequate yearly progress as the annual calculation; the Section 1116 ladder of choice, tutoring, corrective action, and restructuring; the highly qualified teacher definition; and the Secretary’s discretion to condition waivers and approvals on favored standards, assessments, and evaluation systems. The ledger divides cleanly along the line the namable claim draws: measurement survived, response changed authors.
The pattern of the ledger carries an institutional lesson that extends beyond education policy. Reauthorizations are rarely clean breaks; they are negotiations between the coalition that defends the old law’s achievements and the coalition that attacks its failures, and the resulting statute typically preserves what both coalitions value while replacing what only one defended. The disaggregated data survived because civil rights advocates, researchers, and even the law’s critics agreed the data were valuable; the deadline died because nobody outside the statute’s text could defend it; the ladder died because its remedies had failed in practice; and the Secretary’s discretion died because the waiver era had taught legislators of both parties to distrust it. The 2015 law is thus best read not as a repudiation of its predecessor but as an edited version of it, with the edits concentrated exactly where the evidence of failure was strongest and the preservation concentrated exactly where the evidence of value was strongest. That is how legislatures learn, when they learn: not by starting over, but by keeping the instruments that worked and replacing the theories that did not.
The comparison also clarifies what the next argument will be about, and a reader who has followed the four axes can anticipate it. The later law’s design concentrates risk in two places: the quality of state-designed goals and interventions, and the honesty of state implementation. Where a state sets ambitious goals, designs evidence-based plans, monitors them seriously, and directs improvement funds to the schools that need them, the later regime should outperform the earlier one, because locally fitted, research-grounded responses beat federally prescribed formalism. Where a state sets modest goals, approves thin plans, monitors loosely, and treats identification as paperwork, the later regime will underperform the earlier one, because the earlier law’s automatic triggers and common deadline at least forced the uncomfortable facts into view on a fixed schedule. The federal guardrails, the fixed identification thresholds, the evidence tiers, the expanded reporting, and the fenced Secretary, are Congress’s attempt to bound that risk without reimposing the prescription that failed. Whether the guardrails hold is the question on which the later law will ultimately be judged, and it is a question about implementation rather than about text: the statute did its part by writing the requirements down, and the rest belongs to the states, the districts, and the public that reads their report cards.
That allocation of responsibility is the final meaning of the shift from federal prescription to state design. The earlier law told the country what to do and discovered that telling was not enough; the later law tells the country what to measure, whom to identify, what standards of evidence to meet, and what the federal executive may not do, and leaves the doing to the governments closest to the schools. Same measurement, different response: the two statutes measure schools almost identically and differ almost entirely in who decides what happens next, so debates framed as testing versus no testing are describing a difference that does not exist in the statutes. The reader who can state the ledger, explain the reallocation, and defend a verdict with the deciding factor named has passed the One Test, and has understood the comparison every education student, journalist, and school administrator needs.
A final check: the verification flags answered
The brief for this comparison named three verification flags, and the article closes its pre-table argument by answering each directly against the statutory text. First, the testing grade requirements in both statutes. The earlier law required annual reading and mathematics assessments in each of grades 3 through 8, fully implemented by the end of the 2005-2006 school year, with additional assessments at least once in the grades 10 through 12 band. The later law requires the same annual assessments in grades 3 through 8 and at least once in grades 9 through 12, under Section 1111(b)(2)(B)(v)(I), codified at 20 U.S.C. Section 6311. The grades, the subjects, and the annual cadence match exactly; the high school requirement differs only in the band’s lower bound, ninth grade versus tenth, and the substance, at least one administration during the high school years, is identical.
Second, the identification categories. The later law’s categories are defined in Section 1111(c)(4)(D) and Section 1111(d): comprehensive support and improvement covers the lowest-performing 5 percent of Title I schools, every public high school with a four-year graduation rate at or below 67 percent, and schools with additional targeted support needs that fail to improve within the state’s timeframe capped at four years; targeted support and improvement covers schools where one or more subgroups consistently underperform as defined by the state; additional targeted support and improvement covers schools where a subgroup performs at or below the level of the lowest 5 percent of Title I schools, triggering comprehensive support requirements. Each category and each threshold appears in the statutory text as stated.
Third, the citations for the restrictions on the Secretary. The standards bar sits in Section 1111(e)(1)(B)(ii); the evaluation system and effectiveness measure bars sit in Section 1111(e)(1)(B)(iii)(IX) and (X); the Common Core and assessment consortium coercion bar sits in Section 1111(j), which addresses standards, assessments, and partnerships rather than accountability components as such; the act-wide bar on mandating instructional content, standards, assessments, and curricula, including through waiver conditions, sits in Section 8526A, codified at 20 U.S.C. Section 7906a; the Title I-scoped parallel bar sits in Section 1604, codified at 20 U.S.C. Section 6575; and the retained but fenced waiver authority sits in Section 8401, codified at 20 U.S.C. Section 7861. With the flags answered, the comparison proceeds to its summary instrument: the requirement-by-requirement table.
The questions this comparison leaves open
A comparison resolved with a verdict should also name what it cannot resolve, and three questions sit beyond the reach of the statutes and the research alike. The first is whether the evidence tiers will discipline local choices in practice. The four tiers of Section 8101(21)(A) are precise on paper, from experimental studies down to a rationale with ongoing evaluation, but their bite depends on who reviews the district’s evidence, how rigorously, and with what consequence for thin justifications. The earlier law’s history counsels caution: formalistic compliance, plans written to satisfy reviewers rather than to guide action, was the failure mode of the old regime, and no statutory tier system is immune to it. The second open question is whether state monitoring of improvement plans will be real. The later law requires states to monitor implementation, to establish exit criteria, and to take more rigorous action where schools miss those criteria, but monitoring quality varies with state capacity, and the statute’s silence on the content of the more rigorous action leaves the escalation to the same state officials whose earlier efforts the identification was meant to check.
The third open question is the deepest, and it returns to the deciding factor. The earlier regime’s information effects are documented: disaggregated reporting forced uncomfortable facts into public view, and the mechanism studies recorded the resulting movements in spending, staffing, and instructional time. The later regime’s response effects are prospective: state-designed, evidence-based interventions should, in theory, outperform federally prescribed ones, but the national evidence for that proposition did not yet exist in the research available through 2014. A reader who prefers the earlier model is therefore standing on documented ground, while a reader who prefers the later model is standing on a well-reasoned bet. Both positions are defensible, and the comparison’s verdict is structured to honor both: comparability and forced transparency on one side, local fit and evidence discipline on the other, with the reader’s values breaking the tie. What neither side can claim is that the statutes are ambiguous about what they require. The testing is identical, the reporting is broader, the identification is federally fixed, the interventions are locally designed, and the Secretary is fenced. The open questions concern what governments and educators will do with those requirements, which is where every accountability system, in the end, is tested.
The requirement-by-requirement table
| Requirement | Public Law 107-110 (2002) | Public Law 114-95 (2015) |
|---|---|---|
| Testing | Annual reading and mathematics assessments in grades 3 through 8 by the end of the 2005-2006 school year, plus assessments at least once in the grades 10 through 12 band | Annual reading and mathematics assessments in grades 3 through 8 and at least once in grades 9 through 12, under Section 1111(b)(2)(B)(v)(I) |
| Reporting | Assessment results disaggregated by race, ethnicity, economic disadvantage, disability status, and English learner status, published in public report cards | Same disaggregation and public report cards, plus added categories including school-level per-pupil expenditure data and educator qualification information |
| Goal setting | Single federal deadline: 100 percent proficiency for all students and each subgroup by the end of the 2013-2014 school year, enforced through adequate yearly progress | State-designed long-term goals with interim measurements of progress for all students and each subgroup, covering achievement, graduation rates, and English learner progress; no federal deadline |
| Identification | Schools missing adequate yearly progress for consecutive years identified for improvement, corrective action, and restructuring under Section 1116 | Federally defined categories: comprehensive support and improvement (lowest-performing 5 percent of Title I schools, high schools at or below 67 percent four-year graduation rate, unimproved additional targeted support schools), targeted support and improvement (consistently underperforming subgroups), additional targeted support and improvement (subgroup at or below the lowest 5 percent level) |
| Intervention design | Federally prescribed ladder: school choice with transportation, supplemental educational services with 20 percent of Title I Part A reserved, corrective action menu, defined restructuring options | State and district designed improvement plans with stakeholder participation, required to use evidence-based interventions meeting the four tiers of Section 8101(21)(A); state monitoring and stronger state action for schools missing exit criteria |
| Federal discretion | Broad waiver authority used from 2011 to condition relief on state adoption of favored standards, assessments, and evaluation systems | Express prohibitions: the Secretary may not condition plan or waiver approval on specific standards, assessments, or evaluation systems under Section 1111(e)(1)(B); may not coerce Common Core or assessment consortium membership under Section 1111(j); act-wide bar on mandating content, standards, assessments, or curricula under Section 8526A; fenced waiver authority under Section 8401 |
Reading the identification categories closely
The identification system of the 2015 law repays close reading because its precision is the instrument by which the federal government kept its grip on the information function while releasing its grip on the response. Comprehensive support and improvement has three doors, and each is worth examining. The first door, the lowest-performing 5 percent of Title I schools, is a relative measure: it identifies the bottom of the state’s own distribution regardless of the state’s absolute performance level. A state with high achievement overall still has a bottom 5 percent, and the statute requires their identification. The relative design prevents a state from defining its way out of the obligation; no cut score, no trajectory, no redefinition can eliminate the bottom of a distribution. The second door, high schools with four-year graduation rates at or below 67 percent, is an absolute measure, and the choice of 67 percent reflects a judgment about the point at which a school’s holding power has failed so completely that the diploma has lost its meaning for a third of the entering class. The third door, schools identified for additional targeted support that do not improve within the state’s timeframe of no more than four years, creates a ratchet: subgroup-level failure that persists becomes school-level failure by operation of law, and the timetable is capped so that states cannot extend the improvement period indefinitely.
Targeted support and improvement, the middle category, operates on subgroups rather than schools. A school enters the category when one or more of its subgroups consistently underperform, with the state defining the consistency standard. The design preserves the earlier law’s central insight, that averages conceal, while changing the consequence: instead of the whole school entering the federal ladder because one subgroup missed a target, the school writes a targeted plan addressing the subgroup’s needs. Additional targeted support sharpens the point: when a subgroup’s performance falls to the level of the lowest 5 percent of Title I schools, the subgroup’s needs trigger the comprehensive support machinery. The categories thus form a graduated system, from subgroup plans through subgroup-triggered comprehensive support to schoolwide comprehensive support, with the federal statute fixing the thresholds at each step and the state designing the response at each step.
The machinery around identification matters as much as the categories. States must establish exit criteria, the conditions under which an identified school leaves the category, and must monitor implementation of the improvement plans. For comprehensive support schools that fail to meet exit criteria, the state must take more rigorous state-determined action, a requirement that preserves an escalation logic without prescribing its content. Funding follows identification: states must reserve 7 percent of their Title I Part A allocation for school improvement activities, directed to districts serving identified schools. The plans themselves must be developed with stakeholder participation, must be based on a school-level needs assessment, must include evidence-based interventions meeting the statutory tiers, and must identify resource inequities for comprehensive support schools. Each of those requirements is a federal constraint on a state-designed process, and together they define the later law’s theory of action: the federal government guarantees that struggling schools are named, that their plans meet quality standards, and that resources follow the naming, while the states and districts decide what the plans contain.
What the evidence does and does not establish
A verdict that rests on evidence must be candid about the evidence’s limits, and the research on the earlier law has limits worth stating plainly. No researcher can run a controlled experiment on a national statute: there is no America without the 2002 law against which to compare the America with it, so every finding about the law’s effects rests on comparisons across states, across time, or across levels of accountability pressure, each with its own assumptions. The strongest studies exploit the fact that the law’s pressure bit differently in different places: states that had no prior accountability system experienced a sharper change than states that already tested and reported, and schools near the proficiency thresholds faced different incentives than schools far above or below them. Dee and Jacob’s 2011 analysis used National Assessment data and cross-state variation in prior accountability to isolate the law’s contribution, which is why its findings carry the weight they do: significant fourth-grade mathematics gains reaching an effect size of 0.22 by 2007, improvements at both the lower and top percentiles of the distribution, eighth-grade mathematics gains concentrated among low-achieving groups, and no consistent gains in reading. The asymmetry between mathematics and reading is itself informative: mathematics instruction responds more readily to pressure on tested content, while reading achievement, entangled with vocabulary and background knowledge accumulated over years, moves more slowly.
The mechanism studies fill in how the pressure traveled from statute to classroom. Dee, Jacob, and Schwartz documented in 2013 that the law’s accountability pressure was associated with roughly 600 dollars in additional per-pupil spending, financed by state and local sources rather than federal aid, alongside higher teacher compensation, a larger share of elementary teachers holding advanced degrees, no detectable change in class sizes, and a reallocation of instructional time away from science and social studies toward tested reading and mathematics. Reback, Rockoff, and Schwartz reported in 2014 that pressure reduced teachers’ sense of job security and shifted instructional time toward specialist teachers in high-stakes subjects, while finding positive to neutral effects on students’ enjoyment of learning and on achievement. Read together, the mechanism studies describe a system that moved real resources, money, personnel, and time, in response to measured pressure, and that the movements were concentrated where the measurement pointed. That is the documented basis for the verdict’s claim that the earlier regime forced uncomfortable information into public view and produced visible reactions: the reactions are in the spending data, the staffing data, and the schedule data, not merely in the test scores.
What the evidence does not establish is equally important for an honest verdict. It does not establish that the later law’s design works better, because the later law’s core bet, that state-designed, evidence-based interventions outperform federally prescribed ones, had not been tested at national scale in the research available through 2014. The evidence tiers are a plausible mechanism for quality control, but plausibility is not proof, and the history of the earlier law’s formalistic compliance counsels skepticism about any system’s ability to guarantee quality from Washington. The evidence also does not establish that the earlier law’s test-score gains persisted, compounded, or translated into later-life outcomes at the scale the law’s architects hoped; the documented gains are real but narrow, concentrated in early-grade mathematics, and the reading results are a standing caution against overclaiming. Finally, the evidence does not resolve the value question at the heart of the deciding factor: whether comparability across states or local fit of interventions matters more is a judgment about what the education system is for, not a finding any study can deliver. The verdict names the tradeoff; the reader supplies the values.
A note on reading the statutory text
Readers who want to verify this comparison against the laws themselves will find the task easier with a map of where the key provisions live. Both statutes are organized as amendments to the Elementary and Secondary Education Act, so the section numbers that matter are the Act’s section numbers as amended, not the public law’s. Title I, Part A, the compensatory education program for disadvantaged students, is the home of the accountability provisions in both regimes. Section 1111, the state plan section, carries the testing requirements, the reporting requirements, the goals, the identification rules, and the prohibitions on the Secretary in the 2015 law; it is the single most important section in the comparison. Section 1116 carried the earlier law’s intervention ladder. Section 1003 governs school improvement funding, including the 7 percent reservation. Section 8101 collects definitions, including the four evidence tiers. Sections 8526A, 1604, and 8401 carry the prohibitions and the fenced waiver authority.
A reauthorization is amendatory language: it strikes some provisions, inserts others, and redesignates the survivors, which means the 2015 law is best read as a set of instructions for editing the 2002 law rather than as a freestanding document. The practical consequence is that provisions the later law did not touch, such as the testing schedule, appear in the amended text without any marker of their continuity; only a comparison of the before and after reveals what survived. That is why the requirement-by-requirement table in this article is organized by function rather than by section number: the functional organization shows the reader what each regime did about testing, reporting, goals, identification, intervention, and discretion, regardless of where the drafters happened to place the language. A reader who works through the table and then checks two or three rows against the statutory text will have learned the most transferable skill this article can teach: how to read a reauthorization as a record of decisions about who decides.
The verdict: a named deciding factor
A comparison in this series is finished only when the reader can say which design they would choose and why, so this section renders the verdict the brief requires and names the deciding factor explicitly. The deciding factor is this: the earlier regime was better at forcing uncomfortable information into public view and worse at producing responses that helped, and the later regime reversed both strengths and weaknesses. Under the earlier law, the combination of universal annual testing, disaggregated subgroup reporting, a single federal deadline, and a federally triggered consequence ladder made it very difficult for any school system to hide poor performance by its most vulnerable students. The measurement was public, the targets were common, and the consequences were automatic. Under the later law, the measurement and the reporting remain, but the targets are bespoke, the identification thresholds are federally fixed yet the interventions are locally designed, and the entire apparatus depends on the quality of state and district execution in a way the earlier apparatus did not. The earlier model surfaces inequity; the later model, at its best, responds to it more usefully. The reader who values comparability across states should prefer elements of the earlier model, and the reader who values locally designed intervention should prefer the later one.
The evidence for the first half of the verdict, the earlier law’s strength at surfacing information, is the record of what became visible that had been invisible before. Before the disaggregation requirements, a school could report a respectable average while its Black students, its Latino students, its students with disabilities, and its English learners failed at high rates, and no public document would reveal the gap. The earlier law made that concealment unlawful. Dee, Jacob, and Schwartz, in their 2013 study published in Educational Evaluation and Policy Analysis, documented that the law’s accountability pressure raised per-pupil spending by roughly 600 dollars, funded by state and local sources, increased teacher compensation and the share of elementary teachers holding advanced degrees, left class sizes unchanged, and reallocated instructional time away from science and social studies toward tested reading. Those are resource and practice effects, not test score effects, and they show a system responding to measurement pressure by moving money, personnel, and time. Reback, Rockoff, and Schwartz, in their 2014 paper “Under Pressure” published in the American Economic Journal: Economic Policy, found that accountability pressure reduced teachers’ perceived job security and shifted instructional time toward specialist teachers in high-stakes subjects, while showing positive to neutral effects on students’ enjoyment of learning and on achievement. The two studies together describe a regime that forced systems to react visibly: spending moved, staffing moved, schedules moved. Whether the reactions were wise is a separate question from whether they were visible, and on visibility the earlier regime’s record is strong.
The evidence for the second half of the verdict, the earlier law’s weakness at producing helpful responses, is the record of what the prescribed interventions actually did. The take-up figures are the starkest part of that record: federal evaluations reported that about 1 percent of eligible students exercised the school choice option and about 17 percent received supplemental educational services. The choice remedy failed largely because the receiving schools were often full, distant, or unwilling, and because the families most in need of the option were least positioned to use it; the tutoring remedy failed to scale because the supply of approved providers was uneven and the logistics of after-school and summer services defeated participation. The higher rungs of the ladder fared no better in practice: corrective action menus were implemented formalistically in many districts, and restructuring frequently changed governance labels without changing classroom practice. Dee and Jacob’s 2011 findings, significant fourth-grade mathematics gains with an effect size of 0.22 by 2007 and eighth-grade mathematics gains concentrated among low-achieving groups but no consistent reading gains, suggest that the pressure produced real but narrow improvements, concentrated in the most heavily tested subject and grade, rather than the broad transformation the ladder promised. A regime that is excellent at diagnosing and poor at treating has a specific institutional shape, and the earlier law is the clearest American example of it.
The later law’s design can be read as a point-by-point response to those findings, and the verdict must credit the response while noting what it sacrificed. The evidence tiers answer the finding that prescribed interventions lacked research support: no district may spend school improvement funds on an intervention without a research basis meeting one of the four tiers. The state-designed interventions answer the finding that federal prescription produced formalism: districts choose responses fitted to their circumstances, from staffing changes to schedule changes to curriculum adoption, provided the evidence standard is met. The retained identification thresholds answer the danger that local design becomes local evasion: no state may decline to name its lowest-performing 5 percent of Title I schools or its high schools with graduation rates at or below 67 percent, so the uncomfortable information still enters public view by federal command even though the response is locally designed. What the later law sacrificed is comparability and automaticity. Without a common deadline, the public cannot rank states by ambition; without automatic triggers, the identification of a school begins a process of plan-writing and monitoring whose quality varies with the capacity of the state and the district. The later regime’s bet is that local knowledge plus evidence standards will outperform federal prescription, and the honest statement of the verdict is that the bet is plausible but unproven at the level of national evidence, while the earlier regime’s information effects are documented in the research record.
This article takes no position on standardized testing as a practice. The judgments above concern institutional design: which arrangement of measurement, targets, consequences, and executive discretion is more likely to surface inequity and which is more likely to produce useful responses. A reader who opposes standardized testing and a reader who supports it can each accept the verdict as stated, because the verdict compares two regimes that both require the same testing and differ only in what follows it. Every performance claim in this section is attributed to named research, and the article makes no reference to any state plan or dispute. Readers who want the full evidence base that drove Congress to replace the earlier law, including the achievement, resource, and practice findings summarized here, can consult the No Child Left Behind outcomes companion, which assembles the research record in one place.
Why deregulation is the wrong word
The most persistent misdescription of the 2015 rewrite is the framing of the change as deregulation, and the misdescription matters because it leads readers to expect something the statute does not deliver: a smaller federal rulebook. The federal rulebook did not shrink; it changed authors and, in several places, grew more prescriptive. The identification requirements are the clearest example. The earlier law triggered interventions from missed targets that states themselves had set, which meant a state could soften the trigger by softening its trajectory. The later law names the categories and the thresholds in federal text: the lowest-performing 5 percent of Title I schools, high schools with four-year graduation rates at or below 67 percent, subgroups performing at the level of the lowest 5 percent. No state discretion operates on those numbers. A state that wishes to avoid identifying schools cannot do so by redefining the categories, because the categories are defined in Washington. In this respect the later law is more prescriptive than its predecessor, not less.
The evidence standards are a second example. The earlier law prescribed interventions without asking whether they worked; the later law permits any intervention but requires a research basis, with four defined tiers from experimental studies down to a rationale with ongoing evaluation. A district official who lived under both regimes would recognize the difference immediately: under the earlier law, the federal government told the district what to do and asked no questions about evidence; under the later law, districts choose their own responses but must defend each choice with evidence. It is a different kind of regulation, one that constrains the quality of the district’s reasoning rather than the content of its decision. Describing that shift as deregulation confuses the removal of prescription with the removal of rules, and the statute contains too many rules for the confusion to survive a reading.
The reporting obligations are a third example, already documented in the measurement section: the later law added per-pupil expenditure reporting and educator qualification reporting to the disaggregated assessment reporting it retained. The prohibitions on the Secretary are a fourth, and the most telling. A deregulating Congress does not write overlapping, redundant prohibitions on its own executive: Section 1111(e)(1)(B)(ii) on standards, Section 1111(e)(1)(B)(iii) on evaluation systems, Section 1111(j) on Common Core coercion, Section 8526A on instructional content and curricula act-wide, Section 1604 for Title I, and fenced waiver authority under Section 8401. That is not the drafting of legislators who want less federal law; it is the drafting of legislators who want federal law to do a different thing, namely to fence the executive rather than to direct the schools. The accurate description of the 2015 rewrite is reallocation of authority rather than reduction of federal requirements: authority moved from the Secretary to Congress’s own text, from Washington’s prescription to state and district design, while federal requirements around measurement, identification, evidence, and reporting remained and in several cases tightened.
Is the later law simply deregulation?
No. Federal identification requirements, evidence standards, and reporting obligations all remain, and several run more prescriptively than their predecessors. The 2015 law names the school categories states must identify, sets four evidence tiers for interventions, and adds public reporting items. Authority moved from Washington to state capitals, but the federal rulebook did not shrink; it changed authors.
Where to study next
A comparison resolved with a verdict is only useful if the reader knows what to read next, and this article carries the study-path recommendation for the series. The recommended order is deliberate. Begin with the framework both statutes amend, then study the earlier regime in full, then the evidence that drove its replacement, then the later regime, and finally the series study guide that places all five in a learning sequence. The earlier regime’s complete provisions are presented in the No Child Left Behind guide; the later regime’s complete architecture is presented in the Every Student Succeeds guide; the research record on what the earlier law did to achievement, resources, and practice is assembled in the outcomes companion; and the underlying 1965 framework is explained in the Elementary and Secondary Education Act guide. Each of those companions is linked from the section where its subject is treated, so a reader who has followed the four axes in order has already collected the full set.
For readers deciding what to study first across the whole legislation series, the US legislation study guide organizes the series into a recommended sequence, with this comparison positioned after the two statute guides and the outcomes companion. The study guide is the right next step for a student who has finished this article and wants to know which comparison to read second, which statute repays close reading of its text, and how the education titles relate to the series’ other domestic policy subjects.
Two companion tools support the study path: the VaultBook companion and the ReportMedic companion. Between the five linked guides and the two companions, a reader has everything needed to reproduce the One Test from memory: which requirements survived, which did not, why the shift concerned consequences rather than measurement, and which model the evidence favors for surfacing inequity versus producing useful responses.
The closing thought belongs to the deciding factor, because it is the sentence the reader should retain when the details fade. The earlier regime forced uncomfortable information into public view with a bluntness no prior federal education law had achieved, and it paired that bluntness with prescribed responses that the research record shows were often formalistic, underused, or narrow in their effects. The later regime kept the information machinery, fenced the executive that had improvised around the earlier law’s failures, and bet that states and districts, required to identify their struggling schools and to ground their interventions in evidence, would design better responses than Washington could prescribe. Same measurement, different response: the two statutes measure schools almost identically and differ almost entirely in who decides what happens next, so debates framed as testing versus no testing are describing a difference that does not exist in the statutes.
Frequently Asked Questions
Q: What is the difference between No Child Left Behind and ESSA?
The difference lies almost entirely in who decides what happens after test scores arrive, not in the testing itself. Both laws require annual reading and mathematics assessments in grades 3 through 8 and at least once in high school, and both require results disaggregated by race, ethnicity, economic disadvantage, disability status, and English learner status. Where they differ: No Child Left Behind set a single federal deadline of 100 percent proficiency by the end of the 2013-2014 school year and prescribed a federal ladder of interventions, from school choice through restructuring, triggered by missed targets. ESSA requires states to set their own long-term goals with interim progress measures, mandates identification of defined school categories such as the lowest-performing 5 percent of Title I schools, leaves intervention design to states and districts subject to four evidence tiers, and expressly bars the Secretary from conditioning approval on particular standards, assessments, or evaluation systems.
Q: Did ESSA repeal No Child Left Behind?
In substance, yes, though the legal mechanism was amendment rather than standalone repeal. Both Public Law 107-110 and Public Law 114-95 are reauthorizations of the same underlying statute, Public Law 89-10, the Elementary and Secondary Education Act of 1965. ESSA, enacted December 10, 2015, replaced the accountability title of the 2002 law with a new framework: it discarded the 100 percent proficiency deadline, the adequate yearly progress calculations, and the federally prescribed intervention ladder, while retaining the testing schedule, the disaggregated reporting, and the 95 percent participation rule. Because the 2015 act amended the same framework rather than striking a separate law from the books, lawyers describe it as a reauthorization that superseded its predecessor’s accountability provisions. For practical purposes, the No Child Left Behind regime ended when the new provisions took effect.
Q: Which gives states more control, No Child Left Behind or ESSA?
ESSA gives states substantially more control over targets and interventions, while keeping federal control over measurement and identification. Under No Child Left Behind, Congress set the endpoint, 100 percent proficiency by the end of the 2013-2014 school year, and Section 1116 prescribed the interventions in federal text: choice, tutoring, corrective action, restructuring. Under ESSA, states set their own long-term goals with interim measures, and states and districts design the improvement plans for identified schools. But the control is not unlimited: ESSA fixes the identification categories in federal law, including the lowest-performing 5 percent of Title I schools and high schools at or below 67 percent graduation rates, requires evidence-based interventions meeting statutory tiers, and adds reporting categories. The accurate answer is that ESSA moved the design of consequences to the states while retaining federal guardrails around who gets identified and what counts as evidence.
Q: What did ESSA keep from No Child Left Behind?
ESSA kept the entire measurement apparatus and several of the earlier law’s structural features. The annual testing schedule survived identically: reading and mathematics in grades 3 through 8 and at least once in grades 9 through 12, under Section 1111(b)(2)(B)(v)(I). Disaggregated subgroup reporting survived and expanded, with added categories for per-pupil expenditures and educator qualifications. The 95 percent participation requirement survived in Section 1111(c)(4)(E), applying to all students and each subgroup. The requirement that states participate in the National Assessment of Educational Progress in grades 4 and 8 survived. What ESSA discarded was the response machinery: the 100 percent proficiency deadline, adequate yearly progress, the Section 1116 intervention ladder, and the Secretary’s broad discretion to condition waivers on favored policies. The pattern is consistent: measurement stayed, consequences changed authors.
Q: Is accountability weaker under ESSA than No Child Left Behind?
It is different in kind rather than weaker in any simple sense, and the answer depends on which element of accountability one means. On measurement and identification, ESSA is as strong or stronger: testing is identical, reporting expanded, and the identification thresholds, the lowest 5 percent of Title I schools and the 67 percent graduation rate line, are fixed in federal text with no state discretion. On consequences, ESSA is less prescriptive by design: it requires evidence-based plans designed by states and districts rather than imposing a federal ladder. Whether that produces weaker results depends on state capacity and on whether the evidence tiers genuinely discipline local choices, questions the research record had not settled by the close of 2014. The framing of ESSA as deregulation misleads here: federal requirements around identification, evidence, and reporting remain, and several are more prescriptive than their predecessors. Reallocation of authority is the accurate description.
Q: Why did Congress replace No Child Left Behind with ESSA?
Congress acted because the earlier law’s central mechanism had become unworkable and its workarounds had become constitutionally uncomfortable. The 100 percent proficiency deadline of the end of the 2013-2014 school year was unreachable, and as it approached, the number of schools failing adequate yearly progress grew so large that the intervention ladder lost its power to distinguish. The prescribed remedies showed poor results in practice: federal evaluations found roughly 1 percent of eligible students used school choice and roughly 17 percent received tutoring, while restructuring often changed governance labels without changing practice. Meanwhile the waiver era, from the September 2011 announcement through approvals reaching 43 states plus the District of Columbia and Puerto Rico, showed the executive conditioning relief on favored standards and evaluation systems, concentrating in the Secretary a policy-making power that belonged to Congress. The 2015 act, passed 359 to 64 in the House and 85 to 12 in the Senate and signed December 10, 2015, wrote state-designed goals, state-designed interventions, and express prohibitions on the Secretary into the statute itself.
Q: Did waivers from No Child Left Behind lead to ESSA?
The waivers were a major cause, though not the only one. Beginning with the September 23, 2011 announcement and the first approvals on February 9, 2012 to ten states, the executive branch offered relief from the 2002 law’s requirements, including the 2013-2014 deadline, to states adopting favored policies on standards, assessments, and teacher and principal evaluation. Eventually 43 states plus the District of Columbia and Puerto Rico operated under waivers, which meant the law as written had effectively been replaced by an executive-designed regime. The mechanism alarmed legislators of both parties even when they supported some of the underlying policies, because it demonstrated that broad waiver authority plus an unworkable deadline let the Secretary rewrite accountability without Congress. ESSA’s sharpest provisions, the overlapping prohibitions in Sections 1111(e), 1111(j), 8526A, 1604, and the fenced waiver authority of Section 8401, were written with that history in view. The waivers proved the old design could not function as written; the new design was Congress’s answer.
Q: Which should a student study first, No Child Left Behind or ESSA?
Study No Child Left Behind first. The 2002 law created the vocabulary the entire field still uses: adequate yearly progress, disaggregated subgroups, the 95 percent participation rule, and the idea that federal money carries measurable conditions. Its rise and breakdown explain why every provision of the 2015 law looks the way it does: the state-designed goals answer the impossible 100 percent deadline, the evidence tiers answer the prescribed but unevaluated interventions, and the prohibitions on the Secretary answer the waiver era. A student who starts with ESSA encounters provisions that seem arbitrary, such as the 5 percent identification threshold or the four evidence tiers, without understanding the failures they were written to correct. The recommended sequence is the 1965 framework, then the 2002 regime in full, then the research on its effects, then the 2015 rewrite, with the series study guide organizing the order. Comparison requires a baseline, and the earlier law is the baseline.
Q: Did No Child Left Behind require testing in every grade?
No. The statute required annual reading and mathematics assessments in grades three through eight, with that grade-span system in place by the end of the 2005-2006 school year, plus at least one assessment within the grades ten through twelve band. Grades outside those spans were not subject to the federal annual testing mandate. The common impression of testing in every grade likely comes from the law’s salience rather than its text: because the tested grades drove accountability verdicts and dominated school life, the mandate felt universal even where it was not. ESSA preserved the same grade structure under section 1111(b)(2)(B)(v)(I), requiring annual assessments in grades three through eight and at least once in grades nine through twelve. The continuity is exact. Anyone comparing the two statutes on testing should therefore focus on the identical grade spans rather than on any imagined expansion or contraction, since neither occurred.
Q: What was adequate yearly progress, and does it survive under ESSA?
Adequate yearly progress was the annual verdict at the heart of the 2001 accountability system. Each state set a starting point and annual measurable objectives rising toward one hundred percent proficiency by the end of the 2013-2014 school year, and every school, district and state either made adequate yearly progress or did not, subgroup by subgroup, each year. Two consecutive years of missing it triggered improvement status and the section 1116 consequence ladder. The verdict was public, binary and inescapable, which made it the earlier law’s most powerful disclosure mechanism and its most anxiety-producing feature for educators. It does not survive under ESSA. The 2015 act replaced it with state-set long-term goals measured against interim progress, plus categorical identification of schools for comprehensive, targeted or additional targeted support. The annual pass-or-fail judgment for every school in every subgroup ended with the earlier regime, and with it ended the single national timetable that had organized the country’s education politics for fourteen years.
Q: What happened to a school that missed its targets under No Child Left Behind?
A Title I school that missed adequate yearly progress for two consecutive years entered improvement status under section 1116, which began a federally prescribed sequence. The district had to write a school improvement plan, provide technical assistance, and offer pupils the right to transfer to a higher-performing public school with transportation funded from a Title I set-aside. A third missed year added supplemental educational services, free tutoring, with twenty percent of the district’s Title I, Part A allocation reserved for choice transportation and tutoring combined. A fourth year brought corrective action, including staff replacement, new curricula or outside experts. A fifth year required restructuring under governance options such as charter conversion, replacing most staff, private management or state takeover. In practice the early remedies saw limited use, with federal evaluations recording choice take-up near one percent of eligible pupils and tutoring near seventeen percent, and the later rungs were implemented with wide variation across districts.
Q: What are comprehensive support and improvement schools under the 2015 law?
Comprehensive support and improvement is the most intensive identification category under ESSA, defined by Section 1111(c)(4)(D) and Section 1111(d). It covers three groups: the lowest-performing 5 percent of Title I schools in the state, all public high schools with four-year graduation rates at or below 67 percent, and schools previously identified for additional targeted support that fail to improve within the state’s timeframe, not to exceed four years. Identification is mandatory and the thresholds are fixed in federal text; no state may set the bar elsewhere. Once identified, the district develops an improvement plan with stakeholder participation using evidence-based interventions, the state monitors implementation, and schools that miss the state’s exit criteria face more rigorous state-determined action. The category concentrates the law’s strongest requirements on the schools with the weakest results, replacing the old ladder’s automatic triggers with a cycle of identification, planned intervention, monitoring, and escalation.
Q: What are targeted support and improvement schools, and what does additional targeted support mean?
Targeted support and improvement applies where one or more student subgroups consistently underperform within an otherwise unidentified school, with the state defining consistent underperformance. The school develops its own improvement plan addressing the struggling subgroups, using evidence-based interventions. Additional targeted support and improvement is the sharper subcategory: it applies where a subgroup on its own performs at or below the level of the lowest-performing 5 percent of Title I schools in the state. When that line is crossed, the school becomes subject to the comprehensive support requirements for that subgroup’s needs, including the district-level planning, state monitoring, and exit criteria. The two-tier design extends the earlier law’s insight, that schoolwide averages can hide subgroup failure, into the intervention system itself: a school that looks acceptable overall but fails a specific group of students cannot escape identification by pointing to its average.
Q: What limits did the 2015 law place on the Secretary of Education?
The 2015 law fenced the Secretary with overlapping prohibitions written in direct response to the 2011 to 2015 waiver era. Section 1111(e)(1)(B)(ii) bars conditioning plan or waiver approval on adoption of specific academic standards. Section 1111(e)(1)(B)(iii), subclauses (IX) and (X), bars conditioning approval on particular teacher, principal, or school leader evaluation systems or effectiveness measures. Section 1111(j) bars influencing, incentivizing, or coercing adoption of the Common Core or membership in any assessment consortium. Section 8526A imposes an act-wide bar on mandating, directing, or controlling instructional content, standards, assessments, curricula, or programs, expressly including through waiver conditions. Section 1604 adds a Title I-scoped parallel bar. Section 8401 retains waiver authority but excludes allocations, civil rights, and other categories from its reach. Together these provisions make the 2015 law the rare federal statute that distrusts its own administrator in explicit text.
Q: What evidence standards do interventions have to meet under the 2015 law?
Section 8101(21)(A) defines four tiers. Strong evidence means support from at least one well-designed and well-implemented experimental study. Moderate evidence means support from at least one well-designed and well-implemented quasi-experimental study. Promising evidence means support from at least one well-designed and well-implemented correlational study with statistical controls for selection bias. The fourth tier covers interventions with a rationale based on high-quality research findings or positive evaluation suggesting likely improvement, paired with ongoing efforts to examine effects. School improvement plans must use interventions meeting one of these tiers, with the stronger tiers required for certain funding uses. The system constrains the quality of a district’s reasoning rather than the content of its decision: any intervention is permitted, from schedule changes to curriculum adoption, provided the district can cite research of the required rigor. The earlier law prescribed interventions without any evidence requirement at all.
Q: How did the 95 percent participation rule work under each law?
Under No Child Left Behind, testing at least 95 percent of students in each subgroup was part of adequate yearly progress itself: a school that tested too few students failed to make progress regardless of scores, which closed the most obvious gaming strategy, excluding likely low scorers from the testing pool. ESSA preserved the rule in Section 1111(c)(4)(E), requiring states to factor 95 percent participation by all students and by each subgroup into the statewide accountability system, with the statute specifying how non-participation affects a school’s rating. The continuity is deliberate and important. A testing requirement without a participation requirement invites selection of test-takers, and both Congresses understood that measured performance is meaningless if schools choose who gets measured. The rule is one of the clearest examples of a requirement that survived the rewrite intact in both letter and function.
Q: Did research find that No Child Left Behind raised student achievement?
The findings are mixed and subject-specific, which is why this article attributes each claim to its study. Thomas Dee and Brian Jacob, in the Journal of Policy Analysis and Management in 2011, found significant fourth-grade mathematics gains reaching an effect size of 0.22 standard deviations by 2007, with improvements at the lower and top percentiles, and eighth-grade mathematics gains concentrated among low-achieving groups. In reading they found no consistent gains. Dee, Jacob and Nicole Schwartz, in Educational Evaluation and Policy Analysis in 2013, found the law raised per-pupil spending by roughly six hundred dollars financed by state and local sources, raised teacher compensation and the share of elementary teachers with advanced degrees, had no class-size effects, and reallocated instructional time from science and social studies toward tested reading. Randall Reback, Jonah Rockoff and Schwartz, in the American Economic Journal: Economic Policy in 2014, found accountability pressure lowered teachers’ perceived job security while having positive to neutral effects on pupils’ enjoyment of learning and achievement. Mathematics improved; reading did not.
Q: Did school choice and tutoring under the earlier law reach many students?
No. The two remedies that touched students directly had strikingly low take-up, and the figures are among the best-documented failures of the earlier regime. Federal evaluations reported that roughly 1 percent of eligible students used the school choice option, which allowed transfer to a higher-performing school with transportation provided. Roughly 17 percent of eligible students received supplemental educational services, the federally funded tutoring for students from low-income families in schools missing targets for a third year. Choice failed largely on logistics: receiving schools were often full or distant, and the families most in need were least positioned to navigate the transfer process. Tutoring failed to scale because approved provider supply was uneven and after-school and summer scheduling defeated participation. The gap between the statute’s promise and families’ experience demonstrated that a remedy’s existence in federal text guarantees nothing about its delivery, a lesson the 2015 law’s monitoring requirements were written to address.
Q: Why did so many schools miss the 100 percent proficiency deadline?
The deadline was unreachable by design and evadable by redefinition, a combination that guaranteed mass failure. One hundred percent proficiency for every subgroup by the end of the 2013-2014 school year required sustained annual gains no school system had ever achieved, so the number of schools missing adequate yearly progress rose year after year as the date approached, not because schools were collapsing but because the trajectory was impossible. At the same time, each state defined proficiency on its own assessments, creating a direct incentive to set lenient cut scores: a demanding standard meant watching schools fail in growing numbers, while a lenient one produced apparent progress toward the same federal goal. The National Center for Education Statistics mapping studies documented the result, with states’ own results showing more favorable trends than the National Assessment in most comparisons, 12 of 14 in NCES 2010-456 and 8 of 10 in NCES 2011-458. The deadline simultaneously demanded the impossible and rewarded redefinition, which is why it collapsed as a policy instrument.
Q: Did subgroup reporting weaken under ESSA?
No. The disaggregated reporting requirements continued in full under the 2015 act. Results are reported separately for each major racial and ethnic group, for economically disadvantaged pupils, for children with disabilities, and for English learners, at the school, district and state levels, exactly as under the earlier statute. The ninety-five percent participation requirement applies to each subgroup under section 1111(c)(4)(E). What changed was not the reporting but what the reported results triggered. Under the earlier law, subgroup results fed directly into the annual adequate yearly progress verdict, so a struggling subgroup meant a failing verdict for the whole school. Under the later law, subgroup results feed into the identification categories: schools with consistently underperforming subgroups are identified for targeted support, and schools where a subgroup performs at or below the lowest five percent level receive additional targeted support, escalating to comprehensive support if unimproved. The visibility of subgroup performance survived; the annual binary consequence attached to it did not.