The Future of Evidence in Education Network
Necessary but Not Sufficient: Six Design Principles for a Stronger Education R&D System
Six functions a public education R&D system must sustain if evidence is to remain credible, cumulative, and useful across decisions, settings, and time.
- Authors
- Future of Evidence in Education
Convening Members - Report editors
- Kelly Hallberg
Erin Higgins - Series editor
- Vivian C. Wong
- Published
- 2026
Authors and Convening Members
These reports were developed through the Future of Evidence in Education convening held in December 2025. All convening members participated in the full convening and contributed to the ideas developed across the two reports. Members then worked in smaller groups to develop each report. All convening members are authors of both reports.
Convening members
- Rekha BaluUrban Institute
- Beth BoulayMathematica
- Brooks BowdenUniversity of Pennsylvania
- Ben DomingueStanford University
- Dan GoldhaberAmerican Institutes for Research and University of Washington
- Kelly HallbergUniversity of Chicago
- Erin HigginsAlign R&D
- Cara JacksonCenter for Outcomes Based Contracting at the Southern Education Foundation
- Luke MiratrixHarvard University
- Robert OlsenGeorge Washington Institute of Public Policy
- Laura PeckRutgers University
- Jessaca SpybrookUniversity of South Carolina
- Elizabeth TiptonNorthwestern University
- Betsy WolfDistrict of Columbia Office of the State Superintendent of Education
- Brian WrightUniversity of Virginia
- Vivian C. WongUniversity of Virginia
Report teams
This report
Necessary but Not Sufficient: Six Design Principles for a Stronger Education R&D System
- Beth Boulay
- Rekha Balu
- Ben Domingue
- Brooks Bowden
- Dan Goldhaber
- Rob Olsen
- Brian Wright
Report editors: Kelly Hallberg and Erin Higgins Series editor: Vivian C. Wong
Companion report
Evaluating Causal Evidence in Education: A Practical Guide
- Beth Boulay
- Laura Peck
- Jessaca Spybrook
- Elizabeth Tipton
- Betsy Wolf
Report editors: Cara Jackson and Luke Miratrix Series editor: Vivian C. Wong
Suggested citation
Future of Evidence in Education Convening Members. (2026). Necessary but not sufficient: Six design principles for a stronger education R&D system. Future of Evidence in Education Network. [Insert URL]
Acknowledgments
The December 2025 Future of Evidence in Education convening that led to this report was supported by a grant from the Spencer Foundation (Grant #202600090). The views expressed in this report are those of the authors and do not necessarily reflect the views of the Spencer Foundation.
Download this report
How this Report was Developed
In December 2025, we convened a cross-sector group of researchers, evaluators, and evidence leaders for a two-day discussion about the education research and development (R&D) system. Participants brought experience designing and analyzing studies, developing and applying evidence standards, evaluating programs, examining how findings are relevant across settings, estimating costs and benefits, and helping decision-makers use evidence in policy, procurement, and practice.
The meeting sought to develop a shared diagnosis of the challenges confronting education R&D as the federal role in supporting research became more uncertain. We began with decisions that state and local leaders often face, considered what evidence those decisions require, and worked backward to identify the conditions needed for that evidence to remain credible, interpretable, and useful.
As the discussion progressed, it became clear that this work called for two related reports. This report presents six design principles for a stronger education R&D system. It addresses the institutions, incentives, partnerships, documentation practices, synthesis efforts, and sustained investments that shape what evidence is produced, how it is interpreted, and how it can inform decisions across settings and over time. Our companion report presents ten principles for the responsible interpretation and use of evidence. It focuses on how leaders can assess what conclusions the evidence supports, identify relevant questions and limitations, and account for uncertainty when deciding whether to adopt, pilot, or scale a program.
The principles in this report grew out of the practical challenges that the convening members have encountered in building, studying, and sustaining public education R&D. They reflect lessons from work across the education R&D system since the Institute of Education Sciences was established in 2002, including the progress made during that period, the vulnerabilities that have limited its value, and the capacities still needed. Together, the principles offer an account of the functions that must work in concert if the education R&D system is to produce knowledge that remains credible and useful across decisions, settings, and time.
Introduction
State and local education leaders make consequential decisions about programs, practices, products, and policies on tight timelines, often with incomplete information. To make those decisions well, leaders require credible, independent evidence that helps them judge what is likely to work, for whom, relative to what alternative, and under what conditions.
But access to individual studies or summaries of their findings is not enough. Leaders also need an education R&D system that turns study findings into actionable insights. In the United States, responsibility for that work is distributed across many public and private institutions. These institutions’ choices shape the type of evidence that is produced, how evidence is interpreted, and whether it can inform decisions over time. Within this broader system, the Institute of Education Sciences (IES), the research, statistics, and evaluation arm of the Department of Education, has played a central role in funding education research, producing national statistics, conducting evaluations, and reviewing evidence.
The demands on this system have grown as its federal support has weakened. Digital products and AI-enabled tools are entering classrooms faster than conventional evaluation cycles can generate evidence about their effects, and descriptions of their capabilities often exceed what is known about whether they improve student outcomes (Gray & Lewis, 2021; Diliberti et al., 2024, 2025). Districts now maintain access to an average of more than 3,000 digital tools (LearnPlatform by Instructure, 2026; Merod, 2023). At the same time, the federal infrastructure for supporting education research is becoming more fragile. In early 2025, IES was reduced to a fraction of its former staff, and many contracts that supported the agency’s research, data, technical assistance, and evidence review functions were canceled or disrupted. Some disrupted contracts have since been reinstated, and new competitions for regional research capacity and high priority areas are underway, but the agency’s ability to carry out its core functions and its long-term future remain uncertain (Binkley & Vázquez Toness, 2025; Department of Education, 2025; Meyer et al., 2026; Jackson et al., 2025; SAM.gov, 2025).
This report was developed amid current debates about how federal education research should be rebuilt, reorganized, and sustained following recent disruptions at IES. Recreating the system developed over the past two decades would leave important weaknesses in place, while shifting responsibility to philanthropy, states, districts, and other local organizations would leave essential functions without durable support. A recent report, Reimagining the Institute of Education Sciences: A Strategy for Relevance and Renewal, proposes changes intended to connect IES more closely to state and local priorities, make it more responsive to practitioners’ needs, and deliver findings faster (Northern & Opp, 2026). We share these priorities, but also argue that greater responsiveness and speed will not be enough.
We believe that education R&D serves a public interest that extends beyond any single study, organization, or local decision, thereby reinforcing the importance of the federal role in supporting the education R&D system. In the remainder of this report, we argue that the effectiveness and value of the education R&D system depends on a set of functions that a public R&D system must sustain over time. This report is intended for researchers, funders, intermediaries, developers, policy analysts, and others whose decisions shape the production, interpretation, and use of education evidence.
What IES Helped Build and What the Field Still Needs
The creation of IES in 2002 marked a federal commitment to independent, public evidence infrastructure for education (Cook & Foray, 2007). Over two decades, IES raised expectations for rigor in causal research and supported shared standards for judging evidence quality (Whitehurst, 2018; What Works Clearinghouse, 2022). It invested in research-practice partnerships, maintained national data collections, and supported state longitudinal systems that make it possible to track student outcomes, describe the state of education, and conduct policy-relevant research at scale (NCES, n.d.). IES also invested in synthesis and translation through the What Works Clearinghouse (WWC) and its practice guides, which were intended to make evidence more accessible for decision-making.
These investments addressed a specific weakness in the field. Before IES, education research had few shared expectations for when a study could support a causal conclusion, and little incentive or capacity to conduct rigorous impact studies in schools and districts (Cook, 2003). Through competitive grants, training programs, research centers, and investment in methods development, IES changed the norms for causal evaluation and built the field’s capacity to produce credible evidence (NASEM, 2022; Wong, Tipton, & Spybrook, 2025).
These investments remain essential, but the needs of an education R&D system extend beyond the rigor of individual studies. The field has long recognized that leaders need evidence that applies beyond the settings where studies were conducted and keeps pace with emerging programs and products. They also need to understand what implementation requires, which conditions shape results, and how to apply findings when local constraints differ. Yet these broader needs have not been consistently prioritized in how research is designed and funded. High-quality studies considered one at a time cannot meet those needs. The system must build knowledge around shared problems and apply what is learned across studies.
IES has taken important steps to strengthen the review and synthesis of education research. Through the WWC and its grants programs, IES established common standards for reviewing causal studies and created public resources for finding and assessing research, including individual study reviews, systematic reviews, and practice guides. These resources made research findings and judgements about study quality more transparent. Their usefulness for local decisions, however, has been constrained by gaps in the underlying studies, including missing information about implementation, comparison conditions, cost and context, as well as the limits on the research eligible for WWC review. More fundamentally, review and synthesis occur after studies have been designed or completed, and therefore cannot ensure that separate studies examine shared questions or document the information needed for comparison and synthesis.
IES grantmaking illustrates the balance the system must maintain. IES requests for applications established topic areas and methodological expectations, and its grant programs supported implementation research, cost analysis, research-practice partnerships, and other work connecting evidence to practice. Within those boundaries, its research grants programs supported a broad range of field-initiated research, leaving research teams substantial discretion over the specific questions they pursued and the information they documented. That discretion supported scientific independence and field-initiated inquiry, but it did not ensure that separate studies addressed shared questions or documented the information needed for later comparison and synthesis. More targeted R&D centers and research networks created greater coordination within selected priorities, but they did not consistently provide mechanisms for planning and synthesizing bodies of research across the system. Even with these investments, the evidence base left important questions unaddressed and often provided too little information about comparison conditions, implementation, cost, context, product versions, and variation across settings that leaders need for interpreting findings. The system needs to preserve field-initiated inquiry but with an eye toward synthesis across individual studies and needs to continue to strategically use coordinated investments when shared questions, planned variation, and cumulative evidence are needed.
When studies are not designed and documented with later synthesis in mind, gaps in information carry into clearinghouse rating, approved-product lists, and procurement guidance (Goldhaber et al., 2025; Higgins & Doolittle, 2025). Leaders may see a rating or label without enough information to judge whether the evidence applies to their own setting. Different intermediaries may also apply different standards and review criteria, allowing the same body of evidence to point toward different conclusions (Wadhwa, Zheng, & Cook, 2024). Leaders are then left to reconcile those differences without the information needed to do so (Tseng & Coburn, 2019).
Finally, a separate problem arises when consequential decisions must be made before relevant evidence exists. Leaders select and modify policies in response to budget cycles, political pressures, and local needs, while edtech companies may make products available before evidence exists to support their use or plans to generate this evidence are in place. Schools and districts may not pair these choices with plans to document implementation and outcomes or with evaluation designs that provide credible comparisons. As a result, leaders may have little evidence when making a choice, and the field learns little afterward about what was implemented or whether the selected policy or product improved outcomes relative to an alternative.
Six Design Principles for a Stronger Education R&D System
The principles that follow identify six functions that need to be maintained or strengthened if education R&D is to support better decisions, responsible innovation, and improved learning across settings over time. They are not a redesign plan for any single institution. Reimagining IES offers a detailed agenda for renewing IES, including changes to its priorities, infrastructure, and relationships with state and local education systems (Northern & Opp, 2026). This report asks which functions must the education R&D system as a whole sustain and how should responsibility for them be distributed across institutions.
Federal leadership is essential, but no federal agency can perform these functions alone. Agencies, funders, researchers, intermediaries, developers, and education leaders carry out different parts of the work under different mandates and incentives. Their combined efforts have produced important gains, but they have not ensured that all six functions are coordinated or sustained. Shared infrastructure that serves many institutions and persists across funding cycles requires public investment because no single institution has sufficient incentive or responsibility to sustain it. Taken together, the principles identify where a distributed system requires coordination and where essential work may be neglected without clear responsibility and ongoing support.
Figure 1 provides an overview of the six design principles and shows the role each plays in a stronger R&D system. Principles 1, 2, and 3 concern the production and accumulation of evidence. Individual studies produce rigorous evidence suited to the questions leaders face (Principle 1), and shared expectations for how studies are reported and reviewed determine whether those findings can be interpreted and compared across studies (Principle 2). Synthesis draws on the resulting body of work, while prospective planning should determine which studies are needed and what they document so that the evidence base represents meaningful variation in populations, settings, and implementation conditions (Principle 3). Together, these three functions form a cycle in which reporting enables synthesis and what is learned through synthesis informs future studies. Principles 4 and 5 concern the conditions under which this work can be carried out, including research capacity and durable partnerships that connect evidence-building to the decisions education systems face (Principle 4), and public investment in the data systems, research networks, and shared infrastructure that individual grants cannot sustain, along with support for transformative innovation (Principle 5). Principle 6 focuses on the performance of the system itself, whether public investments are producing knowledge that is credible, cumulative, responsive, and useful, and how the system should change when it falls short.
The figure highlights that weakness in any one of these functions limits what the others can produce. Rigorous studies, for example, cannot compensate for study reports that omit information leaders need to interpret findings, and none of these functions can be sustained without the capacity and public investment that support them.
1 Rigorous evidence
Anchor the system in rigorous, independent, decision-relevant evidence.
The education evidence system needs sustained public investment in independent, rigorous research across the full range of questions leaders need to answer. IES established clear norms and standards for causal evaluation in education, including randomized trials, strong quasi-experimental designs, and shared criteria for judging when a study can support conclusions about impact. Those standards were an important achievement that should be protected and built on. At the same time, causal impact evidence alone will not provide answers to all of the questions leaders face. Leaders need evidence about implementation, cost, feasibility, variation across settings, and the conditions that shape results. The field needs clear standards for judging the rigor of each form of evidence on its own terms, along with safeguards that protect scientific judgment from political or commercial pressure.
What is at stake
IES did more than fund individual impact studies; it raised the bar for what qualifies as credible evidence of cause and effect and institutionalized stronger expectations for research design, review, and reporting. Without credible causal evidence, leaders have fewer ways to distinguish evidence of impact from marketing, anecdotes, intuition, or engagement data when making decisions about programs, policies, and practices. Our companion report builds on this foundation by explaining how leaders can judge what conclusions different forms of evidence support and how that evidence should inform decisions.
The next phase of education R&D should preserve those gains while broadening what the field means by rigor. In education research, the work of establishing standards for causal impact studies has often made rigor feel most closely associated with randomized trials and other impact designs. However, that definition is too narrow. A rigorous study should pose clear questions and use sound measures and methods suited to the conclusions it is intended to support. Impact studies should be judged by standards for causal inference. Implementation research, qualitative studies, cost analyses, and mixed methods designs should be judged by standards appropriate to their purposes. They should not be treated as informal supplements to impact evidence, but as a direct part of the evidence infrastructure.
Credible evidence also depends on independence from political and commercial interests that may prefer particular questions, methods, or results. Public officials and agencies have a legitimate role in setting broad research priorities. However, scientific merit review, study design, analysis, interpretation, and reporting also require safeguards against interference. Without those protections, changing political priorities or sponsor interests can distort research findings and limit what leaders are able to learn from them.
Recommendations
- Federal agencies and philanthropic funders should sustain and expand support for rigorous causal research, including randomized trials and strong quasi-experimental designs. Causal work should remain a core public evidence function.
- Federal agencies, funders, and professional associations should support a full range of rigorous study types, including implementation research, qualitative studies, mixed methods designs, and cost and cost-effectiveness analyses alongside causal studies (Bowden, 2023). Funding programs should support these forms of evidence when they are needed to understand an intervention’s impact, what implementation was required, and what it cost. Existing guidance and standards on implementation research, cost analysis, and generalizability provide a foundation, but federal agencies and professional associations should extend this work to forms of research for which standards are not widely shared and update the guidance as new methods and evidence needs emerge.
- Federal agencies and research organizations should prioritize rigorous research designed to inform consequential decisions that recur across districts and states, such as selecting curricula, structuring staffing models, and designing integrated student supports.
Research questions and methods should reflect the information leaders need to make those decisions.
- Federal agencies should establish protections for scientific merit review and for the conduct and reporting of publicly funded research. Subject-matter and methods experts should lead proposal review, and grants and contracts should preserve researchers’ authority to carry out planned analyses and report findings that conflict with policy priorities or sponsor interests.
Examples
IES’s tiered grant structure illustrates how a public R&D system can invest across different kinds of research questions. Exploration, development, efficacy, effectiveness, and measurement grants were designed to support different stages of evidence generation rather than treat all studies as if they served the same purpose (NASEM, 2022). The IES and NSF Common Guidelines for Education Research and Development provide another starting point by distinguishing among types of research and the purposes they serve, including foundational, exploratory, design and development, efficacy, effectiveness, and scale-up research (IES & NSF, 2013). Recent IES guidance on implementation research and cost analysis also recognizes that impact findings are not sufficient on their own (Hill et al., 2023; IES, n.d.). Together, these efforts provide a foundation for developing clearer expectations for rigor across forms of education research, including qualitative and mixed methods studies.
2 Transparent reporting
Set transparent expectations for how evidence is reported, reviewed, and communicated to leaders.
A stronger public education R&D system must do more than produce rigorous studies. Study reports, clearinghouse reviews, evidence tiers, product certifications, and summaries should provide the information leaders need to understand how conclusions were reached and what conclusions the evidence does and does not support. Our companion report offers principles that leaders can use when interpreting evidence for specific decisions. This principle focuses on the reporting, review, and disclosure requirements that determine whether findings can be interpreted by leaders, classified in evidence registries, and compared and synthesized across studies.
What is at stake
A well-designed study can be difficult to use when its report omits information that shapes how the finding should be interpreted. Reports may leave unclear the comparison condition, the version of the product studied, the supports required for implementation, the outcomes and measures examined, or the costs of using the program in practice. A curriculum evaluation may report a positive effect on reading fluency without disclosing that the comparison condition was business-as-usual with no structured reading instruction, that the program required 90 minutes of daily literacy time, or that the study used a researcher-developed assessment overly aligned with the program (Wolf & Harbatkin, 2023). Without this information, a district leader cannot judge whether the finding is relevant to the choice their district faces.
Evidence review organizations make choices about which studies and findings to include, how to classify them, and what information to carry into their summaries. The WWC provides a model for making review standards and protocols public. However, published criteria alone do not ensure that review products preserve all the information leaders need for interpreting results. Depending on their purpose, reviews may exclude some study designs or give limited attention to implementation, comparison conditions, subgroup findings, and other information relevant to a decision. Review products should make their scope and exclusions clear, even when those exclusions are appropriate to the purpose of the review.
Evidence tiers, product certifications, and approved-product lists can help leaders navigate an uneven evidence base, but these designations do not represent the same kind or strength of evidence. Under ESSA’s evidence tiers, Tier 4 means that a program “demonstrates a rationale.” It requires a well-specified logic model grounded in research and an effort to study the program, but no prior evidence that the program improved student outcomes. A product certification may concern how a product was designed or whether it possesses a particular feature rather than whether it has improved student outcomes. When leaders cannot see what criteria a rating uses or what the rating was intended to show, they may assume that ratings from different organizations represent comparable evidence even when they assess different things. A vendor may present a favorable rating as evidence that its product works without explaining whether the rating reflects a logic model, a feature of the product, or evidence of effects on student outcomes (Goldhaber et al., 2025; Higgins & Doolittle, 2025).
Products also change over time. A district considering a tutoring program may be shown evidence from a study of an earlier version of the product that has since been revised. Without clear documentation of the version studied and the changes made since the study, a leader cannot judge whether the prior findings apply to the product currently under consideration.
Recommendations
- Funders and journal editors should require study reports to document the population and setting studied, the intervention and comparison conditions, the outcomes and measures examined, the supports provided during implementation, and the version of the product studied. When cost data are collected, reports should describe the methods and assumptions used (Cost Analysis Standards Project, 2021). Missing information needed to interpret the findings should be grounds for revision.
- Evidence review organizations should make their review criteria, protocols, and scope public. Review products should identify the study designs, outcomes, populations, and forms of evidence eligible for review, explain how the criteria were applied, and distinguish evidence that fell outside the review’s scope from an absence of research.
- Organizations that issue evidence ratings, tiers, or product certifications should state what each designation assesses, what evidence it is based on, and what conclusions it does and does not support. This information should appear alongside the designation in review products and evidence registries rather than only in separate technical documentation.
- Federal agencies, state education agencies, and large district procurement offices should require vendors seeking placement on approved or recommended product lists to identify the version studied, provide the evidence supporting any statement about effectiveness,
and describe substantive changes made since the study. Evidence registries and review organizations should date their reviews and identify the product version to which each review or rating applies.
- Intermediaries that translate and organize research for practitioners and leaders, including curriculum review organizations, evidence registries, and rapid evidence platforms, should preserve the study information needed for interpretation and synthesis. Summaries should identify the population and setting studied, the intervention and comparison conditions, the outcomes and measures examined, the implementation supports provided, and the limits of the conclusions the evidence can support. When the underlying study does not report this information, the summary should make the absence clear.
Examples
The BIRD-E Blueprint and the IES SEER standards illustrate efforts to improve the information available from individual studies. BIRD-E proposes a common data language for describing education research, while SEER encourages researchers to preregister studies, make findings and methods open, document implementation and comparison conditions, analyze costs, use high-quality outcome measures, and support generalization (BIRD-E, n.d.; IES, n.d.). These frameworks show how shared reporting expectations can make findings easier to interpret and synthesize. The WWC provides a model for making review standards and protocols public. Public standards are one part of transparency. Leaders also need to understand the scope of the review and retain access to the information about populations, comparison conditions, implementation, and subgroup findings underlying the rating. Digital Promise offers certifications that assess distinct features, including research-based design, practitioner-informed design, and evidence meeting ESSA Tier 3. Making the purpose of each certification visible helps prevent leaders from treating different designations as equivalent evidence of effectiveness. Finally, the Prevention Services Clearinghouse offers a useful model for addressing changes to programs over time. It identifies the version reviewed and uses a manual-comparison process to determine whether findings from one version can be applied to another (Jackson, Wilson, & Glenn, 2024). These examples show that transparency requires both public review criteria and preservation of the information needed to understand what evidence was considered, what a designation represents, and the conclusions that are supported.
3 Planning and synthesis
Plan and synthesize evidence so the field can learn from variation.
The education R&D system needs stronger ways to build knowledge across studies, populations, and settings. This begins before studies are launched, with research designs that capture meaningful variation across populations, settings, measures, and implementation conditions. It also requires synthesis that can be updated as evidence develops and reach leaders in time to inform the decisions they face. Over time, synthesis should report average effects and examine credible patterns of variation across studies when the evidence permits. These patterns can generate hypotheses about why results differ and help leaders judge what prior evidence suggests for their own settings.
What is at stake
The central question leaders face when using evidence is not whether a program worked somewhere, but whether it is likely to work in their own setting, with their own students, and under their own constraints. Answering that question requires an evidence base that captures variation across settings and helps explain when, where, and under what conditions results are likely to hold. Evidence from one setting can inform decisions in another; otherwise, education research would not be contributing to cumulative knowledge. The question we should ask is not whether the settings that produce a finding are identical, but whether differences between the settings are likely to change results. A stronger evidence base should represent enough variation and document settings well enough to examine which conditions shape results and whether findings are likely to hold elsewhere.
The current R&D system does not systematically vary key study and program features in the way that would be needed to generate this kind of evidence. Studies are planned one at a time, with limited attention to whether their samples, measures, comparison conditions, and implementation documentation will support learning across studies later. Studies are rarely conducted with samples representative of the populations to which findings are meant to apply (Tipton & Olsen, 2022). Sites also self-select into evaluations in ways that make typical study samples systematically different from the broader range of settings where leaders are making decisions (Olsen et al., 2013; Stuart et al., 2017). Schools in small districts, rural areas, and towns are underrepresented in the evidence base relative to large schools in large urban districts (Tipton et al., 2021). As a result, leaders in many settings must draw conclusions from studies conducted in contexts that may not resemble their own. When studies capture too little variation across settings or provide too little information about that variation, later reviews have limited ability to explain where results are likely to hold, where they may differ, and why.
These limitations constrain what can be learned through synthesis. The WWC and related systematic review efforts have helped build a broader knowledge base that has informed policy and practice debates. Meta-analyses can examine variation across studies, but their ability to do so depends on the range of settings, populations, measures, and implementation conditions represented in the evidence base and the information studies report about them. The outcome measure can also affect how large a program’s estimated effect appears. Studies using measures closely aligned with an intervention tend to report larger effects than studies using broader or independent measures (Kraft, 2020; Wolf & Harbatkin, 2023). A synthesis that reports an overall average without showing how the evidence varies may give leaders an incomplete picture of what to expect in their setting.
Cumulative learning also requires synthesis that extends beyond named programs. When evidence is organized mainly around whether one curriculum, product, or intervention outperformed another, the field may learn too little about the recurring features associated with differences in results. Core-components analyses offer one approach. Researchers can document which features are present in each program and examine whether programs that include a feature have different estimated effects from programs that do not. These patterns can identify components that warrant direct testing, but they cannot establish that the component caused the difference because the programs may differ in other ways (McCormick, Wilson, & Dymnicki, 2024). Research on holistic supports for community college students, for example, demonstrates how the field can learn not only whether a particular program improved completion, but whether recurring features such as intensive advising, financial support, structured course-taking, and connections to tutoring or career services are associated with results across settings or generate hypotheses for future research (Weiss, Bloom, & Singh, 2023). Without synthesis across these features of programs, evidence remains tied to individual products and studies rather than contributing to a more general knowledge base.
Even well-designed synthesis may arrive too late for the decisions leaders face. A district considering whether to renew or purchase a new math platform may need to act before credible synthesis about that product is available. The Annenberg Institute’s EdResearch for Action offers one model for making synthesis more responsive to these decision needs. It organizes research syntheses around problems identified by state and district leaders and involves a practitioner advisory board of school and district leaders in developing briefs, with the goal of producing guidance that is clearer, more feasible, and better aligned with the conditions under which leaders make decisions. Leaders may still need to rely on usage data, early assessment indicators, vendor materials, or evidence from related products when synthesis is unavailable. They often cannot wait for a complete evidence base, but the R&D system does too little to assemble and interpret the best available evidence on the timeline when decisions must be made. ESSA Tier 4 recognizes that leaders sometimes select programs without prior impact evidence by requiring research-based rationale and an effort to study the program’s effects. A study conducted as the program is implemented can produce evidence for future decisions, but only if it is designed to support credible conclusions and its findings are reported. Better synthesis can improve the evidence available for an immediate decision, while prospective evaluation can build knowledge from choices made under uncertainty.
Recommendations
- Federal agencies and philanthropic funders should support prospectively planned bodies of studies around high-priority questions, with shared measures, common documentation of implementation and context, and samples chosen to capture meaningful variation across settings. These studies should be designed from the start to support later synthesis, generalizability analysis, and learning about what works, for whom, and under what conditions.
- Federal agencies and evidence review organizations should examine and report variation in effects when the evidence base supports credible analysis, including variation across sites, populations, implementation conditions, resources, and measures. Reviews should describe the populations and settings represented in the evidence base and state when the available studies are too narrow to assess generalizability. Methods for assessing generalizability, including approaches summarized by Tipton and Olsen (2022), should be incorporated when the evidence base supports them.
- Federal agencies and evidence review organizations should synthesize evidence across common program features, implementation supports, and design choices, not only across named interventions. Reviews should help the field identify patterns associated with differences in results and generate hypotheses for future research.
- Federal and state agencies should provide incentives, funding, and technical support for evaluations when districts select programs without prior impact evidence. For interventions used under ESSA Tier 4, agencies should support plans to study effects on student outcomes, and ensure that the findings contribute to the public evidence base.
- Federal agencies should work in consultation with education leaders to commission and maintain living evidence reviews on high-priority topics, including early literacy, tutoring, math instruction, and student support models. These reviews should be updated as new studies become available and, where decisions are time-sensitive, produced in forms that can inform procurement, budget, and adoption cycles.
- AI tools may help translate evidence summaries for practitioners and leaders, but they should not be used as a primary mechanism for research synthesis until they have been evaluated for whether they can extract study characteristics, assess study quality, identify effect heterogeneity, and communicate uncertainty with enough accuracy to support public evidence review.
Examples
The Reading for Understanding Initiative illustrates what sustained investment in cumulative learning can produce. It supported six research teams and generated a body of work large enough to support a synthesis of more than 200 scholarly articles on reading comprehension (Pearson et al., 2020; IES, n.d.). The Special Education Research Accelerator (SERA) offers a related model for planned variation. SERA was designed as a platform for large-scale, multi-site replication studies with diverse samples, and its second phase focuses on generating evidence about the generalizability of intervention effects (IES, n.d.). The EdResearch for Action and the Education Endowment Foundation illustrate models for organizing and presenting research synthesis around questions relevant to education leaders.
4 Capacity and partnerships
Build education-specific capacity for research and evidence use and sustain durable partnerships.
A public education R&D system depends on researchers who understand education settings, public agencies and local systems with the capacity to interpret and use research, and durable partnerships that connect evidence-building to the decisions schools, districts, states, and communities face. These forms of capacity develop through training, experience carrying out research in education settings, and relationships sustained over time.
What is at stake
Education research requires more than methodological expertise. Researchers need a deep understanding of how schools and districts work, how programs are chosen and changed in practice, how outcomes are measured, and what constraints leaders face. IES-funded training programs helped prepare a generation of researchers with strong causal methods and knowledge of education settings (NASEM, 2022; Wong et al., 2025). Disruptions to IES in 2025 put that pipeline at risk, and rebuilding it will take years.
Formal training alone does not prepare researchers to carry out complex studies in schools and districts. IES grant programs also gave researchers experience addressing the practical demands of education research, including gaining access to sites, maintaining partnerships, collecting data, and studying programs under field conditions. Sustaining research capacity therefore requires both training programs and opportunities to conduct rigorous studies.
Expertise also needs to be available where decisions are made. Most trained researchers work in universities and research organizations, while many state agencies and districts have limited capacity to assess and use evidence. A large district with a research office can review studies, manage partnerships, and assess evidence provided by vendors, but a small rural district may have to rely on vendor materials and state guidance. When federal research infrastructure weakens, districts and states with limited research staff have fewer sources of independent support and they become more dependent on summaries, ratings, and intermediary guidance. These resources, however, may not provide enough information for the decisions leaders face (NASEM, 2022).
Research-practice partnerships (RPPs), regional education labs (RELs), local policy labs, and cross-state networks have developed in response to concerns that education research has been produced at arm’s length from practice. By giving researchers and agencies time to understand each other’s work and develop questions grounded in local needs, these partnerships can produce findings that are easier to interpret and use. Yet, such relationships can be difficult to sustain. Part of the reason is that funding, where it exists, is seldom structured to sustain a partnership beyond a single project cycle, a point Principle 5 takes up. Partnerships may also depend on particular individuals, leaving accumulated knowledge vulnerable to staff and leadership turnover. Without routines that preserve institutional knowledge, each new project may have to rebuild relationships that prior work established.
Durable partnerships provide a structure through which researchers, education leaders, teachers, families, and communities can shape research priorities together. Research priorities have often been set by researchers, funders, and federal agencies, with limited input from teachers, parents, and communities. Recent proposals to redesign IES call for a system that is more responsive to educators, families, and policy leaders, but that cannot mean asking for input after the agenda has already been set. When research questions do not reflect the problems practitioners and communities face, the resulting evidence is less likely to be relevant, trusted, or used (Jackson, 2022; Northern & Opp, 2026).
Recommendations
- Federal agencies should rebuild and expand the education research pipeline through training, professional development programs, and grant programs that give researchers experience conducting rigorous studies in education settings. Developing education-specific expertise takes years, and interruptions in training and research opportunities will weaken the field’s capacity long after funding resumes (NASEM, 2022).
- Federal agencies and philanthropic funders should support career pathways into state agencies, district research offices, curriculum organizations, professional development
providers, and education technology firms. Fellowships, loan forgiveness, and other incentives could make these organizations viable destinations for trained researchers and evaluators. Preparation for these roles should combine analytic skills with knowledge of education settings and decision processes.
- Funders should give RPPs enough time to develop shared agendas, trust, and institutional knowledge. RPPs, REL-type partnerships, and cross-state collaboratives need longer awards and renewal options rather than grant structures that require each project to begin again. Five-year grants would provide a more realistic starting point for this work. Awards should also support routines, documentation, and staff development that preserve accumulated knowledge within institutions through staff and leadership turnover.
- Research priorities should be shaped with the people affected by the decisions being studied. Partnerships supported by federal agencies and philanthropic funders should give teachers, students, parents, community organizations, and education leaders a meaningful role in identifying research questions before priorities are set.
- State agencies and districts need staff who can assess research, manage partnerships, and connect leaders with relevant evidence. Federal agencies and philanthropic funders should treat this capacity as infrastructure rather than overhead and should support roles that can be sustained beyond a single project.
- Investments in capacity should reach settings that lack established research offices and partnerships. Otherwise, new resources may flow to the systems best positioned to apply for them while leaving smaller and under-resourced agencies dependent on vendor materials or outside guidance.
- Research partnerships should include developers and vendors when their participation is needed to understand or evaluate products in use. These arrangements should disclose financial interests, give researchers access to the information needed for independent analysis, and protect their ability to report unfavorable findings.
Examples
IES undergraduate, predoctoral, and postdoctoral programs illustrate the kind of sustained investment needed to build education-specific research capacity. Between 2004 and 2020, IES established predoctoral training programs at 21 universities and invested nearly $209 million in that training infrastructure (IES, n.d.). Restoring and expanding those programs, with stronger pathways into state agencies, district research offices, and other education organizations, would help preserve expertise that took decades to develop. IES research grants complemented formal training by giving researchers opportunities to design and carry out rigorous studies in schools and districts. Through these grants, researchers applied their methodological preparation and gained experience conducting studies in education settings.
RPPs, REL collaborations, local policy labs, and cross-state collaboratives illustrate the partnership side of the same principle. Their value is not only that they produce studies, but that they build trust, shared routines, and institutional knowledge over time. The CORE districts collaborative, AERDF, and IES-funded networks such as SEERNet offer examples of coordinated priority-setting at scale. Community-based participatory research projects that center teacher and family priorities in question development, including the Family Leadership Design and the Adapted Measure of Math Engagement, offer another model for making agenda-setting more responsive to the people the evidence system is meant to serve.
5 Public investment
Invest in education R&D at a scale that supports research, public infrastructure, and transformative innovation.
The education R&D system cannot do the work described in the design principles without sustained investment at a much larger scale. Rigorous studies, transparent reporting, durable partnerships, planned synthesis, and system-level feedback together depend on infrastructure that cannot be built through individual grants and short-term projects alone. Education needs public R&D investments large enough to support ambitious questions, shared data and research infrastructure, coordinated studies across settings, and innovation that is developed and evaluated before it spreads widely through schools.
What is at stake
The funding problem concerns both the level and form of federal investment. Competitive grants are essential for supporting individual studies and new lines of inquiry. Some resources, however, must serve many studies and remain available beyond a single award. Because their benefits extend across institutions and funding cycles, no district, research team, firm, or foundation has responsibility for sustaining them for the field.
Federal funding for infrastructure and education research remains limited. In FY 2024, the principal federal agency dedicated to education research and statistics, IES, received $793.1 million, compared with about $47.1 billion for the National Institutes of Health (AERA, 2024; National Institutes of Health, n.d.). For scale, public elementary and secondary schools spent roughly $927 billion in 2020-21, reported in constant 2022-23 dollars (NCES, n.d.). The IES budget is small relative to NIH funding and total public elementary and secondary school expenditures.1
Existing public data resources are one form of infrastructure that requires sustained investment. Through the National Center for Education Statistics (NCES), IES maintains national data collections and supports state longitudinal data systems. NCES data allow researchers and policymakers to describe the condition of education and examine changes in students, schools, and outcomes over time. State longitudinal data systems allow states to connect records across years and education sectors for research and decision-making. The value of these resources depends on continuity, comparability of measures, and secure access to data over a sustained period of time. Interruptions in data collection can create gaps that later funding may not be able to reconstruct.
The growth of administrative and platform data creates a second infrastructure need. States, districts, instructional platforms, and vendors hold data that could help answer questions no
NSF, HHS, and other federal agencies fund research relevant to education. Their investments are not included because their funds are distributed across programs with broader missions and cannot be separated into a comparable education R&D total from published agency budgets.
single source can address. Research use of these data requires secure access, privacy protections, common definitions, and methods for linking records across systems when appropriate. Research networks also need shared measures and agreements that allow studies to build on one another across settings and over time. Individual grants can establish these arrangements for one study, but they are not designed to maintain them for future work. Stronger evidence about implementation, cost, variation, and generalizability depends on sustained investment in shared research infrastructure.
The education R&D system also needs infrastructure for developing and testing new approaches. New products, including AI-enabled tools and digital learning platforms, may reach classrooms before conventional evaluations have produced evidence about how effective they are across students and settings, what implementation requires, and what they cost. Public investment can support shared experimentation platforms and development pipelines that connect early design, testing, improvement, and evaluation before new approaches become widespread. It can also support evaluation of tools already used in schools. These investments would allow evidence to inform the development and use of new products before entering schools.
Recommendations
- Congress and federal agencies should provide sustained funding for education R&D at a level that reflects the size and importance of the education system. Funding should support a robust portfolio of competitive research grants, maintain shared public infrastructure, and sustain coordinated programs of research that extend across studies and settings.
- Federal funding should support infrastructure that no single project, district, state, developer, or philanthropy can maintain. Priorities include longitudinal data systems, secure data access, common measures, coordinated research networks, shared protocols, rapid-cycle testing, and the capacity to study implementation, cost, and differences across settings.
- Federal agencies should create funding streams for large education R&D projects that connect development, testing, improvement, and evaluation from the beginning. These projects should address problems that require coordinated public investment, such as AI-enabled instructional tools, curriculum improvement, tutoring and student supports, assessment systems, and policy approaches for under-resourced settings.
Examples
NCES data collections and the Statewide Longitudinal Data Systems Grant Program demonstrate the forms of public infrastructure that individual research projects cannot create. NCES longitudinal studies provide data that researchers can use to examine education experiences and outcomes over time, while the Statewide Longitudinal Data Systems grants support state capacity to link and use administrative data across education systems. Maintaining these data collections and systems requires sustained public investment beyond the life of any single research project.
IES’s Accelerate, Transform, Scale Initiative illustrates one model for innovation-oriented R&D. Established in response to congressional direction in FY 2023, the initiative was designed to support advanced education R&D intended to produce scalable solutions and was inspired by advanced research projects agencies in other parts of the federal government (IES, n.d.). Digital learning environments, such as OpenStax Kinetic and UpGrade, show how evidence generation can be built into the infrastructure where learning occurs. OpenStax Kinetic was designed as research infrastructure for large-scale studies in authentic digital learning environments, with the capacity to support experimental studies and protect data privacy. UpGrade is an open-source platform for running randomized A/B tests in educational software (Basu-Mallick, Bradford, & Baraniuk, 2023; Ritter et al., 2020). These efforts offer models for connecting development and evaluation through shared research infrastructure.
6 System-level assessment
Measure whether the education R&D system is serving its public purpose.
Education R&D should be held to the same discipline the field asks of the programs and policies it studies. The education R&D system needs mechanisms for assessing whether public investments are producing credible, cumulative, and relevant knowledge, and whether that knowledge is reaching leaders who need it, in forms they can use. Reviewing the performance of the R&D system can help public agencies and funders identify gaps and make better decisions about future investments.
What is at stake
Public investment in education research serves a purpose beyond producing individual studies. It supports independent inquiry into consequential education questions, helps public institutions make better use of scarce resources, and builds knowledge that no single state, district, research organization, or firm has the incentive or capacity to produce on its own. Improving education and outcomes for students is the larger aim, but changes in student outcomes are not a direct measure of the R&D system’s performance. Over the past two decades, the field has made progress in setting standards for individual studies, especially studies that estimate program impacts. While these standards continue to be essential, an R&D system can produce many credible studies and still leave leaders without the evidence they need.
Judging the performance of the system requires asking a different set of questions. Does publicly supported research meet appropriate standards for rigor, independence, and transparency? Are investments addressing consequential choices faced by schools, districts, states, and communities? Do studies build on one another, or do they remain scattered across topics, settings, and products? Are findings and syntheses reaching leaders when decisions are being made, and can leaders find, interpret, and use them? What role does evidence play in policy, procurement, implementation, and resource allocation? These questions extend beyond the quality of any one study to ask whether the research portfolio as a whole is producing knowledge the field can use. Articulating these questions does more than enable assessment; it defines what a public R&D system is ultimately for and gives the field common goals.
Common indicators, such as counts of grants funded, studies completed, programs reviewed, or positive findings, do not answer those questions. They show whether research is being funded, completed, and assessed against evidence standards, but not whether the resulting knowledge accumulates across studies, responds to important decisions, or informs practice. A portfolio can contain many rigorous studies while leaving major questions underexamined. State administrators choosing a program to improve students’ writing, for example, can learn how many studies of writing programs were funded and how many met evidence standards, but not whether the findings answer the implementation questions they are facing, such as what training teachers need, how much instructional time is required, and what costs will be incurred.
Recommendations
- Federal agencies and philanthropic funders should review their research portfolios on a regular basis. These reviews should examine whether funded work meets appropriate standards for rigor and transparency, addresses high-priority decisions, includes populations and settings that have received too little attention, and contributes to a body of knowledge rather than a collection of separate grants. The findings should clarify the state of the evidence and identify priorities for future investment.
- Federal agencies and evidence review organizations should assess whether WWC reviews, practice guides, rapid evidence reviews, and other synthesis products reach state, district, and school leaders and contribute to procurement, adoption, budget, and policy decisions. The assessment should examine whether leaders know these products exist, can access and interpret them, and receive them when relevant decisions are being made. Feedback from leaders and practitioners should identify whether the products address the questions they face and what additional support they need to use them.
- Public agencies and philanthropic funders need a shared set of indicators for judging whether the evidence system is becoming more credible, cumulative, responsive, and useful. Those indicators should distinguish measures of activity, such as studies funded, reviews completed, and standards applied, from measures of performance, such as whether evidence addresses consequential decisions, builds on prior findings, and reaches leaders in forms they can use. They should also assess whether investments in state and local research staff and partnerships have strengthened leaders’ ability to find, interpret, and use relevant evidence.
- Federal agencies and philanthropic funders should examine whether publicly funded state and local innovations are producing evidence that others can interpret, test, or build on. Pilots supported with public funds should document what was tried and under what conditions, what it cost, what was changed, and what others can learn from the effort.
- Research organizations, intermediaries, and professional associations should review the tools made available to leaders for interpreting evidence, including evidence summaries, clearinghouse ratings, and approved-product lists. These reviews should assess whether the tools are accurate, complete, usable and clear about the limits of evidence, as well as whether and how leaders use them and which decisions they inform. The findings should be made public.
Examples
IES’s experience with federal performance measurement illustrates both the value and limits of system-level assessment. Under the Office of Management and Budget’s Program Assessment Rating Tool (PART), a process used in the 2000s to rate federal programs, IES received an “effective” rating in 2007 for its research, development, and dissemination programs. IES continued to track performance after OMB’s PART tool was discontinued around 2009, using performance measures that it reported annually, which went beyond counts of grants and completed studies to track whether findings from funded research had been reviewed by the WWC, met WWC standards, and showed positive effects on student outcomes (IES, n.d.). These measures ask whether funded research is producing credible, documented effects, but they capture only one part of the system’s public purpose. They do not answer whether the research portfolio is addressing the most consequential decisions in education, building knowledge across studies, or reaching leaders in forms they can use.
A future education R&D system would need a broader approach to portfolio review. NIH’s Office of Portfolio Analysis offers one model, supporting agency-wide decisions through tools, methods, and guidance for examining research investments. A similar process in education, led by a public agency or independent body, could assess whether publicly funded studies address consequential decisions, include underexamined populations and settings, and contribute to knowledge that accumulates over time. The 2022 NASEM review of IES provides a precedent for this kind of independent, system-level assessment and suggests how such reviews could be conducted on a recurring basis.
The system also needs ongoing feedback from the people expected to use the evidence. Federal agencies could use recurring convenings with research staff, partnership leaders, and other education leaders, along with documentation from federally supported partnerships, to examine whether leaders know what evidence is available, whether it reaches them when decisions are being made, and how it informs policy and practice. These structures could also identify gaps in the evidence and barriers to its use.
Conclusion
Because the future of federal education R&D is uncertain, the field needs a clearer vision for how the system should be rebuilt when the opportunity comes. This report is meant to support that conversation. Rigorous studies are necessary, but they are not sufficient. Their value depends on a system that preserves what is learned, builds knowledge across studies and connects evidence to decisions that education leaders face. The six design principles offer a starting point for speaking with a more unified voice about the public functions an education R&D system needs to perform.
These principles do not imply that all research should be organized through centralized agendas or coordinated initiatives. Field-initiated research creates space for new questions, methods, and ideas that agencies and funders may not anticipate. Coordinated studies serve a different purpose by examining shared questions across settings, testing for meaningful variation in effects, and filling in gaps that separate projects cannot address alone. A strong education R&D system needs both types of work.
The same balance applies to funding. Competitive grants support individual studies and new lines of inquiry. Public infrastructure that serves many studies and persists across funding cycles requires a different form of investment. Infrastructure cannot depend on individual projects, but investment in infrastructure cannot replace support for new research.
With limited resources and existing institutional barriers, these commitments will create tradeoffs. Stronger documentation requirements add costs for researchers, while faster reviews may need to cover less evidence or answer narrower questions. The concern is not that tradeoffs will be made, but that they may emerge through separate decisions without a shared understanding of their consequences or of what the system is trying to achieve. These principles offer a common basis for making those choices. While they cannot resolve every tension, they can help the field judge whether its decisions preserve room for independent inquiry while building the coordination and infrastructure needed for evidence to accumulate and reach education leaders.
References
- American Educational Research Association. (2024, March). Final FY 2024 appropriations include significant cut for NSF, small reductions for IES and NIH under spending caps. https://www.aera.net/Newsroom/AERA-Highlights-E-newsletter/AERA-Highlights-March-2024/Final-FY-2024-Appropriations-Include-Significant-Cut-for-NSF-Small-Reductions-for-IES-and-NIH-Under-Spending-Caps
- Basu-Mallick, D., Bradford, B. C., & Baraniuk, R. (2023). Secure education and learning research at scale with OpenStax Kinetic. In Proceedings of the Tenth ACM Conference on Learning @ Scale (L@S ’23). https://doi.org/10.1145/3573051.3596187
- Binkley, C., & Vázquez Toness, B. (2025, February 12). DOGE cuts $900 million from agency that tracks American students’ academic progress. Associated Press. https://apnews.com/article/ies-musk-doge-education-cuts-4461d7bdbe9d55c5a411d8465999b011
- BIRD-E. (n.d.). About the Blueprint. https://www.bird-e.org/about-the-blueprint
- Bowden, A. B. (2023). Designing Field Experiments to Integrate Research on Costs. AERA Open, 9.
- Cook, T. D. (2003). Why have educational evaluators chosen not to do randomized experiments? The Annals of the American Academy of Political and Social Science, 589(1), 114–149. https://doi.org/10.1177/0002716203254764
- Cook, T. D., & Foray, D. (2007). Building the capacity to experiment in schools: A case study of the Institute of Educational Sciences in the U.S. Department of Education. Economics of Innovation and New Technology, 16(5), 385–402. https://doi.org/10.1080/10438590600982475
- Cost Analysis Standards Project. (2021). Standards for the economic evaluation of educational and social programs. American Institutes for Research. Available online: https://www.air.org/project/cost-analysis-standards-project-casp
- Digital Promise. (n.d.). Research-Based Design: ESSA Tier 4. https://digitalpromise.org/product-certifications/research-based-design-essa-tier-4/
- Diliberti, M. K., Lake, R. J., & Weiner, S. R. (2025). More districts are training teachers on artificial intelligence: Findings from the American School District Panel. RAND Corporation. https://doi.org/10.7249/RRA956-31
- Diliberti, M. K., Schwartz, H. L., Doan, S., Shapiro, A., Rainey, L. R., & Lake, R. J. (2024). Using artificial intelligence tools in K–12 classrooms. RAND Corporation. https://doi.org/10.7249/RRA956-21
- EdResearch for Action. (n.d.). Design principles for accelerating student learning with high-dosage tutoring. https://edresearchforaction.org/research-briefs/accelerating-student-learning-with-high-dosage-tutoring/
- Education Endowment Foundation. (n.d.). Teaching and Learning Toolkit. https://educationendowmentfoundation.org.uk/education-evidence/teaching-learning-toolkit
- Goldhaber, D., Jochim, A., Lake, R., & Rotherham, A. J. (2025, February 21). Mend, don’t end, the Institute of Education Sciences. The 74. https://www.the74million.org/article/mend-dont-end-the-institute-for-education-sciences/
- Gray, L., & Lewis, L. (2021). Use of educational technology for instruction in public schools: 2019–20 (NCES 2021-017). U.S. Department of Education, National Center for Education Statistics.
- Higgins, E., & Doolittle, F. (2025). Blueprint in Action memo: Innovating with implementation to solve education’s “last mile” problem. ALI Coalition.
- Hill, C. J., Scher, L., Haimson, J., & Granito, K. (2023). Conducting implementation research in impact studies of education interventions: A guide for researchers (NCEE 2023-005). U.S. Department of Education, Institute of Education Sciences, National Center for Education Evaluation and Regional Assistance. https://ies.ed.gov/use-work/resource-library/report/evaluation/implementation-research-impact-studies-guide-researchers
- Institute of Education Sciences. (n.d.). Accelerate, Transform, Scale Initiative: IES’s strategy for bringing bold, innovative ideas to education scale. U.S. Department of Education. https://ies.ed.gov/accelerate-transform-scale-initiative-iess-strategy-bringing-bold-innovative-ideas-education-scale
- Institute of Education Sciences. (n.d.). Cost analysis: A starter kit. U.S. Department of Education. https://ies.ed.gov/use-work/resource-library/resource/cost-analysis-starter-kit
- Institute of Education Sciences. (n.d.). Developing infrastructure and procedures for the Special Education Research Accelerator. U.S. Department of Education. https://ies.ed.gov/use-work/awards/developing-infrastructure-and-procedures-special-education-research-accelerator
- Institute of Education Sciences. (n.d.). Performance measures. U.S. Department of Education. https://ies.ed.gov/about/national-center-education-research-ncer/performance-measures
- Institute of Education Sciences. (n.d.). Predoctoral Interdisciplinary Research Training Programs in the Education Sciences. U.S. Department of Education. https://ies.ed.gov/funding/research/programs/research-training-programs-in-the-education-sciences/predoctoral-interdisciplinary-research-training-programs-in-the-education-sciences
- Institute of Education Sciences. (n.d.). Reading for Understanding Research Initiative. U.S. Department of Education. https://ies.ed.gov/funding/research/programs/reading-for-understanding-research-initiative
- Institute of Education Sciences. (n.d.). Special Education Research Accelerator Phase 2: Identifying Generalization Boundaries. U.S. Department of Education. https://ies.ed.gov/use-work/awards/special-education-research-accelerator-phase-2-identifying-generalization-boundaries
- Institute of Education Sciences. (n.d.). Standards for Excellence in Education Research. U.S. Department of Education. https://ies.ed.gov/use-work/standards-excellence-education-research-seer
- Institute of Education Sciences & National Science Foundation. (2013). Common guidelines for education research and development (NSF 13-126). U.S. Department of Education; National Science Foundation. https://www.nsf.gov/pubs/2013/nsf13126/nsf13126.pdf
- Jackson, C. (2022). Democratizing the development of evidence. Educational Researcher, 51(3), 209–215. https://doi.org/10.3102/0013189X211060357
- Jackson, C., Bassok, D., Boulay, B., Kurlaender, M., Page, L., & Tipton, E. (2025, February 24). Cutting research funding would make education less effective and efficient. Brookings. https://www.brookings.edu/articles/cutting-research-funding-would-make-education-less-effective-and-efficient/
- Jackson, C., Wilson, S. J., & Glenn, M. (2024). How does the Prevention Services Clearinghouse rate the design and execution of studies? (Handbook of Standards and Procedures, Version 2.0). Office of Planning, Research, and Evaluation, Administration for Children and Families, U.S. Department of Health and Human Services. https://acf.gov/opre/report/fact-sheet-how-does-prevention-services-clearinghouse-rate-design-and-execution-studies
- Kraft, M. A. (2020). Interpreting effect sizes of education interventions. Educational Researcher, 49(4), 241–253. https://doi.org/10.3102/0013189X20912798
- LearnPlatform by Instructure. (2026). EdTech Top 40: K–12 edtech engagement. https://www.instructure.com/edtech-top40
- McCormick, M., Wilson, S., & Dymnicki, A. (2024). Identifying core components in fatherhood programs: A meta-analytic approach (OPRE Report 2024-09). Office of Planning, Research, and Evaluation, Administration for Children and Families, U.S. Department of Health and Human Services.
- Merod, A. (2023, July 10). Districts used 2,591 ed tech tools on average in 2022-23. K-12 Dive. https://www.k12dive.com/news/school-districts-ed-tech-use/685995/
- Meyer, K., Perera, R. M., Reber, S., & Valant, J. (2026, February 20). Education FAQs: Checking in on the Department of Education. Brookings. https://www.brookings.edu/collection/why-we-have-and-need-a-us-department-of-education/
- National Academies of Sciences, Engineering, and Medicine. (2022). The future of education research at IES: Advancing an equity-oriented science. The National Academies Press. https://doi.org/10.17226/26428
- National Center for Education Statistics. (n.d.). Fast facts: Expenditures. U.S. Department of Education, Institute of Education Sciences. https://nces.ed.gov/fastfacts/display.asp?id=66
- National Center for Education Statistics. (n.d.). National Center for Education Statistics. U.S. Department of Education, Institute of Education Sciences. https://nces.ed.gov/
- National Center for Education Statistics. (n.d.). Statewide Longitudinal Data Systems Grant Program. U.S. Department of Education, Institute of Education Sciences. https://nces.ed.gov/programs/slds/
- National Institutes of Health. (n.d.). NIH Data Book: Appropriations history by institute/center (FY 1938–present). https://report.nih.gov/nihdatabook/report/5
- National Institutes of Health, Office of Portfolio Analysis. (n.d.). Office of Portfolio Analysis. https://dpcpsi.nih.gov/opa
- Northern, A., & Opp, A. (2026). Reimagining the Institute of Education Sciences: A strategy for relevance and renewal. U.S. Department of Education. https://ies.ed.gov/ies/2026/02/reimagining-ies
- Office of Management and Budget. (n.d.). Assessing program performance using the PART. George W. Bush White House Archives. https://georgewbush-whitehouse.archives.gov/omb/performance/
- Olsen, R. B., Orr, L. L., Bell, S. H., & Stuart, E. A. (2013). External validity in policy evaluations that choose sites purposively. Journal of Policy Analysis and Management, 32(1), 107–121. https://doi.org/10.1002/pam.21660
- Pearson, P. D., Palincsar, A. S., Biancarosa, G., & Berman, A. I. (Eds.). (2020). Reaping the rewards of the Reading for Understanding Initiative. National Academy of Education.
- Ritter, S., Murphy, A., Fancsali, S. E., Fitkariwala, V., Patel, N., & Lomas, J. D. (2020). UpGrade: An open source tool to support A/B testing in educational software. Proceedings of the First Workshop on Educational A/B Testing at Scale (Learning @ Scale 2020).
- SAM.gov. (2025). Regional Educational Laboratory Program: 2027. U.S. General Services Administration. https://sam.gov/opp/5451117f9f9744bba876b7b330d64f0a/view
- Stuart, E. A., Bell, S. H., Ebnesajjad, C., Olsen, R. B., & Orr, L. L. (2017). Characteristics of school districts that participate in rigorous national educational evaluations. Journal of Research on Educational Effectiveness, 10(1), 168–206. https://doi.org/10.1080/19345747.2016.1205160
- Tipton, E., & Olsen, R. B. (2022). Enhancing the generalizability of impact studies in education (NCEE 2022-003). U.S. Department of Education, Institute of Education Sciences, National Center for Education Evaluation and Regional Assistance. https://ies.ed.gov/use-work/resource-library/report/guide/enhancing-generalizability-impact-studies-education
- Tipton, E., Spybrook, J., Fitzgerald, K. G., Wang, Q., & Davidson, C. (2021). Toward a system of evidence for all: Current practices and future opportunities in 37 randomized trials. Educational Researcher, 50(3), 145–156. https://doi.org/10.3102/0013189X20960686
- Tseng, V., & Coburn, C. (2019). Using evidence in the U.S. In A. Boaz, H. Davies, A. Fraser, & S. Nutley (Eds.), What works now? Evidence-informed policy and practice (pp. 351–368). Policy Press.
- U.S. Department of Education. (2025, March 11). U.S. Department of Education initiates reduction in force. https://www.ed.gov/about/news/press-release/us-department-of-education-initiates-reduction-force
- Wadhwa, M., Zheng, J., & Cook, T. D. (2024). How consistent are meanings of “evidence-based”? A comparative review of 12 clearinghouses that rate the effectiveness of educational programs. Review of Educational Research, 94(1), 3–32. https://doi.org/10.3102/00346543231152262
- Weiss, M. J., Bloom, H. S., & Singh, K. (2023). What 20 years of MDRC RCTs suggest about predictive relationships between intervention features and intervention impacts for community college students. Educational Evaluation and Policy Analysis, 45(4), 569-597.
- What Works Clearinghouse. (2021). WWC reporting guide for study authors. U.S. Department of Education, Institute of Education Sciences, National Center for Education Evaluation and Regional Assistance. https://ies.ed.gov/ncee/wwc/Docs/ReferenceResources/WWC_Author_Guide_Jul2021.pdf
- What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0. U.S. Department of Education, Institute of Education Sciences, National Center for Education Evaluation and Regional Assistance. https://ies.ed.gov/ncee/wwc/Docs/referenceresources/Final_WWC-HandbookVer5_0-0-508.pdf
- Whitehurst, G. J. (2018). The Institute of Education Sciences: A model for federal research offices. The Annals of the American Academy of Political and Social Science, 678(1), 124–133. https://doi.org/10.1177/0002716218768243
- Wolf, B., & Harbatkin, E. (2023). Making sense of effect sizes: Systematic differences in intervention effect sizes by outcome measure type. Journal of Research on Educational Effectiveness, 16(1), 134–161. https://doi.org/10.1080/19345747.2022.2071364
- Wong, V. C., Tipton, E., & Spybrook, J. (2025, April 10). How federal investments in education research help students succeed. Brookings. https://www.brookings.edu/articles/how-federal-investments-in-education-research-help-students-succeed/
The Future of Evidence in Education Network · Sidenotes appear in the right margin on wide screens; on narrow screens, tap a numbered reference to open its note.