The Future of Evidence in Education Network

Evaluating Causal Evidence in Education: A Practical Guide

Ten principles for judging whether the evidence behind a program is strong enough, relevant enough, and complete enough to inform the decision in front of you.

Authors
Future of Evidence in Education
Convening Members
Report editors
Cara Jackson
Luke Miratrix
Series editor
Vivian C. Wong
Published
2026

Authors and Convening Members

These reports were developed through the Future of Evidence in Education convening held in December 2025. All convening members participated in the full convening and contributed to the ideas developed across the two reports. Members then worked in smaller groups to develop each report. All convening members are authors of both reports.

Convening members

  • Rekha BaluUrban Institute
  • Beth BoulayAnnenberg Institute at Brown University
  • Brooks BowdenUniversity of Pennsylvania
  • Ben DomingueStanford University
  • Dan GoldhaberAmerican Institutes for Research and University of Washington
  • Kelly HallbergUniversity of Chicago
  • Erin HigginsAlign R&D
  • Cara JacksonCenter for Outcomes Based Contracting at the Southern Education Foundation
  • Luke MiratrixHarvard University
  • Robert OlsenGeorge Washington Institute of Public Policy
  • Laura PeckRutgers University
  • Jessaca SpybrookUniversity of South Carolina
  • Elizabeth TiptonNorthwestern University
  • Betsy WolfDistrict of Columbia Office of the State Superintendent of Education
  • Brian WrightUniversity of Virginia
  • Vivian C. WongUniversity of Virginia

Report teams

This report

Evaluating Causal Evidence in Education: A Practical Guide

  • Elizabeth Tipton
  • Jessaca Spybrook
  • Betsy Wolf
  • Laura Peck
  • Beth Boulay

Report editors: Cara Jackson and Luke Miratrix Series editor: Vivian C. Wong

Companion report

Necessary but Not Sufficient: Six Design Principles for a Stronger Education R&D System

  • Beth Boulay
  • Rekha Balu
  • Ben Domingue
  • Brooks Bowden
  • Dan Goldhaber
  • Rob Olsen
  • Brian Wright

Report editors: Kelly Hallberg and Erin Higgins Series editor: Vivian C. Wong

Suggested citation

Future of Evidence in Education Convening Members. (2026). Evaluating causal evidence in education: A practical guide. Future of Evidence in Education Network. https://www.edevidencenetwork.org/evaluating-causal-evidence-in-education-a-practical-guide

Download this report

How this Report was Developed

In December 2025, we convened a cross-sector group of researchers, evaluators, and evidence leaders for a two-day discussion focused on the education evidence system. Participants brought their years of experience designing and analyzing studies, developing and applying evidence standards, evaluating programs, studying how findings carry across settings, estimating costs and benefits, and helping decision-makers use evidence in policy, procurement, and practice.

The purpose of the meeting was to develop a shared diagnosis of the challenges confronting the system as the federal role in supporting education research becomes more uncertain. We began with decisions that state and local leaders often face, asked what kinds of evidence those decisions require, and worked backward to identify the conditions needed for that evidence to remain credible, interpretable, and useful.

As the discussion developed, it became clear that the work needed to be organized into two related reports. This report presents ten principles for the responsible interpretation and use of evidence. It focuses on how people can assess what evidence can and cannot support, what questions to ask, and what remains uncertain when deciding whether to adopt, pilot, or scale a program. A companion report focuses on how the education evidence system can produce evidence that is more cumulative and useful over time. It examines the institutions, incentives, documentation practices, and synthesis efforts that shape what evidence gets produced, how it is interpreted, and whether it can inform future decisions.

The principles in this report grew out of the practical challenges the convening members have encountered when interpreting results from studies conducted in the field. Evidence is often incomplete, uneven, or not presented in ways that line up with the decisions leaders need to make. The report brings together the kinds of questions researchers and evaluators ask when they read evidence, weigh uncertainty, and judge what a finding can and cannot support. The ten principles of this report are meant to support consideration of whether evidence is incomplete, mixed, or overstated. The aim is to lay bare our habits of thinking. We hope illuminating these habits help leaders decide when evidence is strong enough to inform a decision and know where caution is still warranted.

Introduction

State and local leaders are asked to make high-stakes decisions on tight timelines about programs, products, practices, and policies all the time. These decisions have costs for schools in terms of time, money, and social capital. Evidence is useful when it helps determine whether a new approach is likely to deliver the benefits it promises, whether the benefits are worth the costs, and whether the approach is a good fit for their context. With evidence, leaders can make these high-stakes decisions with greater confidence, and better results.

Even programs that have earned an “evidence-based” label are not necessarily the right choice for a particular school, district, or state. The evidence behind that label may come from a different setting, with different students, under different conditions, against a different comparison, or on outcomes that do not match the decision leaders face. Summaries and evidence ratings can be useful, but they often leave out information leaders need to judge whether the evidence supports the decision in front of them.

This report is meant to help readers judge whether the available evidence is strong enough, relevant enough, and complete enough to inform a decision. We do not rank programs or certify “what works.” We offer principles that can help leaders ask better questions before they adopt, pilot, scale, or reject a program. The aim is to make decisions more grounded in the evidence that exists, while making clear what the evidence can and cannot show. The table below summarizes the principles, and the following sections explain how to use them. We have a small number of technical terms; we define these in boxes labeled “Key Term” throughout the report as these terms come up.

Table 1 · Ten Principles for Assessing Evidence

Scroll sideways to see the full explanation column.

# Short name The principle The explanation
On reading the evidence initially
1 Implementation costs Assess the full costs and conditions of implementation. Some programs may be effective but require resources beyond what is feasible.
2 Evidence transparency Look for enough study information to judge the credibility of the evidence. We need study details to judge credibility and apply the remaining principles.
On evaluating a study’s findings
3 Comparison group Identify the comparison for the intervention. Without a point of comparison, we cannot know whether a change is due to the program, or something else.
4 Interpreting effect sizes Take the outcome measure, comparison condition, study scale, and precision into account when comparing effect sizes. A large effect size can reflect a weak comparison group, a narrow outcome measure, or a small scale rather than the program working well.
5 Practical significance Take both practical and statistical significance into consideration. Statistical significance can indicate the extent to which evidence shows an effect, but not whether the effect is large enough to matter for the decision at hand.
On assessing the relevance of evidence
6 Meaningful outcomes Evaluate whether the reported outcomes are relevant to a given context. Evidence on one outcome does not mean the program will improve other outcomes that leaders care about.
7 Contextual considerations Assess the extent to which your context and planned implementation are like those of the study. Success in one setting does not guarantee success in another, and results may vary across different implementation conditions.
8 Body of evidence Consider the body of evidence when multiple studies exist. Evidence from multiple studies can show whether results hold across settings, populations, outcomes, and implementation conditions.
On evaluating more complex claims
9 Subgroup claims View subgroup findings with caution and check whether they were planned, are justified, and are supported across settings. Subgroup findings can be due to chance alone; they are more credible whether they were planned in advance, tied to the program’s theory of action, and appear across multiple studies or settings.
10 Personalization claims Apply the same evidence standards to personalized and algorithmic tools. Complex tools still need evidence that they improve outcomes relative to a clear alternative.

The Decision Problem

Education leaders rarely face a simple choice between an evidence-based program (such as a tutoring provider, new curriculum, training program, or a technology product) and no program at all. More often, they are choosing from several plausible programs. In making these choices, they may draw on implementation reports, user data, internal analytics, recommendations from colleagues, vendor materials, pilot results, formal evaluations, and studies conducted in other settings. Each source offers some information about whether a program is useful, promising, or likely to improve student outcomes, but that information varies in quality, relevance, and credibility.

How can a leader navigate these various sources of information? Even if they narrow the options to one promising solution, what should they consider before moving forward? In our view, such a decision boils down to two fundamental questions:

What would implementing the program require in my school or district?

If implemented, should we expect to see benefits that would justify the investment?

Answering these questions requires examination of evidence. In this report, evidence refers to information that helps leaders assess what a program requires, whether it is likely to causally improve outcomes, for whom, under what conditions, and with what tradeoffs or uncertainty. Leaders need evidence about the staffing, time, training, materials, technology, disruption, and other demands required to implement the program to assess costs. They also need evidence about whether the program is likely to produce enough benefit to justify those demands.

Evidence on causal effects can drive real change. As one example, the Institute of Education Sciences (IES)-supported evaluations of CUNY’s Accelerated Study in Associate Programs (ASAP) found that a package of academic and financial supports substantially improved college completion. These evaluations showed a greater share of students graduated under ASAP than would have graduated under CUNY’s prior services and helped support expansion of ASAP in New York and replication in other states, including Ohio. Strong causal evidence does not remove all uncertainty or show exactly where a program will work, but it can give leaders a stronger basis for deciding which programs deserve serious consideration.

Key term

Causal effect

The extent to which a program causes student outcomes to improve is known as the causal effect of a program, or the causal impact. To determine a program’s causal effect, we ask “to what extent did a program improve outcomes relative to what would have happened otherwise?” Studies that answer this question are sometimes referred to as impact studies or causal inference research.

Over the last twenty years, the federal government has invested in evaluating the causal effects of programs. Doing so serves a public good, since the results of publicly accessible evaluations can inform decision-making nationwide. Program evaluations have been funded through the IES (Wong et al., 2025) and the Investing in Innovation program, now known as the Education Innovation and Research program (Boulay et al., 2018; Goodson et al., 2024). Many evaluated programs initially seemed promising, practical, or appealing. They often had advocates for broader use. Unfortunately, while all programs are built with good intent, hoped-for benefits are not always achieved. Some programs that were evaluated produced positive effects, but many did not (Barshay, 2024). Before investing time, money, and effort in changing practice, leaders should seek reasonable reassurance that the change is likely to benefit students.

One step toward reasonable reassurance is to look closely at what the evidence does and does not show with regard to causal effects. Program materials often make evidence sound more settled than it is. The principles on Table 1 are meant to help leaders ask the questions that can protect consequential decisions from overclaimed findings, incomplete reporting, and evidence that may not apply to their setting.

Reading the Principles

The principles in this document are what we ourselves use to interpret evidence. We explain each principle in turn, along with why it matters. We also include questions leaders might ask as they reflect on the principle in their own context. The idea is to give a practical way to ask what the evidence supports, what it does not support, and what still needs to be judged.

In Appendix A, we offer additional readings and resources. The Future of Education Evidence Network website offers even more resources, including case studies that show how the principles can be used to think through real-world decisions.

Ten Principles of Evidence

The principles are arranged in a sequence of what questions arise when thinking through an evidence-based decision. We group by larger themes. The first set, Initial Read of the Evidence (Principles 1 and 2), asks what a program would require in practice and whether the evidence is reported with enough transparency to be understood. The second set, Evaluating Evidence of Causal Effects (Principles 3 through 5), focuses on whether the evidence supports a conclusion that the program caused changes in student outcomes. The third set, Assessing the Relevance of Evidence (Principles 6 through 8), considers whether the findings are meaningful for a decision in a specific context. The final set, Evaluating More Complex Claims (Principles 9 and 10), applies the earlier principles to claims that require extra care, including personalized programs and claims about effects for particular groups of students.

Principles 1–2

Initial Read of the Evidence

Before assessing whether evidence shows that a program improves student outcomes, leaders need a first read of both the program and the evidence. They need to understand what is required to implement the program in practice, including staff time, training, materials, data systems, and the full cost of delivery. Leaders also need enough information about the evidence to understand what was studied and how the conclusions were reached. The first two principles focus on these initial judgments.

1 Implementation costs

Assess the full costs and conditions of implementation.

In practice

A district is considering a new digital learning program with affordable licenses, and vendor materials suggest that other districts have seen gains. However, once leaders look more closely, they realize that successful implementation would require teacher training, time for staff to integrate the program into existing instruction, technical support, monitoring to make sure students use the program as intended, and someone responsible for reviewing data and adjusting use over time. The program may still be worth adopting, but all these costs need to be accounted for when deciding if it is worth it.

The full cost of program implementation is more than the price of edtech licenses or a set of curriculum materials. It also includes the staff time, training, and oversight needed to deliver the program with enough integrity to have a chance of improving outcomes. If the program is not enough of an improvement over current practice, or implementation requirements exceed local capacity, then results may be disappointing even if the program worked elsewhere.

Questions for interpreting evidence

Does adopting this program represent a meaningful change from current practice? If the program is not substantively different from what is already in place, then benefits of switching may be small. In that case, strengthening implementation of the current program may lead to better outcomes at a lower cost than switching to a new but similar one.

What resources are required? How much will effective implementation cost, and is it feasible? Consider start-up costs (training, set-up), recurring costs (coaching, data review), and the staff time required to coordinate and troubleshoot. Adopting programs that are not well-aligned to other ongoing initiatives may cause additional coordination work. The time required from teachers, coaches, and administrators, along with materials, training, and ongoing support, should all be counted as part of the full cost.For a practical guide to estimating program costs, see IES’s Cost Analysis: A Starter Kit.

What will the program displace? New programs that require a substantial reallocation of time or money have high opportunity costs: they reduce attention to other priorities and may require cutting back on other educational opportunities. If the time and resources come from other successful initiatives, then overall impact includes both gains from the new program and losses from the reallocation of resources. Name what will be reduced or dropped.

Are the program’s components aligned with existing evidence about effective practice? A program may look promising as a package, but specific activities, strategies, and resources can vary in how much they align with what is known from prior research. Look for whether the instructional practices of the program are consistent with evidence summaries such as relevant IES practice guides (see Appendix A for additional resources).

What components are essential and where is there room for local adaptation? It is natural to adapt a program to fit local needs and conditions. Flexibility can help address community members’ concerns and improve local fit, but adaptations could also inadvertently remove or weaken components integral to the program’s effectiveness. When assessing how a program will be implemented, identify which components are essential and which can be modified. Ask whether the available evidence applies to the version the district would implement. If existing policies or resource constraints make it difficult to implement essential components, consider whether those barriers can realistically be addressed. If not, the program may not be a good choice for the district.

Who has to do what? Which roles are responsible for delivery (teachers, coaches, administrators, vendors), and what is the expected routine (frequency, duration, scheduling)? Studies may point to implementation features, such as regular attendance, ongoing coaching, minimum dosage, staffing qualifications, or time allocated for delivery, that appear important to maintaining effectiveness. Use that information to plan for a strong implementation.

Will district and school leaders, teachers, and students use the program as expected, or resist it? Low participation, uneven use, or burnout can undermine even well-designed programs. If past initiatives have struggled with uptake or sustained use, that is a signal to plan for additional supports. If engagement data are not available, ask what evidence exists about uptake in local settings and what supports were used to achieve it. If educators, or parents, or others involved have substantial reservations about the program, then it may be difficult to implement effectively.

How will program implementation be monitored locally to ensure conditions for success? Describe what “implementation as intended” means for the program and what indicators of implementation quality will be tracked (such as participation rates and usage) and how (logs, observations, brief surveys). Identify who is responsible for reviewing the data and how often.

What would count as evidence that implementation is off-track, and what mid-course corrections would be taken? Plan in advance for how concerns will be escalated and handled if participation is low, delivery is uneven, or supports are not being provided. Do not wait until the end of the year to decide what signals would require attention or how to respond.

Takeaway

Map the full implementation demands against knowledge of local context and capacity, and identify what the program would displace. Include the amount of staff time and data collection needed to monitor implementation effectively as part of the full cost.See Hill et al.’s Conducting Implementation Research in Impact Studies of Education Interventions: A Guide for Researchers. For more information, see Appendix A.

2 Evidence transparency

Look for enough study information to judge the credibility of the evidence.

In practice

A district is considering a tutoring program that neighboring districts say has been successful. The vendor shares a short summary describing short-term gains in reading, along with teacher testimonials and a few slides on student progress. But the materials do not include a full report, do not explain how students were selected to participate, do not describe what the program was compared to, and do not say much about what implementation required.

To evaluate the credibility and relevance of evidence, leaders need to know what was studied, how it was studied, what outcomes were measured, what the program was compared to, and what implementation required. If the underlying study is not publicly available, or if findings are only reported in broad summary form, then we cannot know whether the evidence actually indicates that a program is effective at improving student outcomes.

When studies are reported clearly and publicly, people can inspect the evidence and determine whether it meets the standards to be described as a causal effect.The What Works Clearinghouse (WWC) provides one widely used set of standards for reviewing the quality of education research. When evidence is produced by an organization that benefits from the program’s adoption, leaders should pay especially close attention to what study details are reported and whether similar findings have been produced by neutral researchers.

Questions for interpreting evidence

Have studies of the program been published or made publicly available? Look for full reports, not only marketing summaries. It is a good sign when program developers are willing to engage in evaluation, describe how the study was carried out, and provide detailed results.

Do studies include detailed information on study context and program implementation? Such detail could include program components, who delivered the program, how often it occurred, resources required, and costs. If implementation is described only in general terms, then it is difficult to know what caused changes in outcomes. Context and implementation information is also necessary for assessing the full costs of implementation (Principle 1: Implementation costs) and for determining whether the study is relevant (Principle 6: Meaningful outcomes and Principle 7: Contextual considerations).

If adaptions of the program exist, do any studies speak to the adapted versions? Ongoing evaluation is important when programs evolve over time. Evidence on an old version of a program tells leaders about that version of the program. While adaptations may have improved the program, it is also possible that the program was altered in ways that make it less effective than the original version (e.g., it was oversimplified to control cost).

Did researchers clearly specify the research questions, methods, and analyses, and document the study plans before conducting analyses? Researchers make many decisions while collecting, cleaning, and analyzing data, and these decisions can subtly shape the findings. Plans written in advance, such as a publicly posted protocol, analysis plan, or pre-registration plan, reduce the risk of incorrectly concluding there was an impact when there was none (“false positives”), which might occur if researchers’ decisions are influenced by whether they generate interesting results. Some journals use Open Science Badges to note when authors have shared data or materials or registered their plans in advance. These are useful markers of transparency, but readers should still examine what was shared in advance and whether it maps to the findings reported.

Do studies include a clear and thoughtful discussion of limitations and caveats? Such a discussion signals a cautious and self-reflective analyst. For example, does the report address any differences between the program and comparison group that may call into question whether the program’s apparent effectiveness might be explained by other factors (see Principle 3: Comparison group)? If outcome measures are narrow or short-term, is this mentioned as a caveat (see Principle 4: Interpreting effect sizes)? If the sample has limited diversity or a narrow range of scores, are findings appropriately described as limited to those students? If results vary across sites, subgroups, or outcomes, is that variation reported? Every study has caveats; failing to discuss limitations indicates a lack of transparency.

Are studies conducted by independent researchers with no professional or financial stake in the success of the programs? Developer studies may selectively highlight favorable findings while minimizing or omitting less favorable results. And indeed, studies conducted or commissioned by program developers find larger positive effects on student outcomes than independent studies (Wolf et al., 2020). Pay close attention to what is reported, what is not reported, and whether results have been reproduced by others.

Takeaway

Seek full, publicly available study reports or peer-reviewed articles rather than marketing summaries and apply extra scrutiny when the organization promoting a program also conducts its own evaluation. Trust findings only when the study details are transparently reported.

Principles 3–5

Evaluating Evidence of Causal Effects

In the next set of principles, we focus on evaluating evidence that is framed as establishing whether a program is effective at improving student outcomes. These principles are intended to help leaders sort through whether the available evidence can credibly be taken as meaning the program did meaningfully improve student outcomes, at least for the study’s population.

3 Comparison group

Identify the comparison for the intervention.

In practice

A group of children received a supplementary online math program, and their math test scores increased from the beginning to the end of the year. While this shows change over time, many reasons other than the program could explain that change. Students’ scores might have improved just as much using a different program, or they may have improved because teachers had more or better professional development in the same timeframe. To know if the supplementary math program caused the improved math test scores, we need to know how test scores would have changed if students had not had the supplementary math program (the “counterfactual”—see callout box).For an overview of causal inference in education and social sciences, see Shadish, Cook, and Campbell’s Experimental and Quasi-Experimental Designs for Generalized Causal Inference.

To produce credible evidence of the causal effect of the program, we need a comparison group of similar students. A comparison group allows us to attribute differences in outcomes to the program rather than to other changes that occurred at the same time. A comparison group that is like the group assigned to the program allows us to attribute differences in outcomes to the program rather than to pre-existing differences between groups. Knowing whether a program causes a desired change can help justify spending money to adopt and implement the program.

Key term

Counterfactual

Researchers gauge the program’s effect relative to what would have happened in the absence of the program or under the alternative that would otherwise be offered. We often call this “what would have happened otherwise” the counterfactual. We generally use the outcomes of a comparison group to assess what that counterfactual would likely have been.

Questions for interpreting evidence

Does the study have a comparison group? Studies that lack any comparison group (e.g., pre-post studies, or before-and-after assessments of the same students) can provide descriptive evidence of change over time, but do not provide evidence of the program’s causal effect. Scores often improve over time because of normal growth and instruction, even without a new program. Changes in staffing, schedules, other supports, or broader events can also explain before-and-after changes. A comparison group helps separate change that would have happened anyway from change plausibly attributable to the program.

Were students assigned to groups randomly? When random assignment or a lottery is used to determine who is assigned to the program, the left-out participants form a comparison group that is usually quite similar to the group assigned to the program. This creates a clean comparison for assessing the benefits of treatment. By contrast, if people self-select into programs, or educators intentionally assign certain students to programs, the factors that drive those choices can cause the program and comparison groups to be inherently different (this is called “selection bias”). In brief, random assignment into the program avoids the problem of selection bias.

Is the comparison group arguably similar to the group who got the program? When random assignment is not feasible, check for evidence that the comparison group is similar before the program was introduced, especially on prior achievement and other characteristics related to the outcomes. If the groups are not comparable – for example, if higher performing students are more likely to participate, or if educators choose who is offered the program – then differences in outcomes may reflect underlying differences in the groups rather than the effect of the program.Researchers refer to similarity between comparison and treatment groups as baseline equivalence (see the WWC Standards Brief for Baseline Equivalence for more details).

Are certain people excluded from the study, and, if so, does that change interpretation of the findings? Excluding people from the study is fine when based on a pre-existing characteristic that can be measured in both the program and comparison group (such as excluding English learners from the study). In contrast, exclusions based on experiences that occur after the program begins can undermine the similarity of the groups. For example, some studies exclude treated students from the analysis based on number of sessions attended, lessons completed, or other measures of their engagement with the treatment. In the comparison group, we cannot identify which students would have participated at that level had they been offered the program. This creates an unfair comparison if students who remain in the treated sample are systematically different from the students excluded.

Takeaway

Ask whether the study includes a comparison group and whether the comparison group was similar to the treated group before treatment was given. Pre-post studies without a comparison group describe change over time, not causal effects. Assess whether the comparison group looks like the group assigned to the program on characteristics such as prior achievement; if they are not similar, then differences in characteristics may explain any differences in outcomes.

4 Interpreting effect sizes

Take the outcome measure, comparison condition, study scale, and precision into account when interpreting effect sizes.

In practice

A study of a reading program reports an effect size of 0.35. While initially promising, further investigation revealed the evaluation of the program was done in a district where teachers had not previously been well supported or trained on reading instruction, and furthermore the test used was only measuring phonological awareness. Given these two concerns of contrast and an overly narrow outcome measure, the stakeholders considering the program decided to dig deeper into whether that program was a good fit for them.

Key term

Effect size

Effect sizes describe how much a program changed an outcome, such as student achievement, in standardized units. In many education studies, effect size is measured in standard deviation units and shows how much a program changed an outcome relative to the spread of scores in that study. An effect of 0.20, for example, means that the group assigned to the program scored 0.20 standard deviations higher than the comparison group on that outcome, in that sample.Effect size can be measured in other ways (e.g., the raw mean difference, the correlation coefficient, the odds ratio); we focus here on standard deviation units since it is commonly used in education research. Because the standard deviation reflects how much scores vary within the study sample, the same difference in scores can produce different standardized effect sizes in samples with different amounts of variation. Effect sizes from different studies may therefore differ because of both differences in program impact and differences in the measures and samples used.

A large effect size estimate might reflect a real and meaningful gain but can be inflated by factors other than a program’s causal impact. In the illustration above, for example, the low quality of instruction made it easier to show improvement; the program might not give as high a return in a district where reading instruction is stronger. Similarly, overly narrow outcomes are easier to change, but improving on a narrow outcome might not translate into better overall reading ability. Other drivers of effect size not related to the actual program include highly selected samples and contexts, or intensive supports provided during the evaluation. Studies might also be too small to estimate effects precisely.

An effect size estimate is an estimate produced by a particular study, using a particular outcome, for a particular group of students, relative to a particular comparison condition, and all estimates have uncertainty. For that reason, the size of a reported effect should be interpreted in context, not treated as a definitive summary of how well a program works. Because effect sizes can be driven by aspects of a study not related to the actual effectiveness of the program, effect sizes should generally not be used on their own to rank different interventions, programs, or practices against one another. They are only one element—albeit an important one—in assessing the utility of programs.

Questions for interpreting evidence

How narrow or broad is the outcome? Some outcomes are narrow and closely aligned to the program (such as letter-sound fluency), while others capture broader constructs (such as reading achievement). All else equal, we are more likely to see larger effects on outcomes closely aligned with the program and measured soon after implementation, and smaller effects on broader outcome measures or outcomes measured after more time has passed.

Who is included in the study? First, some studies drop students who did not participate or receive the full dosage from the treatment group sample, which may result in larger effect sizes than you would see if the full sample of students offered the treatment were retained in the study.Dropping students in this way can also make the estimate entirely invalid; treat such studies with caution. Second, a standardized effect size expresses the difference between groups relative to the variation in scores, so the amount of variation in the study sample can affect the reported effect size. In particular, the same difference will appear larger in a more homogenous sample with more similar scores (for example, when a study only includes low-performing students) and smaller in a sample with a wider range of scores (as when a study includes all students).

What is the program compared against? Often the comparison is described as “business-as-usual,” but this can mean very different things. Ask what students in the comparison group received. An effect size is the change relative to this contrast condition. For example, if students in the treatment group are receiving one-on-one tutoring, and students in the comparison group receive small-group tutoring, the contrast is between the types of tutoring—not between tutoring and not being tutored at all. We would not expect to see as large an effect when the contrast is small, even if the program is of high quality.

At what scale and under what conditions was the program studied? Small pilot studies can provide more training, monitoring, technical assistance, or researcher involvement than a district could sustain during routine implementation. These conditions can contribute to larger effects than might be found when a program is implemented across more schools with less support. When interpreting effect size, examine the level of implementation support provided in the study and whether the program has produced similar effects in studies conducted at a larger scale. Ask whether the program could be similarly supported if rolled out in a new context without the original researchers’ involvement.

How precise is the estimate? Sample size does not determine whether a program’s true effect is large or small, but it does affect how much uncertainty surrounds the estimate. Estimates from small studies can vary a great deal and may appear unusually large or small by chance. Look at the confidence interval, not only the reported effect size, and ask whether it includes effects that would lead to different conclusions about the magnitude of the effect.

Do other ways of describing the result tell the same story? If the study translates findings into months of learning or other terms, check how those findings were constructed and whether they are consistent with the reported effect size. Translations may rest on additional assumptions or simplifications that can misrepresent findings (Jackson, 2025; Pane & Baird, 2019). Just as effect sizes are relative to some reference group, translations are indexed by some reference group as well. Be sure to check if translated findings are relative to similar students, at similar grade levels, on similar tests, to your target context.

Takeaway

Be suspicious of large effect sizes. Based on our collective experience, effects above 0.30 on broad achievement outcomes are highly unusual and should prompt a closer look at the outcome measure, the comparison condition, and the analytic choices. A large effect is not automatically wrong, but it could reflect a narrow measure, a weak comparison condition, a high degree of implementation support, inappropriate translation, or be simply a product of the uncertainty in the evaluation itself. Do not compare effect sizes across studies without considering differences in outcomes, study samples, comparison conditions, and settings and scale.

5 Practical significance

Take both practical and statistical significance into consideration.

In practice

A tutoring program increased the proficiency rate by two percentage points and the result was statistically significant. That gain might be too small to justify the cost of the program, especially if tutoring is expensive, difficult to staff, or displaces other supports. On the other hand, if the program is easy to adopt and relatively low cost, that same gain might be considered worthwhile. This “is it large enough to be worth it” question is one of practical significance.

A statistically significant finding means the study found evidence that the program affected outcomes beyond the kinds of differences that might show up by chance. A statistically significant effect can still be a very small effect, particularly in a large study. Statistical significance is only one piece of the puzzle: a study can find evidence of an effect, but the effect may or may not be large enough to justify the effort involved in implementation.

Practical significance asks whether the estimated effect is large enough to matter for the decision at hand. This judgment depends on the costs, including how much effort, expense, or disruption are expected to be incurred in a particular context. A practically significant effect is one in which the size of the expected benefit justifies the expected effort and expense.

Questions for interpreting the evidence

Is the size of the effect worth the cost? When evaluating whether an effect size is “big enough,” consider if the program is hard or easy to implement. A small effect may be good enough if effort and expenses are low. Even a large effect might not be worth an enormous effort or substantial reallocation of resources from other successful programs. Ask if the expected benefit justifies the expected effort and expense.

Could a large sample be making a very small difference look more important than it is? Be cautious when a report emphasizes statistical significance when the estimated effect is very small. When samples are large, tiny differences can be statistically significant. Consider the practical value of a small effect.

How precise is the estimate? To assess practical significance, especially if the study is small, look for confidence intervals. These are ranges that show how much the true effect could reasonably differ from the study’s estimate. Small studies often produce wide confidence intervals, which indicates more uncertainty about the size of the effect. Consider whether the program would be worth adopting if the true effect is near the low end of the range.

Consider the comparison condition when interpreting the estimated effect. If the study compares the program to a strong business-as-usual, we expect the benefit of the program to be smaller. If the same program is compared to a weak business-as-usual (such as low-quality instruction or limited resources), the benefit of the program might appear larger. Consider whether the business-as-usual in the decision context is relatively weak or strong.

Takeaway

Assess whether the estimated benefit is large enough to matter given what the program costs and what it would replace, and take sample size, the plausible range of true effects, and comparison conditions into account when interpreting statistical and practical significance.

Principles 6–8

Assessing the Relevance of Evidence

Determining whether something that worked in one setting will work somewhere else is difficult. Even when a study provides credible evidence of improved outcomes in one setting, education leaders still have to evaluate whether those findings are likely to hold in their own context. The next principles shift from asking whether evidence indicates that a program worked somewhere to whether the evidence is broad and relevant enough to suggest it would likely work in a new setting.

6 Meaningful outcomes

Evaluate whether the reported outcomes are relevant to your context.

In practice

A study of a mathematics program reports a large, positive effect on an outcome measure that captures students’ ability to multiply fractions. While encouraging, this does not necessarily mean that students’ overall mathematical ability improved. Perhaps the program only focused on fractions, or that was the only part of the program that worked; we do not know whether the program would improve outcomes on a more general measure of mathematical ability.

Narrow or specialized outcomes can be easier to change than broader outcomes such as overall reading or math ability. A program may improve a skill practiced in the program, such as a curriculum-based assessment, a developer-created task, or a short measure of a specific skill, without improving the broader outcomes leaders care about, such as reading comprehension, math achievement, course performance, attendance, or graduation. The same concern applies to researcher-developed measures that are closely tied to the program being studied. In those cases, the program may appear to work because students were taught the same content or skills measured by the test. Leaders should consider whether the outcomes measured align with the breadth of skills they seek to improve.

Questions for interpreting evidence

Do the outcome measures align with how the findings about the program are reported? If stated findings are about student learning, the measures should be of student learning. Usage metrics (e.g., minutes used, lessons completed, or in-program mastery) or satisfaction measures (e.g., user ratings) describe participation and experience. They do not, on their own, measure student learning. Measures also need to match how findings are reported in breadth. For example, if the finding is described as the program improves reading achievement, the outcome measure should capture reading achievement broadly rather than focusing on a narrow subdomain, such as phonemic awareness.

Are the outcome measures well-constructed (valid) and consistently measured (reliable)? Look for evidence that outcome measures were carefully developed, ideally independently of the study, and for technical documentation of validity and reliability consistent with the Standards for Educational and Psychological Testing (see Appendix A). When outcome measures are created by the study team or embedded in the product, ask what evidence exists that the measure captures the intended construct and is not overly tailored to the program. Independent measures provide a useful check since estimated effects tend to be larger on measures that are more closely tied to the intervention (Wolf et al., 2020; Wolf & Harbatkin, 2023).

Are the outcomes measured associated with longer-term outcomes of interest? For short-term outcomes (observed immediately after the end of a program), check for evidence of predictive validity: are those outcomes at least associated with other policy-relevant or longer-term (12 or more months) outcomes that matter? For example, if the short-term outcome is an interim assessment, does it correlate with end-of-year assessments, grade retention, course completion, or high school graduation? If the link is unknown or weak, short-term gains are insufficient evidence for longer-term outcomes.

Takeaway

Check whether the outcomes measured match how findings are described. To determine relevance, consider whether measured outcomes align with the goals leaders have for the program. Finally, give more weight to broader, longer-term, and independently measured outcomes; such measures are more likely to speak to goals in a variety of settings.

7 Contextual considerations

Assess whether your context and planned implementation are similar to those studied.

In practice

A tutoring program had a positive effect on literacy in a district with low student absenteeism. Students received tutoring during the school day from trained full-time tutors, with close oversight from school leaders. Would that program work in a new district? What if the district planned to use the same tutoring program, but delivered after school, using part-time staff? Or what if the district was dedicated to implementing the program as designed, but was also struggling with high student absenteeism? Even if the program itself is the same, differences in staffing, scheduling, and student participation can easily affect whether the district should expect similar results.

“Will this work here?” is a key question for education leaders. No matter how rigorous the causal evidence is, whether we can reasonably expect to see the same results in a new setting depends on a variety of factors. Answering this question requires establishing how similar the study setting is to the setting in which the decision is taking place, with respect to the conditions that allowed the program to succeed. This means looking for information about the study context, including features of the setting that may have shaped implementation quality. The more similarity between the context of the available evidence and the context of the decision being made, the greater the confidence that the program will have similar benefits in the new context.

Questions for interpreting evidence

How similar are the students and settings to those in the study or studies that demonstrate effectiveness? Pay attention to who was served and where the program was implemented. If the students in the study differ in important ways from those in the local context to be served – for example in grade level, prior achievement, language status, or special education identification – treat the evidence as less directly applicable.

Is the study’s comparison condition similar to current practice? The benefit of a program depends not only on the program itself, but also on the comparison condition. A program may look more effective when compared to a weak alternative than when compared to a strong existing support. When interpreting evidence from another setting, ask what the program replaced in the study and whether that is like what it would replace in the local context.

What features of the context matter for the program to be successful? Consider contextual features, such as state or local policies, teachers’ union contracts, and qualifications of the staff implementing the program. If these conditions differ from the study setting, implementation may look different, and outcomes may then differ as well.

Would the program be as well-resourced and implemented with the same level of quality as in the evidence examined? If not, how might program implementation be compromised? If the program is scaled back, be explicit about what will change (time, staffing, training, supports). Expect “lighter touch” versions of a program to have smaller effects.

Takeaway

Just because a program worked in one setting does not mean it will be effective in another. Compare the local setting to the studied setting on features that are most likely to affect implementation and outcomes: if the local setting is quite different from the study context on features that matter, the program may not work the same way. Additionally, changing the program components or omitting aspects of the program may also alter program effectiveness.

8 Body of evidence

Consider the body of evidence when multiple studies exist.

In practice

A district reviews evidence on a tutoring program. One study shows positive effects in a setting where tutors received substantial training and coaching, the tutoring occurred multiple times per week, attendance was high, and the program replaced weak or inconsistent supports. A second study shows smaller effects in a setting with less training, fewer sessions, and lower participation. Looking across the studies is more informative than focusing on either one alone. One takeaway might be that the program appears more promising in settings with stronger staffing, training, and participation.

Even the best research can only tell us what happened in the places being studied. A single study in a single place can be informative, but it leaves open the question of whether the results depend on features of that setting, the comparison condition, or the supports used to deliver the program at that time. The full body of evidence from multiple contexts, however, can distinguish what is stable about a program from what is conditional on where, when, and how it is delivered. When results are positive across settings that differ in meaningful ways, we can have more confidence that the program has added value beyond a single local success.

Key terms

Generalizability, transferability, and external validity

Evaluators sometimes use terms such as generalizability, transferability, or external validity to describe whether results from one setting are likely to be useful in another. Findings have external validity if they generalize or transfer to other populations or settings. In this report, the question is whether the evidence base includes settings that resemble the setting of the users of the evidence, and whether results hold when conditions vary.Tipton and Olsen (2022) explain generalizability as the question of whether findings from a study sample can be extended to a defined target population. Here, we use these terms in a more practical sense to refer to whether evidence from one setting is likely to apply in another.

Questions for interpreting evidence

How many studies exist, and how much do the settings differ? Evidence from multiple districts, schools, or years is more informative than multiple studies from the same organization in the same setting. Look for variation in student populations, staffing models, schedules, and the comparison condition.

Are the results consistent in direction, even if the size varies? Some variation is expected because study estimates are imprecise and subject to noise. But differences in results may also reflect systematic differences across studies. Look for whether the results tend to point in the same direction across studies. A single null or slightly negative estimate is less concerning if the sample is small or the estimate is close to zero. More concern is warranted when a large, well-designed study shows clearly negative results that are hard to reconcile with positive findings elsewhere.

Are the studies using outcomes that are relevant and similar enough across studies? If one study uses a narrow measure closely aligned to a program and another uses a broader outcome, differences in effect size may reflect the outcome measures used rather than differences in program effects (see Principle 4: Interpreting effect sizes and Principle 5: Practical significance). All else equal, attend more closely to studies that speak to the intended use of the program in terms of the outcomes of interest and the timeframe over which changes in outcomes will be measured.

Is the comparison condition similar across studies, and similar to the local setting under consideration? A program can look more effective when compared to weak practice and look less effective when compared to strong practice. For each study, ask what the comparison group received, and ascertain which study looks the most similar to the treatment contrast in the local context (see Principle 7: Contextual considerations).

Do the studies describe implementation conditions in enough detail to learn from differences? If multiple studies report different impacts, that could be due to differences in implementation. Look for those details. For example, in the case of tutoring, determine the tutor qualifications, degree of training and coaching, group size, session length and frequency, participation rates, and how tutoring was scheduled within the school day. If such details are missing, it is difficult to know whether differences in results could be due to differences in delivery.

Takeaway

Look at the body of evidence for consistency in direction across multiple studies and attend to whether variation in results is explained by differences in implementation or study quality. If a program shows positive results across multiple contexts, it is more likely – but not guaranteed – to work in a new setting. If a program has variable results across contexts, check if study designs, contextual factors, or program implementation differences could explain the variation. In the best case, a multi-context evidence base can help identify what conditions appear necessary for success.

Principles 9–10

Evaluating More Complex Claims about Subgroups, Personalization, and Algorithms

Leaders may wish to know whether a program works better for some students than others, or whether the treatment can be effectively targeted to individual students in need. These are reasonable questions, but they are harder to answer than questions about overall average impacts.

9 Subgroup claims

View subgroup findings with caution unless they are planned, justified, and supported across settings.

In practice

A leader in a school district where many students are reading below grade level is considering adopting a new program. They have also found they are underserving multilingual learners and want to know if the program is particularly effective for that subgroup. If it is, the leaders plan to adopt as specified, and if not, they would look for additional supports for this population.

This is a decision about a subgroup. Subgroup analyses estimate a program’s impact within a defined group of students or examine whether program impacts differ across groups defined by gender, race/ethnicity, baseline achievement, disability status, multilingual learner status, or other characteristics. Subgroup findings can inform decisions in principle, but many studies do not have enough students in the sample to precisely estimate effects within many subgroups of interest. In addition, when researchers examine many subgroups, subgroup differences can appear by chance. The challenge is knowing when the evidence credibly indicates true differences in the effectiveness of the program.

Interpreting differences in the effect of a program across subgroups can also be tricky, even when they are estimated credibly. Found differences across subgroups (such as race or ethnicity, gender, disability status, multilingual learner status, or other identities) suggest groups benefit differently from treatment, but they do not explain why. For example, subgroups may react differently to treatment because the subgroups had unequal experiences, opportunities, access to supports, implementation, or other social and institutional conditions. Found subgroup differences should prompt examination of the conditions that may account for the pattern, not attribution of the pattern to characteristics inherent in the groups themselves.Racial and ethnic categories should not be treated as inherent explanations for differences between groups. Such differences should be interpreted in light of the historical, social, institutional, and implementation conditions that may be associated with these categories. See the American Psychological Association’s Journal Article Reporting Standards for Race, Ethnicity, and Culture and Gillborn, Warmington, and Demack (2018) on the treatment of racial categories in quantitative research.

Questions for interpreting evidence

Is the subgroup finding based on the right comparison? Subgroup findings are not simple comparisons between groups. A subgroup finding should estimate either the program’s impact within a subgroup or whether the program’s impact differs across subgroups. An impact finding within a subgroup requires comparing students in that subgroup who were offered the program with similar students in the same subgroup who were not. A finding that a program works better for one subgroup than another requires comparing the impact for one subgroup with the impact for a different subgroup, not comparing average outcomes across the groups. In the latter case, the finding could simply reflect baseline differences rather than differences caused by the program. See Principle 3: Comparison group.

Was the subgroup of interest identified before the results were known? Subgroup findings are more credible when researchers publicly state which subgroups will be examined before they conduct the analysis. This is the idea behind preregistration, which records the planned analyses in advance. This reduces the risk that a pattern found only after looking at the data will be treated as if it were expected all along. Since researchers might check many subgroups, preregistration can increase confidence that a reported finding is not just due to random chance.Looking at subgroups increases the possibility of false discoveries (findings due to chance or noise in the data rather than a true effect) which is one reason why subgroup analyses warrant caution (von Hippel & Schuetze, 2025).

How many subgroup comparisons were examined? Evaluation reports should state which subgroup analyses were conducted and how many comparisons were tested. The more comparisons a study examines, the higher the likelihood of false positives. Look for a statement that the study adjusted for multiple comparisons.

Was the subgroup tied to the program’s theory of action? A subgroup analysis should have a reason. For example, a phonics intervention may have a rationale for examining whether effects differ by students’ baseline decoding skills. Subgroup findings are more credible when they are supported by a clear rationale related to how the program works.

Does the subgroup finding appear in more than one site, cohort, or study? A subgroup finding is more credible when the pattern appears more than once. A finding that appears in one study but not others may reflect measurement error or sample composition rather than a real and stable difference in effects.

Could the subgroup finding reflect differences in implementation or access to program supports? One group may appear to benefit more because those students received more time, stronger delivery, or added support. For example, a reading program may appear to work better for students below grade level because those students in the treated group also received more small-group instruction, more time on the program, or closer monitoring from reading specialists. If effects are smaller for Latine students, leaders might examine whether implementation differed across the schools serving those students and, where relevant, whether the program provided adequate multilingual supports. The subgroup finding alone cannot establish which condition explains the difference.

Takeaway

Subgroup findings can help leaders think about fit, targeting, and service priorities, but they require more evidence than average impact findings. Treat a one-time subgroup finding as a possibility to examine, not as a firm basis for deciding who should receive a program. Findings are more credible when the subgroup was specified before results were known, the number of subgroup comparisons is limited, researchers can identify a rationale for expecting different results for a subgroup, and the same pattern appears across multiple sites, cohorts, or studies.

10 Personalization claims

Apply the same evidence standards to personalized and algorithmic tools.

In practice

A district is considering an adaptive math platform, described as a program that uses state-of-the-art AI-driven algorithms to customize instruction for individual students. Such labels describe what the tool is designed to do. They do not in themselves show that the tool helps students learn more.

To determine if the program works, the same evidence principles apply. The core question is whether using the tool improves outcomes compared with what it would replace. Evidence that a tool changes content, predicts risk, or recommends supports is evidence of the tool’s function, not its impact on student outcomes. A tool can sort students into groups, identify students as high risk, or recommend lessons without improving student learning. For leaders, the task is to look past the description of the tool and ask what decision, assignment, instruction, or support changed for students and educators as a result of using it, and, ultimately, whether those changes improved student learning.

Questions for interpreting evidence

Who or what produced the evidence? Ask whether the tool was tested with actual students and educators or only with simulated users, AI-generated responses, benchmark tests, synthetic data, or developer-created scenarios. Tests using simulations or AI-generated responses can show the tool performs a function, produces accurate answers, or behaves as intended under selected conditions, but such tests are not evidence that using the tool improves outcomes for students. Evidence of impact requires measuring outcomes for real students who experienced the tool and comparing those outcomes with what would have occurred under a credible alternative.

How was the tool used, and what was it compared to? First investigate how the tool was used: what changed for students or educators when the tool was introduced? Did students use the adaptive math platform during classroom instruction, or for homework? Did teachers use the tool’s recommendations to form groups or assign lessons? Then identify the comparison group and ask what students and educators in that group experienced. The estimated impact reflects the difference between those two conditions, not the effect of the algorithm in isolation.

What could go wrong for students or educators? Personalized and algorithmic tools can create errors or tradeoffs. A tutoring screener may miss students who need help or assign tutoring to students who do not need it. An adaptive platform may place students in content that is too easy, too hard, or disconnected from classroom instruction. A tool may also increase screen time, add teacher burden, or displace essential instruction. Ask whether the evidence examines these risks and whether the benefits are large enough to justify the tradeoffs.

Was the tool tested in a setting like the context in which the evidence is being used? A tool tested in one district may not work the same way elsewhere (see Principle 7: Contextual Considerations). Results may change when students, assessments, data systems, or staffing differ. Some schools may have less capacity to use the tool effectively. When the evidence comes from a setting unlike the district’s own, leaders should consider what factors make it more or less likely that they see similar results.

Takeaway

Personalization, adaptivity, and algorithms do not change the evidence standards. Evidence about use or process is not evidence of causal impact on student outcomes. The core elements of what evidence demonstrates effectiveness remain: ask what was tested, what the comparison is, whether student outcomes changed, what could go wrong, and whether the settings are comparable.

Closing

The ten principles reflect how we assess the strengths and limitations of evidence in our work. We hope they can help others make those judgments as well. In practice, the evidence available for many programs will be incomplete. Studies may describe implementation but not outcomes. Some lack details about what the program was compared to or how to deliver the program well.

And yet, leaders still must make decisions. These principles are meant to help identify what the evidence supports and what it does not support. Applying the principles may show that there is little evidence to support a program, strong evidence that it works across a variety of contexts, or something in between. In many cases, the evidence will be mixed, with some of the principles uncovering important concerns.

The strength of evidence can help guide what to do next. For a program that shows promise but leaves real questions, a small pilot may help leaders learn whether it can be implemented well before expanding. Even when evidence suggests that a program should work, actual results will depend in large part on how the program is delivered and what it replaces. Therefore, monitor the implementation! This might include collecting data on whether people are using the program as intended and what supports they need. Leaders can draw on the monitoring data to troubleshoot implementation challenges and identify where adaptations may be needed.

Monitoring involves attending to the evidence generated by the local leaders and staff. In that sense, monitoring shifts the work from using research to producing evidence. Monitoring can include examining outcomes or comparing outcomes in places that adopted the new program with outcomes in places that did not. Good monitoring designs can help identify which supports, versions, or delivery conditions are most important for improving student outcomes (Peck, 2019). As leaders monitor implementation and outcomes, they can use the same principles described above. Even basic questions can be useful at the start. Who is using the program? Do people find it useful? Where are the pain points? When leaders want to judge whether the program is improving student outcomes, the principles described above again apply. From start to finish, these principles can help leaders navigate the hard work of building something new.

In the end, the ten principles are habits of thought that help us evaluate the evidence of many programs, products, practices, and policies that we have seen offered as solutions to problems in education. We hope they are useful to you just as they have been useful to us.

Appendix A: Resources

General resources / introduction

Books featuring practitioner-facing guidance on finding, interpreting, and using education research include Common-Sense Evidence: The Education Leader’s Guide to Using Data and Research by Nora Gordon and Carrie Conaway and Educational Goods: Values, Evidence, and Decision-Making by Harry Brighouse, Helen F. Ladd, Susanna Loeb, and Adam Swift.

The principles presented here are consistent with the Standards for Excellence in Education Research from the Institute of Education Sciences, which set expectations for how studies are documented and communicated, including pre-registration, open materials and data, documentation of treatment implementation and contrast, cost analysis, high-quality outcome measures, and attention to generalization.

To search for causal evidence of program effects, see the Institute of Education Sciences’ What Works Clearinghouse or Johns Hopkins University’s Evidence for ESSA.

The IES-NSF Common Guidelines for Education Research and Development outlines six types of education research.

Related to the ASAP studies, MDRC has an Intervention Return on Investment (ROI) Tool for Community Colleges (mdrc.org/intervention-roi-tool). See also Minaya et al. (2026) and Weiss et al. (2019) in the reference list.

Initial read of the evidence  ·  Principles 1 & 2

IES’s Cost Analysis: A Starter Kit is a practical guide that provides a checklist, definitions, a template, and a three-phase approach to establishing cost estimates. The Starter Kit recommends identifying the full set of program “ingredients,” including personnel time, materials, facilities, training, and other supports; pricing those ingredients; and distinguishing start-up, recurring, total, and incremental costs. The guide also notes that resources already on hand, such as staff time or existing space, count as costs because they are being reallocated from other uses.

Conducting implementation research in impact studies of education interventions: A guide for researchers recommends specifying implementation research questions in advance; documenting how the intervention was delivered, in what context, and with what contrast; and collecting data that can be used to assess fidelity, adaptations, and variation in implementation across sites. Aimed at researchers, the guide could also provide a structure for district leaders to document integrity of implementation.

The What Works Clearinghouse (WWC) Procedures and Standards Handbook describes criteria WWC uses to review studies, assess design quality, and synthesize findings, and its reporting guides also emphasize clear, complete, and transparent reporting by study authors.

IES has practice guides in math, literacy, behavioral interventions, and more. Recommendations are derived from systematic reviews of research, and the strength of the evidence base is clearly indicated on the landing page of each practice guide. Practice guides also include examples and common obstacles and advice for navigating implementation challenges.

Other types of evidence summaries can be found in the Association for Education Finance & Policy’s Live Handbook, Annenberg’s EdResearch for Action, and the U.K.’s Education Endowment Foundation. In the tutoring space, the National Student Support Accelerator at Stanford provides research-based resources, including the District Playbook for High-Impact Tutoring, and Proven Tutoring at Johns Hopkins’ Center for Research and Reform in Education has an evidence clearinghouse for reading programs in grades K–8 and math programs in grades K–10.

What Teachers Need to Know About Education Research by Jackson, Ming, and Farley-Ripple (2025) provides questions to ask when considering whether to adopt a new program.

Evaluating evidence of causal effects  ·  Principles 3–5

Experimental and quasi-experimental designs for generalized causal inference by Shadish, Cook, & Campbell explains how randomized experiments, quasi-experiments, natural experiments, and nonexperimental designs attempt to approximate the missing counterfactual by creating or constructing a comparison, and it discusses the strengths and limits of these different approaches for supporting causal claims. The WWC standards provide a more applied framework for how these design ideas are used in practice to review education studies and evaluate the strength of causal evidence across different research designs.

Experimental Evaluation Design for Program Improvement by Laura Peck explores a variety of experimental evaluation design options and suggests opportunities for experiments to be applied in more varied settings.

Many studies described as random assignment or quasi-experimental do not demonstrate baseline equivalence (Jackson et al., 2024), leaving it unclear whether the groups are truly similar. This is one reason why Principle 2: Evidence transparency is critical: it allows others to investigate the extent to which the study design and analysis enables credible causal inference.

Effect Size Basics: Understanding the Strength of a Program’s Impact provides a brief primer. Interpreting Effect Sizes of Education Interventions by Kraft (2020) contains useful guidelines for interpreting effect sizes from causal research on education interventions and additional questions to ask. A related ASCD article features a conversation between Kraft and John Hattie.

In A Framework for Learning From Null Results, Jacob et al. (2019) argue that null findings should be examined in light of implementation, statistical power, outcome choice, and the size of effects the study was able to detect.

Average Effect Sizes in Developer-Commissioned and Independent Evaluations by Wolf et al. (2020) finds that program evaluations carried out or commissioned by developers produced average effect sizes that were substantially larger than those identified in evaluations conducted by independent parties. The authors provide evidence that interventions evaluated by developers were not simply more effective than those evaluated by independent parties.

Assessing the relevance of evidence  ·  Principles 6–8

The Joint Committee on Standards for Educational and Psychological Testing of the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education publishes the Standards for Educational and Psychological Testing, which has served as the field’s guidance on testing in the United States since 1966. The 1999 edition is available online.

EdInstruments is a database of common measures, including surveys, assessments, and observation protocols.

Wolf and Harbatkin (2023) demonstrate that reported effect sizes often differ by outcome measure type, including whether measures are narrow, researcher-developed, or closely aligned with the program.

References

  • American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
  • American Psychological Association. (2023). Journal article reporting standards for race, ethnicity, and culture.
  • Barshay, J. (2024, March 18). PROOF POINTS: Only a quarter of federally funded education innovations benefited students, report says. Hechinger Report.
  • Boulay, B., Goodson, B., Olsen, R., McCormick, R., Darrow, C., Frye, M., Gan, K., Harvill, H., & Sarna, M. (2018). The Investing in Innovation Fund: Summary of 67 Evaluations: Final Report (NCEE 2018-4013). Washington, DC: National Center for Education Evaluation and Regional Assistance, Institute of Education Sciences, U.S. Department of Education.
  • Brighouse, H., Ladd, H. F., Loeb, S., & Swift, A. (2018). Educational goods: Values, evidence, and decision-making. University of Chicago Press.
  • Gillborn, D., Warmington, P., & Demack, S. (2018). QuantCrit: education, policy, ‘Big Data’ and principles for a critical race theory of statistics. Race Ethnicity and Education, 21(2), 158–179. https://doi.org/10.1080/13613324.2017.1377417
  • Goodson, B. D., Harvill, E., Sarna, M., Brown, K., & McCormick, R. (2024, February). Federal efforts towards investing in innovation through the i3 Fund: A summary of grantmaking and evidence-building (NCEE 2024002). National Center for Education Evaluation and Regional Assistance, Institute of Education Sciences, U.S. Department of Education.
  • Gordon, N., & Conaway, C. (2020). Common-sense evidence: The education leader’s guide to using data and research. Harvard Education Press.
  • Hill, C. J., Scher, L., Haimson, J., & Granito, K. (2023). Conducting implementation research in impact studies of education interventions: A guide for researchers (NCEE 2023-005). Washington, DC: U.S. Department of Education, Institute of Education Sciences, National Center for Education Evaluation and Regional Assistance. Retrieved from http://ies.ed.gov/ncee.
  • Institute of Education Sciences. (2020). Cost Analysis: A Starter Kit. Version 1.0. (IES 2020-001). U.S. Department of Education. Washington, DC: Institute of Education Sciences. Retrieved from https://ies.ed.gov/pubsearch/
  • Jackson, C. (2025, June 10). Who Cares About Effect Sizes? Education Week. Retrieved from edweek.org
  • Jackson, C., Ming, N., & Farley-Ripple, E. (2025, March 4). What Teachers Should Know About Education Research. Education Week. Retrieved from edweek.org
  • Jackson, C., Wilson, S. J., & Glenn, M. (2024). How Does the Prevention Services Clearinghouse Rate the Design and Execution of Studies?, Handbook of Standards and Procedures, Version 2.0, OPRE Report 2024-339, Washington, DC: Office of Planning, Research, and Evaluation, Administration for Children and Families, U.S. Department of Health and Human Services.
  • Jacob, R. T., Doolittle, F., Kemple, J., & Somers, M.-A. (2019). A framework for learning from null results. Educational Researcher, 48(9). https://doi.org/10.3102/0013189X19891955
  • Kraft, M. A. (2020). Interpreting Effect Sizes of Education Interventions. Educational Researcher, 49(4), 241–253.
  • Minaya, V., Scott-Clayton, J., & Thomas, J. K. R. (2026). The Returns to Degree Completion at CUNY’s Community Colleges. Community College Research Center.
  • Pane, J. F., & Baird, M. D. (2019, July 3). Did One Program Really Cost Students 276 Years of Learning? Education Week.
  • Peck, L. R. (2020). Experimental evaluation design for program improvement. SAGE Publications.
  • REL West. (2021). Effect Size Basics: Understanding the Strength of a Program’s Impact.
  • Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.
  • Soland, J. (2025, December 8). A more expansive approach to studying what works in education. Brookings Institution.
  • Tipton, E., & Olsen, R. B. (2022). Enhancing the Generalizability of Impact Studies in Education (NCEE 2022-003). Washington, DC: U.S. Department of Education, Institute of Education Sciences, National Center for Education Evaluation and Regional Assistance.
  • von Hippel, P. T., & Schuetze, B. A. (2025). How Not to Fool Ourselves About Heterogeneity of Treatment Effects. Advances in Methods and Practices in Psychological Science, 8(2). https://doi.org/10.1177/25152459241304347
  • Weiss, M. J., Ratledge, A., Sommo, C., & Gupta, H. (2019). Supporting Community College Students from Start to Degree Completion: Long-Term Evidence from a Randomized Trial of CUNY’s ASAP. American Economic Journal: Applied Economics, 11(3), 253–297.
  • What Works Clearinghouse. (n.d.). Standards Brief for Baseline Equivalence.
  • What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0. U.S. Department of Education, Institute of Education Sciences, National Center for Education Evaluation and Regional Assistance (NCEE).
  • Wolf, B., & Harbatkin, E. (2023). Making sense of effect sizes: Systematic differences in intervention effect sizes by outcome measure type. Journal of Research on Educational Effectiveness, 16(1), 134–161. https://doi.org/10.1080/19345747.2022.2071364
  • Wolf, B., & Klager, C. (2026). Subgroup Effects, Compositional Effects, and Treatment Effect Heterogeneity in the What Works Clearinghouse Study Data. Journal of Research on Educational Effectiveness, 1–32. https://doi.org/10.1080/19345747.2026.2638432
  • Wolf, R., Morrison, J., Inns, A., Slavin, R., & Risman, K. (2020). Average effect sizes in developer-commissioned and independent evaluations. Journal of Research on Educational Effectiveness, 13(2), 428–447. https://doi.org/10.1080/19345747.2020.1726537
  • Wong, V. C., Tipton, E., & Spybrook, J. (2025, April 10). How federal investments in education research help students succeed. Brookings Institution.

The Future of Evidence in Education Network · Sidenotes appear in the right margin on wide screens; on narrow screens, tap a numbered reference to open its note.