Why Tutoring at Scale Doesn’t Always Mean Impact at Scale

After COVID-19, there was a rare consensus in education policy: tutoring would be the key to recovery. Researchers, policymakers, and even tech philanthropists rallied behind this. The evidence seemed solid. Meta-analyses showed impressive gains, often the equivalent of half a year of learning. But what if those average effects gave us the wrong idea about what really works in real schools?

A recent, ambitious working paper by Matthew Kraft, Beth Schueler, and Grace Falken I found via Best Evidence in Brief, takes a hard look at that assumption. Rather than simply pooling results from dozens of studies, they ask a sharper question: What should we realistically expect from tutoring programs when rolled out at scale in public schools, aiming to improve test scores?

Their answer is both sobering and incredibly useful.

They analysed 265 randomised controlled trials—three times more than the largest previous meta-analysis—and found that average effects shrink considerably when the tutoring programs look more like those currently implemented in U.S. schools. Once you limit the sample to large-scale programs (1,000+ students), based in the U.S., and evaluated with third-party standardised tests, the average effect drops to just one-third or even one-quarter of the full-sample effect. In numbers: from 0.42 standard deviations in the complete set to 0.14–0.21 in the more realistic subset.

Still, these are not meaningless effects—especially at scale. But they do call for a reset of expectations.

So why do tutoring effects decline as programs grow?

The authors explore four explanations:

  1. Publication bias? Probably not. The patterns don’t suggest that small studies are only published when they show large results.
  2. Program design? Definitely. Larger programs tend to use larger group sizes and offer less tutoring time.
  3. Different student populations? Likely. Small-scale programs often focus on students who stand to gain the most—larger programs serve a broader, more diverse group.
  4. Implementation quality? Almost certainly. It’s harder to maintain high quality when scaling up quickly and across schools.

There is one piece of good news: programs that stick to a “bundle” of high-impact design features—such as three sessions per week, small group sizes, and in-person delivery during school hours—tend to hold up better when scaled. The decline in effectiveness is still there, but it is smaller.

What about cost-saving adaptations? The findings are mixed. Online tutoring generally shows weaker effects, but the evidence base is still emerging. Higher student–to–tutor ratios usually mean smaller gains, but not zero. Peer tutoring seems surprisingly promising—possibly a low-cost way to scale, but more comparative research is needed.

The key takeaway? It’s tempting to scale up a program that worked beautifully with 50 students and expect similar results. But education doesn’t scale like that. The more we try to replicate small successes across entire systems, the more we must ask: what actually travels well? What breaks in the process?

This study isn’t an argument against tutoring. It’s a call for smarter scaling, designing with reality in mind, and being more precise when we talk about “evidence-based” policy. If we want tutoring to help millions of students, we need to move beyond the moonshot narrative and build something that works, even at ground level.

Abstract of the working paper:

U.S. public schools are engaged in an unprecedented effort to expand tutoring in the wake of the COVID-19 pandemic. Broad-based support for scaling tutoring emerged, in part, because of the large effects on student
achievement found in prior meta-analyses. We conduct an expanded meta-analysis of 265 randomized
controlled trials and explore how estimates change when we better align our sample with a policy-relevant
target of inference: large-scale tutoring programs in the U.S. aiming to improve standardized test performance.
Pooled effect sizes from studies with stronger target-equivalence remain meaningful but are only a third to a
half as large as those from our full sample. This result is driven by stark declines in pooled effect sizes as
program scale increases. We explore four hypotheses for this pattern and document how a bundled package of
recommended design features serves to partially inoculate programs from these attenuated effects at scale.

Leave a Reply