A previous summer saved me from making a mistake on this blog. I came across a meta-analysis on AI in education and bookmarked it to write about later. By the time I returned to it, however, the study had started to fall apart. Although it had appeared in a top journal, many of the underlying studies – and even some of the journals in which they had been published – turned out to be of rather questionable quality (and that is putting it mildly).
It became a recurring theme throughout the past school year. Journals publish an enormous amount of research on AI in education. But much of these studies leaves a lot to be desired. A Stanford report reached the same conclusion. And the low point came when a journal had to retract a widely discussed meta-analysis.
This summer, a new meta-analysis appeared. Is this one better, even if the underlying research often still isn’t?
Let me be clear from the outset: this meta-analysis by Gülsen Gençdal and Nazım Çoğaltay is not in the same league as the retracted one. Methodologically, it is a solid piece of work. The authors conducted a systematic literature search, used clear inclusion criteria, limited themselves to experimental and quasi-experimental studies, and carried out the usual checks for heterogeneity and publication bias.
The problem, however, remains the same: garbage in, garbage out. A meta-analysis does not become strong simply because it is conducted correctly. It can never be stronger than the studies it includes.
And that, in my view, is still where this study struggles. The authors conclude that AI has a large positive effect on academic achievement (d = 0.80). That sounds impressive. Perhaps even too impressive. Average effects of that magnitude are unusual in educational research. This is especially the case in a field that is still so young and where many studies rely on small samples, short interventions and researcher-developed outcome measures. We have seen this pattern before. The early enthusiasm surrounding growth mindset also came from similarly small studies reporting remarkably large effects.
There is a second, perhaps even more surprising, problem. The study does not really examine whether AI works. Instead, it combines a huge variety of very different interventions under a single label: intelligent tutoring systems, adaptive learning platforms, machine learning, deep learning and generative AI are all treated as if they belong to the same category. From a pedagogical perspective, however, they often have very little in common.
It is a bit like studying the effects of “transport” by analysing bicycles, cars, trains and aeroplanes together. There are some rather obvious differences in distance, emissions and comfort.
As a result, the average effect size tells us surprisingly little. Schools are not asking whether “AI” works in general. They want to know whether a particular AI application, for a particular learning goal, with a particular group of students, adds value compared to a good alternative.
So yes, this meta-analysis is certainly better than some of its predecessors. But it does little to change the conclusion I reached over the past school year. What we need today is not more AI research, but better AI research. Only then – perhaps after a few more years – will meta-analyses be able to provide reliable answers to the questions schools are actually asking.