This blog has an AI problem

This blog has an AI problem. No, not because it was written by AI (although I translated it with AI’s help). Or rather, I have an AI research problem. The problem is that I am increasingly unsure which new studies on AI and education are still worth discussing here.

I have a monthly podcast in Dutch on educational research, and my co-host, Rinke, has already commented on it a few times, sighing, ‘Yet another study on AI.’ Obviously, it has been difficult to ignore artificial intelligence in recent years. And since ChatGPT became available to the general public at the end of 2022, there has been a real explosion of research into what generative AI could mean for learning and education. And by explosion, I really do mean an explosion.

Stanford has a repository collecting research on AI and K-12 education. When researchers took stock in October 2025, it contained 818 papers. A few months later, there were already more than 1,100. By now, we have moved quite a bit further, and the number has passed 1,500 studies. At first sight, that sounds fantastic. A new technology appears and it may have major consequences for education. And the researchers? They throw themselves at it.

But I increasingly notice a different feeling when I read research for my work or for this blog and come across yet another study on AI and education: haven’t we studied this quite a lot already?

A lot of research, but how much are we learning?

A first problem is that much of the research seems to tell us essentially the same thing over and over again. Let me summarise:

  • Pupils and students can complete certain tasks faster with AI.
  • AI can help with feedback. Students appreciate the convenience.
  • If AI takes over the thinking, there is a risk that people will think less for themselves.
  • If, on the other hand, AI is used to support one’s own thinking, there can be positive effects.
  • As with so much in education, the precise results depend strongly on what the learner does, what the AI does, what task is being performed and how the whole thing is embedded in teaching.

These are important findings concerning pupils and students. There is also research into teachers’ use of AI, which can save time and can have a negative effect on trust between teachers and students.

As a strong advocate of replication research, I certainly don’t want to argue for less replication now. Quite the opposite. We need more replications across different subjects, ages, countries, and circumstances. And with technology changing so quickly, a study using a new model may indeed yield different results from one using an older version.

But… there is a difference between replicating research to make our knowledge more robust and adding yet another study because AI happens to be a topic that rather easily attracts attention. The number of publications is therefore not directly proportional to the amount of knowledge we gain.

And then there is the quality

There is something else striking about that Stanford analysis. Of the 818 papers in their repository at the time, only 20 ultimately met the criteria for studies that allowed sufficiently strong causal claims.

That certainly does not mean that the other 798 studies are worthless. A qualitative study can be extremely valuable without investigating causality. Descriptive research can tell us how pupils and teachers use AI. Survey research can reveal interesting patterns. But the difference between more than 800 publications and 20 studies with strong causal evidence does put the size of the research field somewhat into perspective.

And anyone who reads a lot of research on AI soon encounters the problems I have mentioned before: small samples, short interventions, self-reporting, researcher-made tests, insufficiently strong control groups and sometimes rather large conclusions based on rather limited data. Those large conclusions then tend to be popular on LinkedIn.

But there is another, rather peculiar problem. Research on AI sometimes has an AI problem of its own. In 2024, for example, a review on ChatGPT and the changing role of teachers was published in Heliyon. The article was later retracted. The post-publication investigation found, among other things, irrelevant references and large sections of text without proper attribution. There were also concerns about the undisclosed use of generative AI in writing the article. The author disputed the reasons for the retraction. But this is certainly not the only example.

We already have a lot of meta-analyses

Generative AI, as we now know it, has only been widely used since the end of 2022. The research field is therefore only a few years old. Yet it has already become difficult to keep track of the meta-analyses on generative AI and education.

One meta-analysis from 2026, for example, combines 35 experimental studies on ChatGPT and learning outcomes. Another analysis of generative AI in higher education combines 57 studies and 97 effect estimates. Yet another, specifically looking at undergraduate students and ChatGPT, found 66 studies published between early 2023 and May 2025. And by now there are separate meta-analyses on, for example, AI feedback and even on generative AI in medical education.

Again, to be clear: conducting a meta-analysis in a young research field is not wrong by definition. If there are enough comparable studies, bringing them together can be extremely useful. But there is a characteristic of meta-analyses that we often overlook. A meta-analysis cannot make the underlying research better than it is. Garbage in, garbage out, as a fellow researcher once described it to me.

If you combine dozens of small, short or methodologically weak studies, you do not suddenly get strong evidence automatically. And when different meta-analyses repeatedly use largely the same primary studies, we do not suddenly have several independent streams of evidence either.

Even a meta-analysis can disappear again

There is an example that illustrates this almost too perfectly. In May 2025, a meta-analysis appeared with the promising title The effect of ChatGPT on students’ learning performance, learning perception, and higher-order thinking. The article received a great deal of attention and was cited hundreds of times. In April 2026, the meta-analysis was retracted. Again, this does not mean that meta-analyses on AI are unreliable. That would be precisely the same kind of overly broad conclusion I am warning against here.

But it does show why we need to be careful with the idea that the words meta-analysis automatically mean that we now have a definitive answer. In fact, this week another interesting systematic review and meta-analysis appeared, this time on generative AI in STEM education. The researchers found 85 eligible studies, 49 of which could be included in the meta-analysis. At first sight, the average effect of generative AI on cognitive learning outcomes was positive. But you had better look at the analysis a little longer.

The differences between the studies were enormous: I² = 96.32 per cent. The prediction interval ranged from g = -1.52 to g = 3.20. In other words, what might be expected in a new setting could range from a fairly negative to a very large positive effect. The researchers also found indications of publication bias. When they tried to correct for this in an additional analysis, the overall positive effect could even largely be explained by that publication bias. Well, what do we know then?

There is another problem: technology doesn’t wait

And then there is something educational research struggles with anyway: it takes (a lot of) time. You have to design the study, find participants, carry out the intervention, clean and then analyse the data, write an article, convince reviewers and then wait (a long time) for it to appear.

But meanwhile, the AI company behind the application you studied hasn’t been sitting still, and yet another new version has been released. A study using, for example, an LLM from 2023 can be methodologically excellent and still raise the question of how well the results apply to systems from 2026. Models change, interfaces change and, above all, the way people interact with them changes.

All this creates an annoying paradox: we want fast research because the technology changes quickly. But fast research increases the risk of weaker research. By the time we have strong, long-term research, there is a risk that the technology has changed again.

I genuinely don’t see a simple solution to this.

So perhaps what we need most isn’t more research

I certainly don’t want to argue that we should stop researching AI and education. Quite the opposite, although I genuinely notice that I am starting to avoid it myself. At the same time, the technology is too important, and the possible consequences for learning and education are too great to ignore.

But what I wanted to point out in this blog is that, with today’s approach, we sometimes learn very little. A thousand papers do not mean that we know a thousand times more. Ten meta-analyses do not mean that we have ten independent confirmations. And an impressive effect size tells us little if we don’t know which studies lie underneath it, what exactly was measured, and what the AI condition was compared with.

What I increasingly hope to see are fewer quick variations on the same question and more solid longitudinal research, stronger experiments, good replications, realistic educational settings and, above all, research that builds on what we already know. But yes, that takes time, and that in turn has other disadvantages, as I just described.

I asked AI for a solution to this problem, but I didn’t get a satisfactory answer there either.

Leave a Reply