Let’s start with the obvious: generative AI is rapidly entering classrooms, whether we like it or not. Recent surveys show that many teachers are already using AI, often to prepare lessons, create classroom activities, or generate tests. That makes sense, and an earlier Education Endowment Foundation (EEF) evaluation suggested that this does not necessarily come at the expense of quality.
A new randomised study by researchers at the University of Pennsylvania adds an interesting nuance. The authors are not arguing against AI. Rather, they suggest that benefits for teachers do not automatically become benefits for students.
One of the strongest study designs so far
I’ve criticised the quality of many AI studies before, but this one appears considerably stronger. The researchers conducted a randomised field experiment across 14 schools in Turkey. Nearly 200 teachers were randomly assigned to either an AI group or a control group. The AI group received access to a GPT-4o-based tool specifically designed to support teachers with lesson planning, differentiation, assessment and feedback.
The researchers then looked not only at how teachers used AI, but, more importantly, at what students noticed.
What did they find? Looking first at the teachers, around two-thirds of all AI conversations were about producing teaching materials:
- preparing lessons;
- creating exercises;
- writing exams;
- developing syllabi.
Teachers were much less likely to use AI for differentiation, feedback, or thinking through pedagogical decisions. The limited use of feedback surprised me a little; the rest did not.
Another striking finding was how brief the interactions were. The median conversation consisted of just two prompts. This suggests that AI was mainly being used to generate a finished product quickly rather than to develop ideas collaboratively or critically refine them. In other words, tha AI use was fairly basic.
The most striking findings
In line with the earlier EEF evaluation, average examination results did not decline, although scores were already remarkably high in both groups.
Student motivation, however, did.
Students whose teachers used AI rated their subject as:
- less interesting;
- less important;
- less enjoyable.
The effect is certainly not dramatic, but at around 0.11 standard deviations it is large enough to take seriously.
The differences between teachers were equally interesting. Among teachers who had been lower-performing before the study, students also showed lower examination scores and lower confidence in their own abilities. This pattern did not appear among higher-performing teachers.
The researchers also found that teachers who had already been frequent AI users before the experiment became more sceptical about AI afterwards. Less experienced users, in contrast, became more optimistic. The study therefore suggests that experienced users may also become more aware of AI’s limitations. That is certainly something I recognise myself.
Taken together, these findings suggest that AI is not necessarily the great equaliser some hope it might be. If anything, AI seems to become problematic when it replaces pedagogical thinking rather than supports it.
Why might this happen?
The authors propose two possible explanations.
The first is that AI may gradually displace the teacher’s personal voice. Good lessons are not simply accurate explanations or well-designed exercises. They also reflect the teacher’s style, humour, examples and personality. When multiple teachers rely on the same AI tools, some of that individuality may be lost. The researchers suspect that students notice this sooner than we might expect.
The second explanation strikes me as even more relevant. AI may encourage some teachers to think less deeply about their lessons because it takes over part of the cognitive work. Yet this process of thinking through lesson structure, generating examples and anticipating misconceptions is a crucial part of teachers’ professional development. Teachers who outsource too much of that process simply spend less time developing their own pedagogical expertise.
Does this mean AI is bad?
No. And the authors, Alp Sungu, Benjamin Lira and Angela Duckworth, are careful to say exactly that.
Their study does not suggest that AI has no place in education. It does, however, remind us not to automatically assume that saving teachers time will translate into better outcomes for students. These are simply two different questions.
AI can reduce administrative work, provide inspiration and help with routine tasks. But when it starts to replace pedagogical thinking rather than support it, genuine risks may emerge.
One important caveat
As always, this study has its limitations. The experiment lasted only one semester, took place in Turkey, and evaluated just one particular AI system. It is therefore entirely possible that different AI tools, different ways of using them, or longer implementation periods would produce different results.