Gamification shows once again that the question ‘what works?’ is too simple.

This research has been sitting in an open tab in my browser for a while. I read it back when the post about the umbrella review of fifteen meta-analyses on project-based learning was published, but I didn’t want to write a similar post two days in a row. At the time, the conclusion was quite painful. While almost all meta-analyses reported positive effects, they were of such low methodological quality that we could draw few reliable conclusions from them. This meta-meta-analysis by Kanadli and Sancar-Tokmak on gamification provides a nice addition to that story. Not because gamification and project-based learning are the same thing. Nor does this new research encounter the same problems. But because it shows once again how misleading a seemingly simple question can be: does it work?

By now, we can no longer call gamification a new idea. In short, it means using elements from games in a context that is not a game itself. Think, for example, of points, badges, levels, leaderboards, challenges, or other game mechanics in a learning environment. It is often forgotten that these elements originate in educational research or theory. Think of the token economy, scaffolding, or the zone of proximal development.

And by now, quite a bit of research has also been done into the concept. So much, in fact, that not only do meta-analyses exist, but now there is also a new meta-analysis of existing meta-analyses. The researchers found 28 (!) meta-analyses on gamification, of which 10 ultimately met their inclusion criteria.

Neither researcher simply took the average effects from those ten meta-analyses and then recalculated an average. Instead, they return to the effect sizes of the original, primary studies included in these meta-analyses and re-analyze them within a single statistical framework. Where necessary, they even recalculate the effect sizes.

Thus, for academic performance, they report 105 effect sizes from 97 studies. For motivation and engagement, it concerns 42 effect sizes from the same number of studies.

So: does gamification work?

If you look only at the average calculated in this way, the answer seems quite unambiguous: yes. For academic performance, the researchers initially find a more-than-respectable average effect size of g = 0.559. For motivation and engagement, they find almost the same: g = 0.527.

But is that the end of the story? No, not really.

To start with, the differences between the studies are enormous. For academic performance, for instance, the researchers’ effect sizes range from -1.613 to +4.035. The prediction interval, which indicates the range of true effects we might expect across different contexts, ranges from -0.491 to +1.609.

These differences are pretty huge. Averages such as these may suggest that gamification works, but they obscure cases where it works very well, cases where it doesn’t matter much, and cases where it could be significantly detrimental to learning.

The researchers therefore also examined outliers. When they remove nine outliers for academic performance, the average effect drops to g = 0.503. This still does not resolve the heterogeneity, which remains high at I² = 76%. The prediction interval also still ranges from -0.289 to +1.294.

We see something similar for motivation and engagement. After removing three outliers, the average effect drops from g = 0.527 to g = 0.398. Here, heterogeneity is I² = 77%, with a prediction interval ranging from -0.169 to +0.965.

As a side note, removing outliers requires some caution. We cannot simply assume that extreme effects are wrong. In an additional analysis, the researchers identified extreme values that they retained because they were plausible and likely not the result of a data error. Personally, I would therefore not attach too much importance to the precise difference between, say, 0.559 and 0.503. I find the enormous variation between studies much more interesting.

And then there is publication bias.

Regarding academic performance, the researchers find evidence that positive results are likely overrepresented in the published literature. When they try to correct for this, the estimated effect for academic performance drops to g = 0.316, with a prediction interval from -0.778 to +1.410. For motivation and engagement, the estimate, adjusted for potential publication bias, ultimately drops to g = 0.183. The effect does not disappear, but it does become considerably smaller.

But there is a need for some caution. The statistical method the researchers use is not some magical machine that tells us what the ‘real’ effect is. The authors themselves explicitly describe the adjusted estimate as exploratory.

On average, there appears to be a positive effect. However, the magnitude of that effect varies greatly. And in some specific circumstances, the effect can even be negative. Such an average effect does not tell us exactly which circumstances those are.

But how good are all those meta-analyses actually?

Then we come back to the umbrella review from earlier. In project-based learning, this was exactly the big surprise/disappointment that got me blogging. There were fifteen meta-analyses, but a formal assessment of their methodological quality showed that they all had to be rated as critically low.

Therefore, I feel one important step is missing in this new study on gamification. As far as I can see, Kanadli and Sancar-Tokmak do not perform a formal quality assessment of the ten meta-analyses on which their new analysis is based, for example, using AMSTAR 2 or ROBIS. That does not mean that they simply accept the results of those meta-analyses as true. On the contrary. Their method is quite interesting in this regard. As I wrote earlier, they go back to the effect sizes from the primary studies. Furthermore, they attempt to resolve dependencies between effect sizes. The researchers also conduct sensitivity analyses and recalculate effect sizes when they identify inconsistencies.

However, this does not tell us how reliable the underlying research basis is. After all, the quality of a meta-analysis is about more than just calculating the final effect size. In fact, the researchers circumvent, or rather overcome, this question through their approach.

Perhaps we simply ask a question that is too simple too often (bis)

The question ‘does gamification work?’ is not wrong. And the answer, based on this study, is not that we know absolutely nothing, either. On the contrary, the best short answer seems to me to be, on average, probably a bit (sexy answer, right?). But as soon as you know this, the interesting questions actually just begin. Which form of gamification? What purpose does it need to serve? Which students do you want to use it for? And in what context? How long will you be using it? Which game elements will you use? What will the teacher be doing? And how good is the research on which we base our answer?

The researchers ultimately reach a similar conclusion: the effectiveness of gamification seems to depend heavily on context and implementation. No shit, Sherlock.

Leave a Reply