In the scientist’s mind

When you read a scientific article, research can sometimes seem like a remarkably emotionless process. First, the researchers formulate a hypothesis. Next, they conduct an experiment. Then they analyse the data and draw a conclusion. Anyone who does research themselves knows, of course, that it rarely feels that way. On the contrary, it is a process of trial and error, of eureka moments, and of quite a bit of cursing, not always under your breath. And that is without even mentioning the ordeal that publishing can be (damn you, second reviewer!).

A new article by Jonna Brenninkmeijer and colleagues offers a rather exceptional glimpse into this process. The researchers closely followed a team of psychologists attempting to replicate an influential experiment from the 1980s. You can find that study here, by the way. They observed meetings, analysed documents, and followed the project from its preparation to the eventual publication. The result is almost a glimpse inside the mind (and heart) of a scientist.

The original experiment they wanted to replicate was quite small by today’s standards. Thirty-two four-year-old children participated, eight per condition. One person gave the instructions, and one person coded what the children did. One of the four instructional methods (scaffolding) produced clearly better results, and the underlying educational strategy subsequently became widely known. This new team of researchers wanted to repeat the study as accurately as possible. But this time with more than 250 children, four people giving the instructions, and five coders to ensure substantial statistical power.

And then the problems begin.

To start with, it turns out that a lot of information about the original experiment has been lost. There is no original protocol left, no recording, and even the original puzzle has disappeared. The researchers therefore have to reconstruct exactly what happened at the time. Replication sounds simple, but despite the enormous influence the original research had on education, the available information proves surprisingly scarce.

But they soon run into many more issues. Take something as simple as whether or not a parent is present. A four-year-old reacts differently when mum or dad is in the room. But the researcher also behaves differently when a parent is watching. So what do you do? Do you allow the parent to be present because that was probably also the case in the original experiment? Or do you put the parent behind an observation window because that gives you a better-controlled experiment?

When the researchers actually start working with the children, reality sets in. They try to establish very precisely when and how a researcher is allowed to help a child. But four-year-olds have the annoying habit of not always behaving according to a research protocol. Sometimes a child does nothing. At other times, they explicitly ask for help. Sometimes they simply start playing with the puzzle pieces. As a researcher, you want to react like a human being, but does the protocol allow you to?

At one point, one of the researchers notes that some of the children’s responses feel completely natural. But, according to the experimental protocol, they should be counted as mistakes. This becomes a constant choice between sticking to the protocol and having a warm, supportive interaction with the flesh-and-blood child sitting live in front of you. All of this makes it an emotionally intense process.

For example, the researchers discovered that the person who gave all the instructions in the original experiment was the original researcher’s wife. Moreover, she had been intensively trained. A researcher who was later trained by them recalls how the couple literally looked over her shoulder during her training and constantly corrected her. One of the researchers conducting the replication reacts particularly strongly to this. She realises that something was apparently passed on that was never described in the original scientific publications. She has protocols, decision rules, and coding schemes, but no one is standing over her shoulder to teach her what the original researchers apparently took for granted. So is this no longer a true replication?

Coding the videos also turns out to be much more difficult than expected. What, for example, constitutes an action by a child? If a child picks up a puzzle piece and puts it somewhere else, does that count as a correct action, an incorrect action, or not as a meaningful action at all? And what counts as help? Putting a puzzle piece in the right place is clearly help. But what about laying out pieces? Giving a compliment? Smiling in encouragement?

Five coders have to assess such behaviours in the same way. Initially, they struggle to reach sufficient agreement. One of the researchers says that she worked extremely hard on the project for twenty months. She worries about her career. She does not want to make any more sacrifices. At one point, she even says that she never wants to do a project like this again. At another, empty cells in the dataset keep her awake at night.

All of this will feel genuinely familiar to anyone involved in research, but at the same time, these are sentences you rarely read in the final methods section of a scientific paper. And the interesting thing is that these emotions are not merely noise that the rational scientist needs to suppress. Sometimes they tell the researchers that something is amiss in what they are doing.

Why can’t we reliably align our coders? Why can’t we implement this protocol naturally? Or why can’t we reconstruct exactly what the original researchers did? And if a team of expert researchers struggles to implement a teaching strategy according to the protocol after years of preparation, what does that mean for a teacher who has to use the same strategy in a regular classroom? This last question is crucial for evidence-informed practice, by the way.

Eventually, one important part does work out. The researchers calculate how closely their instructors followed the prescribed instructions and arrive at approximately 70 per cent. That is roughly the same percentage as in the original study. This leads to a sense of relief. This is until the results come in.

Nothing.

Or better, close to nothing: the four different instructional methods produce approximately the same results. The strong effect observed in the original study does not recur. The replication has failed, or rather, the original finding does not replicate.

That, too, comes with emotions. The researchers had already considered the possibility that they would not find the original effect. But apparently, they had not expected to find any difference at all between the conditions. And then they have to face the most difficult question: what does such a failed replication mean?

The answer could simply be that the original study was flawed. Remember, it was an experiment with only eight children per condition, one instructor, and one coder. The new study was much larger and did not find the effect. But the researchers also begin to question something more fundamental.

Perhaps the essence of this educational strategy does not lie in a collection of steps that can be neatly captured in a protocol. Perhaps part of the original success lay in the interaction between that particular researcher and those particular children. Think of experience, timing, tacit knowledge, and coaching.

The authors of this new study draw some fairly substantial philosophical conclusions from this. According to them, the case shows how method, emotions, and even our ideas about what a psychological phenomenon actually is can become intertwined. During the study, the researchers conducting the replication themselves changed their minds about what exactly they were trying to replicate.

I would be more cautious here. This remains a single, albeit very well-developed, case study. I do think there is also a much simpler possible explanation: perhaps the original, very small study was simply wrong, or the observed effect was greatly overestimated. We do not need to revise our entire view of psychological phenomena to explain that.

In recent years, much discussion has centered on how to make science more reliable. This led to the introduction of preregistration, larger sample sizes, and the promotion of open data, replications, and transparency. Regular readers of this blog know that I consider all of this more than justified. But sometimes it can feel as if these measures could somehow remove the human being from science.

Of course, they cannot. Scientists still have expectations and doubts. They should, btw. Researchers are also loyal to colleagues and theories. They can be afraid of making mistakes. And they feel relieved when something works and disappointed when months of work seem to come to nothing. Sometimes they lie awake worrying about their dataset.

I don’t think good science requires any of those things to disappear, although some scientists would love to sleep better. Good science tries to ensure that knowledge remains reliable despite all of them.

Leave a Reply