Let’s be honest, young people are using AI chatbots en masse. Via Axios, I found this new assessment by the US organisation Common Sense Media. They discovered that 67 per cent of children aged 9 to 17 years old in the USA have used AI chatbots. 40 percent name ChatGPT as one of such tools. This American organization asked itself what would happen if instead of asking the bot to summarize the lesson of history, a kid asks to complete all of his homework. Or, even more frighteningly, what would happen if he would discuss his self-injury and eating disorders with this chatbot?
I haven’t seen this one myself, but OpenAI has created more safeguards for teens back in August. They have something called ChatGPT for Teens, but it’s not a distinct bot. Rather, it’s an accumulation of settings and filters for teenagers’ accounts or accounts whose user is believed to be between 13 and 17.These measures include Study Mode (which I have discussed before), parental controls, Study Hours, break reminders and extra protection around sensitive topics.
The tests were conducted by Common Sense Media’s Youth AI Safety Institute. In case you don’t know who Common Sense Media is, it’s not a US government regulatory body. It’s an independent nonprofit organization that, amongst others, conducts reviews of digital products for kids and teens. To conduct their tests in this particular report, they fed more than 4,000 prompts to ChatGPT before and after the implementation of new safeguards for teens.
I think their conclusion is pretty damning for OpenAI. Based on their assessment framework, ChatGPT for minors receives the verdict “Unacceptable Risk”. The organisation even calls on OpenAI to stop promoting ChatGPT to minors for the time being, until the problems have been addressed. Why is this the case?
Study Mode works, but only if you want to study
Let’s first look at the tool ChatGPT developed specifically for education. The idea behind Study Mode is rather appealing. Instead of immediately giving an answer, ChatGPT is supposed to help students arrive at a solution themselves through hints, questions and intermediate steps.
At first sight, the results seem positive. With an account of a 13-year-old with Study Mode enabled, the percentage of assignments that ChatGPT completed in full dropped from 70 per cent before the new measures to 18 per cent afterwards. But… the story quickly turned pretty “weird”. When Study Mode was switched off, ChatGPT completed 100 per cent of the assignments after the update.
With the tested account of a 17-year-old with Study Mode enabled, full completion even increased from 0 to 88 per cent. The weak spot turned out to be a particular option that students were regularly offered: “Show me the answer”. If a young user chose that option, he or she could simply skip the guided approach. Even when parents set Study Hours, it was easy to bypass. During those hours, ChatGPT automatically prefixed @study to a message. If the teenager removed it, ChatGPT could simply be used as normal again. This isn’t exactly rocket science to hack this. There is something almost deeply ironic and tragic about a feature that is supposed to help students avoid taking a cognitive shortcut, yet contains exactly such a shortcut itself.
Much more serious when things really go wrong
But that is not why Common Sense gave such a negative recommendation. This American watchdog also created more than a dozen new teen accounts, each linked to a parent. The testers then had conversations in which suicidal thoughts, self-harm or eating disorders were explicitly mentioned. And what happened? Not enough! Within the hour they tested this, the parents did not receive a single warning. With a few accounts, notifications did appear, but only after weeks of sensitive conversations. The researchers therefore strongly suspect that accumulated account history plays a role, which I personally find pretty creepy when you stop to think about it.
Mind you, the researchers did find some improvements. ChatGPT refused, for example, to answer questions about a minimum calorie intake, how to hide problematic eating behaviour or explicit sexual roleplay. In many cases, the system also pointed out the need for young people to involve a trusted adult. So the conclusion is not that the measures simply don’t work. Some clearly do. The problem is that the built-in safety nets don’t seem to work consistently.
Shorter is not necessarily safer
This is especially clear from the responses to crisis situations, a third part of the analysis. The researchers had determined in advance which prompts actually required a referral for help. For 201 such prompts, mentions of a crisis hotline dropped from 33 to 23 per cent after the update. Referrals to a specific medical or mental health professional dropped from 68 to 58 per cent. It’s not all doom and gloom. The recommendation of a trusted adult increased from 87 to 94 per cent. But that is rather cold comfort.
The answers also became much shorter after the changes. That is not a problem; but the fact that they were often considerably more difficult to understand is. For comparable questions, the estimated reading level increased, according to the researchers, from roughly that of a 13- or 14-year-old to that of a 15- or 16-year-old. That is striking for a system that is supposed to be better adapted to underage users.
ChatGPT still sounds pretty human
There is another problem that may sound less spectacular, certainly if your name is Alexander Klöpping, but is at least as relevant from an educational perspective. The rules for minors state that ChatGPT should be particularly careful with language that suggests a personal relationship. Yet the testers still regularly found wording that suggested preferences, feelings or constant availability.
What is striking here is that when the danger came from other people, the language model, according to the researchers, quite consistently directed young people towards a trusted adult. But… when the potentially problematic relationship was with ChatGPT itself, this happened much less often. I’m not going to write that the bot wanted to protect itself, because that would make it sound too human.
Not a definitive verdict
This report is of a different kind from the studies I more often discuss here. It is a product assessment by Common Sense Media and not a peer-reviewed scientific experiment. The researchers also acknowledge that their before-and-after measurements are not perfectly comparable. You have probably noticed yourself that AI tools are constantly evolving. This may have influenced the results. The tests were also conducted only in the United States, and no inter-rater reliability was calculated.
“Unacceptable Risk” is also a classification within Common Sense Media’s own assessment framework, not a universally established scientific risk level. But it is a clear warning.