Four students at a classroom table working on paper while a laptop between them glows with an abstract blue pattern of connected light, a lecturer writing on a whiteboard in the background
← Back to Insights

They looked like they were learning. Then the AI was taken away.

2 September 2026 · 6 min read · By Paul Robinson · LearnFrame Insights

Nearly a thousand teenagers sat a maths exam in Turkey. Some of them had spent the practice sessions beforehand working with a chatbot. In the exam, with the chatbot gone, they scored 17 percent worse than the students who had never been given one at all.

They had not been lazy. During practice they had looked like the best students in the room.

That result was published in the Proceedings of the National Academy of Sciences on 25 June 2025. It got a fraction of the attention the other study got. It deserved more. Every Head of Education at a professional body has members running the same experiment on themselves right now, and nobody is going to score the exam.

The good news first

Start with the study everybody heard about.

In autumn 2023, two Harvard physics lecturers, Gregory Kestin and Kelly Miller, ran a trial with 194 of their own students. Each student did two lessons in consecutive weeks. One was in a classroom, taught by a good instructor, using the kind of hands-on group teaching that decades of research says works best. The other covered the same material at home, alone, with an AI tutor the lecturers had built themselves.

Then everyone was tested.

The students learned more than twice as much from the AI tutor. They did it faster: a median of 49 minutes against 75 in class. And they said they found it more engaging. The paper came out in Scientific Reports in June 2025.

That is the headline that travelled. AI tutor beats Harvard classroom. If you sat in a conference in the last year, someone put it on a slide.

Now the second study

The Turkish study was run by a team led by Hamsa Bastani at the Wharton School, and it was built to ask a harder question. Not “does the AI help while you are using it?” but “what are you left with when it is gone?”

Nearly a thousand secondary school students were split three ways for their maths practice sessions. One group got a chatbot that behaved like ordinary ChatGPT. One group got a chatbot that had been given the teachers’ worked solutions and told to give hints, never the answer. One group got no AI at all.

During practice, the AI groups raced ahead. The plain chatbot group scored 48 percent better than the students with no help. The hints group scored 127 percent better.

Then the researchers took both chatbots away and ran an exam.

The plain chatbot group scored 17 percent worse than the students who never had a chatbot. Worse than nothing. The hints group came out level with the no-AI group. The damage was gone, but so was the gain.

The authors’ word for what happened is “crutch”. The students with the plain chatbot had used it to get answers. It felt like learning. It produced good scores. It left nothing behind.

It felt like learning. It produced good scores. It left nothing behind.

Same technology, opposite result

Put the two studies side by side and something jumps out.

Both used the same underlying model. Both put a chatbot in front of a student with a problem to solve. One beat a well-taught Harvard classroom by a factor of two. The other produced a group of students who would have been better off with no help at all.

The difference was not the AI. It was what a human decided the AI was allowed to say.

Harvard’s tutor was fed the lecturers’ own solutions and given a short set of rules. Give one step at a time. Never give the whole answer in one message. Keep it to a few sentences. Ask the student to try first. Those rules were written by two people who had spent years teaching that exact course to that exact kind of student.

The Turkish hints tutor had the same idea in a weaker form: the teachers’ solutions, and an instruction not to hand over the answer. That was enough to stop the harm. It was not enough to create a gain.

The plain chatbot had no rules at all. It answered whatever it was asked. That is what it is built to do, and that is what did the damage.

So the active ingredient was expert judgement, applied in advance, about what a learner should be made to do for themselves. The AI amplified whatever design it was given. Good design, twice the learning. No design, worse than nothing.

Why this lands on a professional body

Think about what your members are doing this afternoon.

A solicitor drafting a clause. An accountant checking a treatment. A pharmacist answering a query. An engineer sizing a component. Many of them now do the first pass with a chatbot, and it is the plain kind. No hints. No rules. No expert’s solution behind it. Just answers, on demand, that look right.

That is the arm of the Turkish study that came out 17 percent worse.

Nobody is going to run the exam. There is no moment when the tool is taken away and the profession finds out what its members can still do without it. It shows up instead, one case at a time, as a confident draft that nobody checked, from a professional who used to be able to do it unaided.

And here is the part that should worry a Head of Education more than the members.

Many professional bodies are about to buy or build a CPD programme about AI. A lot of those will arrive with an AI assistant built into the corner of the screen, because that is what the market is selling and it demos beautifully. If that assistant simply answers questions, the body has bought the plain chatbot arm. The programme will look excellent while people are inside it. Completion will be high. Satisfaction will be high. The bill comes later, in what the members can do once they close it.

What separates a good one from a bad one

The two studies hand a Head of Education a clear test to put to any supplier, and it has nothing to do with which model the supplier uses.

Who decided what the learner has to do for themselves, and when did they decide it?

In every arm of either study that did no harm, the answer was the same: a subject expert, before the first learner ever logged on. The expert wrote the solutions. The expert wrote the rules. The AI followed them.

If the answer is “the AI works it out as it goes” or “the learner can ask it anything”, that is the plain chatbot. It will score well on every measure you have except the one that matters.

This is also the honest answer to a question I get asked a lot, which is whether AI makes building a programme cheaper. In production, in places, yes, and I have written about where. But the thing that made Harvard’s tutor work was not produced faster by AI. It was two lecturers deciding, line by line, what a student should be allowed to get without earning it. That is design work. It is slow, it is expert, and it is the part that decides whether anything is left behind.

The exam nobody runs

The Turkish researchers did the one thing no professional body can do. They took the tool away and measured what was left.

You cannot take the tool away from your members. You can decide what they practise with. A module that answers their questions will feel helpful and leave nothing. A module that makes them decide, and then shows them how their decision compared with an expert’s, is the hints tutor with the missing piece added.

That is how our Module 1 works: three situations, no answers given, and a score for how you decided.

A concrete next step

Want to know what your own programme could tell you that it currently does not? The Programme Design Diagnostic is a fixed-fee, independent read on a single programme: a board-ready diagnosis and a costed build scope, in two to three weeks.

The Diagnostic →

Building member or staff CPD on AI?

AI for Professional Practice is open now, and it is built to be licensed by professional bodies, training providers and corporate academies, branded as yours and re-verified quarterly. Early conversations shape the licensing round.

Licence the Programme → Arrange a Conversation →

Sources: Kestin, Miller, Klales, Milbourne and Ponti, “AI tutoring outperforms in-class active learning”, Scientific Reports 15, 17458, June 2025 (194 students, autumn 2023, learning gain more than double, median 49 minutes against 75). Bastani, Bastani, Sungu, Ge, Kabakci and Mariman, “Generative AI without guardrails can harm learning: evidence from high school mathematics”, Proceedings of the National Academy of Sciences 122 (26), 25 June 2025 (nearly one thousand students; 48 and 127 percent better in practice; 17 percent worse in the exam for the unguarded group). Both retrieved 2 September 2026.

For where AI does and does not speed up production, see what a custom eLearning build actually looks like. For how a decision-based result should be scored, see one test, five ways of scoring it. For what AI literacy means for staff, see AI literacy in plain English. See more insights from LearnFrame.

About the author

Paul Robinson is the founder of LearnFrame, which designs and builds custom eLearning programmes for professional certification bodies, regulated training providers, and corporate academies. He has worked in digital learning for three decades, since the 1995 Nasdaq IPO of CBT Systems. Connect on LinkedIn.