- Home
- The Thinking Wire
- The AI Tutor Improved the Practice Session and Changed Nothing on the Test
The AI Tutor Improved the Practice Session and Changed Nothing on the Test
Students given a GPT tutor with guardrails performed 127% better on their math practice problems than students who had no model at all. Then they sat an unassisted exam, and the paper reports the result in one flat sentence: “Student performance in the GPT Tutor arm was statistically indistinguishable from that of the control arm.”
That is a randomized controlled trial, run in collaboration with a large high school in Turkey during the fall semester of the 2023-2024 academic year, across roughly 1,000 students. Bastani et al. published it in PNAS as “Generative AI without guardrails can harm learning: Evidence from high school mathematics.” The GPT Base arm performed statistically significantly worse than control, by 17%, on the unassisted exam. That number is real and correctly belongs to the Base arm, not to AI users in general.
The number that should worry us more is the one where nothing happened.
The Remedy We Implied Does Not Survive the Test
We have argued twice that AI generates cognitive debt, and that the debt is now reaching the talent pipeline. Both pieces carried the same implicit fix, and it is the obvious fix: do not let the model produce the answer, make it teach. Turn the generator into a tutor. Add guardrails, add Socratic prompting, withhold the solution, ask the student to reason.
The GPT Tutor arm in the Turkish trial is the guardrailed version of that intervention. It worked, spectacularly, inside the session. It produced no measurable advantage in the student.
A 127% lift that decays to zero the moment the tool is removed is not a small disappointment. It means the practice session was measuring the pair, the student plus the model, and reporting the score as if it belonged to the student. Any organization measuring an AI-assisted onboarding program on in-session output is reading the same instrument.
What the Session Measures Is Not Who Is Learning
The mechanism shows up more clearly in the qualitative work. ACM’s “The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers” observed novice programmers working with AI assistance and recorded a mismatch between what participants believed they were doing and what they did. In the authors’ words: “Participants thought it was like having a personal tutor. From the data in our study … we observed that they did not, in fact, use GenAI tools like a personal tutor. In fact, it was quite the opposite.”
No sample size is available for it, and it is qualitative, so treat it as a description of behavior rather than an effect estimate. The behaviors it describes are specific enough to recognize on any team. Heavy-AI participants “often skipped crucial planning stages.” They finished with what the paper calls an “illusion of competence.” One participant “had to rely on the LLM to fix the error that the LLM introduced in the first place,” which is the cleanest short description of a compounding dependency I have read.
Anthropic’s 2026 study on how AI assistance affects the formation of coding skills lands on the same mechanism from the other direction: “Cognitive effort, and even getting painfully stuck, is likely important for fostering mastery.” Struggle is the input to expertise. A tutor that reduces struggle inside the session is optimizing against the thing that produces the outcome.
The Skill That Actually Correlated
One group in the ACM study did well. The paper describes what separated them. The successful mitigated-AI group had developed “negative expertise, the ability to ignore incorrect or unhelpful GenAI suggestions.”
That is a hiring criterion and a training target, and it is close to the opposite of the enablement programs I have seen. Knowing when to reject the output requires already holding a model of what correct looks like. This is why the capability is so awkward to build with the tool in hand: you acquire it by being right about something the model was wrong about, repeatedly, which requires having been in a position to know.
François Chollet frames the ceiling that makes this permanent rather than transitional: “LLMs are a static database of skills. They are interpolation engines. Software engineering, however, is an exercise in adaptation and novel problem-solving. You cannot interpolate your way through a completely unique system failure.” Your outage at 3am is not in the database. The person handling it either has negative expertise or is guessing with a confident co-pilot.
Six Questions, Available Monday
Lars Faye’s essay “AI Coding will Prevent Expertise” (published in July 2026, so read it as an argument rather than a fresh finding) contains the most useful artifact I have seen on this. It is a six-question checklist a developer runs before accepting AI assistance on a task. Four of them, in his words:
- “If I did not have access to an AI tool, could I still accomplish this task?”
- “Am I using the model to deepen my understanding, or expedite the answer?”
- “If I had to audit and verify the generated output, could I adequately explain what was happening?”
- “Is this a truly rote task that’s been done 100 times before, or a task that requires executive decision-making somewhere in the process?”
The checklist works because it separates two things that get collapsed into one word. Faye’s distinction: “Cognitive debt is abdicating your judgment and decisions, whereas cognitive offloading is delegating the mechanical or tedious.” Offloading a migration script you have written forty times is free. Offloading the decision about whether the migration should exist is the debt.
And there is the line every engineering leader should take to their next architecture review: “If you can’t properly audit the accuracy of the generated code, then you can’t audit the accuracy of the generated concept.” Audit capability is the ceiling on delegation. A team that cannot audit the code has, without deciding to, delegated the design.
Do This Now
Take one of your most recent engineering hires. Give them a task inside your production codebase, in a scope they have already shipped in with AI assistance, and remove the assistance for the duration. Watch what happens in the first thirty minutes. You are not testing whether they are slower. Of course they are slower. You are testing whether they can start.
That is the unassisted exam, and it is the only reading that tells you whether the last year of accelerated output built capability or rented it. The practice score is the analogue of your velocity chart. In the trial it moved 127% without the student moving at all.
Then put Faye’s four questions in the pull request template, where the decision actually happens. The first one is enough on its own to change behavior, because it is the only question in the set that a developer cannot answer honestly while the tool is open.
This analysis synthesizes AI Coding will Prevent Expertise (Lars Faye, July 2026), Generative AI without guardrails can harm learning: Evidence from high school mathematics (Bastani et al., PNAS, 2025), and The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers (ACM, 2024).
Victorino Group helps engineering organizations measure whether AI-assisted delivery is building capability or renting it. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation