What Are We Buying When We Buy an AI Tutor?
Instructional vehicles need properly-trained drivers and a properly-designed institutional and pedagogical highway system.
The promotional materials make enormous promises. An AI interactive space, sealed off from the open internet and tuned for practice, will supposedly do what a teacher with thirty students cannot: meet every learner where they are, respond instantly, and never tire. Districts and universities are buying into these systems on the premise that the space itself is the instructional advance, as if placing students inside an AI-mediated environment will reliably produce better learning.
The first several years of evidence give us a more specific answer. AI tutors can raise achievement under the right instructional conditions, sometimes impressively. The strongest studies, however, do not show that an interactive AI environment works by itself. They show that carefully designed scaffolding, close alignment between practice and assessment, supplemental time, and teacher mediation can produce gains when AI is one part of the instructional arrangement.
The strongest case for AI tutoring
In the fall of 2023, Gregory Kestin, a physics lecturer at Harvard, ran a randomized controlled trial comparing a custom AI tutor with his own active-learning classroom, a format that physics education research has spent decades validating. One hundred ninety-four students rotated through both conditions, learning one topic in class and another at home with the tutor. Students learned more than twice as much with the AI tutor in less time, according to the published study in Scientific Reports (Kestin et al., 2025).
A second strong result comes from Edo State, Nigeria, where the World Bank evaluated Microsoft Copilot, running on GPT-4, as an English tutor for first-year senior secondary students. Over six weeks, students in nine public schools used the system in after-school sessions. The reported gains were 0.23 standard deviations in English and higher on a combined measure that included digital skills. The authors describe the intervention as highly cost-effective and suggest that, if effects scaled across a full school year, the gains could equal one and a half to two years of ordinary schooling (World Bank, 2024).
These are significant findings, and I’d argue that they should make skeptics pause. They also require careful reading. The Harvard trial tested a tutor built by the instructor and his team around the exact learning objectives, problems, and feedback structure of the unit. The Nigeria trial tested supervised, teacher-led, supplemental use of Copilot in a low-resource setting. In both cases, AI was embedded in a pedagogical design that constrained what the tool did and how students interacted with it.
What the research designs actually show
Kestin’s tutor was not a general-purpose chatbot dropped into a physics course. It had the correct solutions in advance, released help one step at a time, required students to attempt work before receiving explanations, and targeted the same material measured by researcher-built pre- and post-tests. The comparison condition was not weak instruction; it was Kestin’s own active-learning classroom. The result is therefore best read as evidence for a tightly scaffolded AI tutor aligned with a short-term assessment, not as evidence that AI spaces in general outperform classrooms.
The World Bank study has a similar structure. Students did not simply log into Copilot and learn independently. Teachers led the after-school sessions and guided students through their interactions with the system. The largest gains appeared among female students and students with higher baseline performance, which complicates the sales pitch that AI tutors will automatically rescue the students furthest behind. The “two years of schooling” figure is an extrapolation from six weeks, while the measured finding is narrower: supervised, teacher-mediated AI practice added to existing instruction improved English scores in a specific setting.
This distinction is significant because districts are rarely buying Kestin’s exact physics tutor or the World Bank’s implementation model. They are buying platforms, access, dashboards, and contained environments. The evidence supports the instructional arrangements that made the AI useful. It does not establish that the environment alone carries the effect.
Scaffolded use and unstructured use
Khanmigo is useful here because it was built with far more pedagogical intention than most district-level AI deployments. Sal Khan has since acknowledged that the first version did not change student learning as much as many people hoped (Khan Academy, 2026). That concession is not a failure of imagination; it is a useful correction to the idea that a safe, conversational, always-available AI tutor will automatically produce deep learning.
The independent evidence is mixed rather than catastrophic. One controlled study comparing Khanmigo with a Google search group and a paper-only group found learning gains across conditions, with no statistically significant differences among the three groups (Yaylali and Mehta, 2025). Students liked Khanmigo and saw it as a helpful supplement, but the quantitative result did not show that the AI tutor outperformed the alternatives. That is exactly the pattern districts should care about: engagement and perceived helpfulness can rise without producing a distinctive learning effect.
There is also reason to be cautious about unstructured AI assistance. A widely discussed experimental finding summarized in the literature reports that students using AI assistance solved more practice problems correctly while the tool was available, then performed worse on conceptual understanding when the assistance was removed (Jose, 2025). The mechanism is not mysterious. A tool that helps students get through tasks can also reduce the amount of thinking they do for themselves, especially when the design rewards speed, completion, or correct answers rather than durable understanding.
The older lesson: method travels through medium
Richard Clark made the classic version of this argument in 1983: media are vehicles for instruction, and a delivery truck does not change the nutrition of the groceries it carries (Clark, 1983). His claim was not that technology is irrelevant. It was that researchers often confuse the medium with the instructional method delivered through it. When a new medium appears to improve achievement, the actual cause may be the tutoring sequence, feedback loop, practice schedule, teacher mediation, or assessment alignment packaged inside the medium.
Research on intelligent tutoring systems broadly supports that caution. A 2014 meta-analysis found that intelligent tutoring systems outperformed large-group classroom instruction and performed about as well as individualized human tutoring across the included comparisons (Ma et al., 2014). A 2015 review reported substantial average effects for intelligent tutoring systems, though effects were much smaller on standardized tests than on assessments more closely tied to the tutored material (Kulik and Fletcher, 2015). That gap is important. The more an assessment resembles the system’s own practice environment, the easier it is for the system to look powerful; transfer to broader measures is harder.
This older literature does not contradict the newer AI tutoring studies. It helps explain them. The strongest effects tend to appear when the tool is tied to specific content, sequenced practice, immediate feedback, and a nearby measure of learning. Those conditions can be valuable, but they are instructional conditions. Calling them an “AI interactive space” hides the part of the design that drives the instruction.
Who benefits?
Justin Reich, in Failure to Disrupt, argues that educational technologies often get adopted in ways that extend existing practice rather than transform it. He also emphasizes the Matthew effect: new tools often benefit students who are already positioned to take advantage of them (Reich, 2020). The early AI tutoring evidence fits that warning more than the promotional language does.
The Harvard students were already high-performing learners at an elite university, comfortable with tests, self-paced study, and abstract problem-solving. In Nigeria, higher-performing students gained more than lower-performing students. These findings do not mean AI tutoring cannot help struggling learners. They mean the help is not automatic. Students who already know how to ask questions, persist through confusion, interpret feedback, and connect examples to underlying concepts are better prepared to benefit from an AI tutor. Students without those habits may need more human structure, not less.
What the space is actually buying
If the instructional method is doing the work, the purchasing question changes. A district should not ask whether an AI interactive environment is effective in the abstract. It should ask what instructional method the environment enforces, what evidence supports that method, which students benefit, which students stall, and whether the measured gains transfer beyond the platform’s own practice tasks.
The Harvard result is a scaffolding study. The Nigeria result is a supplemental-instruction study with human mediation. Khanmigo’s early results suggest that a carefully branded and pedagogically serious tool can produce modest or indistinct learning effects when students use it as a general helper. The older intelligent tutoring literature says the same thing in a different vocabulary: the medium can deliver instruction, but the method determines what students learn.
Students need assessment that captures what they know beyond a purpose-built pre-post test. They need instruction in how to use AI with the rigor their discipline requires. They need scaffolding calibrated to their current understanding, and that calibration depends on knowledge of both the subject and the student. An interactive space may support those tasks, but it cannot replace them by existing.
The truck is still not the groceries. Institutions buying AI tutoring environments are buying delivery capacity, and delivery capacity can matter. The hard question is whether the system contains instruction worth delivering.
Nick Potkalitsky, Ph.D.
Check out some of our favorite Substacks:
Mike Kentz’s AI EduPathways: Insights from one of our most insightful, creative, and eloquent AI educators in the business!!!
Terry Underwood’s Learning to Read, Reading to Learn: The most penetrating investigation of the intersections between compositional theory, literacy studies, and AI on the internet!!!
Alejandro Piad Morffis’s The Computerist Journal: Unmatched investigations into coding, machine learning, computational theory, and practical AI applications
Michael Woudenberg’s Polymathic Being: Polymathic wisdom brought to you every Sunday morning with your first cup of coffee
Rob Nelson’s AI Log: Incredibly deep and insightful essay about AI’s impact on higher ed, society, and culture.
Michael Spencer’s AI Supremacy: The most comprehensive and current analysis of AI news and trends, featuring numerous intriguing guest posts
Daniel Nest’s Why Try AI?: The most amazing updates on AI tools and techniques
Jason Gulya’s The AI Edventure: An important exploration of cutting-edge innovations in AI-responsive curriculum and pedagogy



Nick — one of the most useful pieces you've written this year. The Kestin analysis in particular deserves wider attention. The media-vs-method distinction lets us read a "2x gain" study honestly without either dismissing the result or overclaiming what it shows.
For readers who stay with it: what Kestin built was, in Complex Adaptive Systems terms, a rule-governed adaptive scaffold. Not a general-purpose chatbot, not an "AI space" — a specific architecture that constrained what the tool did at each turn, aligned to the assessment, refused to answer before student attempt, and held its solution structure against student pressure. That is what worked. Extending Kestin's finding to "AI tutoring in general" is exactly the medium-vs-method confusion Clark 1983 warned about, and you're right to name it.
Where I'd push slightly on Clark's metaphor: the truck-and-groceries framing assumes the medium is neutral. In AI-tutoring architecture, the medium is not neutral. The harness — the orchestration and local-rule layer above the model — IS exactly what makes the method operational at scale. A harness that enforces anti-fold discipline, requires completion-claim verification, and holds the productive-struggle band at each turn is doing methodological work at the substrate level. Kestin's team built one instance of that harness, tuned to one physics unit. The architectural question is whether we can specify the pattern so it generalizes across domains without needing to be custom-built each time.
To be honest about the evidentiary asymmetry: Crucibl has structured-process evidence from two production semesters at Ensign College, but not a Kestin-caliber RCT. That within-class efficacy study is in preparation. The architectural claim is that the pattern generalizes; the RCT is what will falsify it.
Wrote a longer version of this argument this week — "The Lighthouse and the Ship" is on my Substack. Your piece is required reading, as always.
— Chris Wasden, EdD
Insightful read, Nick!
What are we buying when we buy an AI tutor?
My mind immediately thought: AI slop 🤖🤭