Khanmigo AI tutor gives modest math gains in two-year Tennessee trial

Khanmigo AI tutor gives modest math gains in two-year Tennessee trial

A working paper reports some of the first large-scale experimental evidence on AI tutoring. The authors open with the promise: generative AI has been promoted as the technology that could transform education by giving every student a personal tutor. They test it in a two-year cluster randomized trial in 18 Tennessee middle schools. Randomly assigned students used Khan Academy together with its AI tutor, Khanmigo, which was configured to coach rather than give answers. The sessions took place during existing daily remedial mathematics sessions.

The effect was positive but small. Assignment to Khanmigo raised math achievement by 1.3 national percentile ranks per term, which the authors express as about 0.06 to 0.08 standard deviations over a school year. For students who actively participated for a full year, the implied effect reaches 0.14 standard deviations. The authors say these gains resemble those from Khan Academy practice without AI assistance.

The paper offers one explanation for the modest result: students used the tutor infrequently and, when they did, rarely engaged it in substantive mathematical dialogue. Access was not the problem. 96 percent of students tried Khanmigo at least once. But the median student messaged it on only a third of the days they practiced, and in only 17 percent of the exercise sessions in which they made a mistake. The messages students did send were mostly bare answers or clicks on suggested prompts.

The authors' conclusion is that the binding constraint appears to be engagement. Realizing the promise of AI tutoring, they write, will require getting students to use it, not just giving them access.

Key facts

  • The trial ran for two years across 18 Tennessee middle schools, with students randomly assigned to use Khan Academy with the Khanmigo AI tutor during existing daily remedial math sessions.
  • Assignment raised math achievement by 1.3 national percentile ranks per term, about 0.06 to 0.08 standard deviations over a school year; the implied effect of a full year of active participation is 0.14 standard deviations.
  • The authors say the gains resemble those from Khan Academy practice without AI assistance.
  • 96 percent of students tried Khanmigo at least once, but the median student messaged it on only a third of the days they practiced, and in only 17 percent of the exercise sessions in which they made a mistake.
  • The authors conclude that engagement appears to be the binding constraint: students must be persuaded to use AI tutors, not just given access.

Why it matters

The idea that generative AI can give every student a personal tutor has been widely promoted. This paper tests it in a randomized trial rather than a demo, and the authors describe it as some of the first large-scale experimental evidence on the question. The headline result is sober: a real but small gain, similar to what Khan Academy practice delivers without AI. The more useful finding is where the shortfall comes from. Students did not reject the tool; nearly all tried it. They just did not talk to it much, or about much.

Who it affects

Schools and districts weighing AI tutors for math, especially for remedial sessions like the ones studied here. Khan Academy and other makers of tutoring tools, whose products depend on students actually engaging with the AI. Researchers studying AI in education, who now have a randomized benchmark for effect size.

How to use it

The practical reading comes from the authors' own conclusion: giving students access is not enough, and getting them to use the tutor is the hard part. The paper reports that students' messages were mostly bare answers or clicks on suggested prompts, and that the median student used the tutor on a third of practice days. Anyone deploying such a tool has those figures as a baseline to compare against.

How solid is it

The design is strong on paper: a cluster randomized trial running two years across 18 schools, reported in a working paper. The numbers come from the paper's abstract, and the effect sizes are stated with the level specified (per term in percentile ranks, per school year in standard deviations). The 0.14 standard deviation figure is an implied effect of a full year of active participation, not the effect of assignment. The source does not report confidence intervals, p-values or statistical significance, and it does not say how many students took part.

Risks and caveats

The engagement explanation is offered by the authors as one explanation, not as a tested cause. The source does not say whether the comparison to Khan Academy without AI was a separate arm of the trial or an outside benchmark. The Khanmigo in the study was configured to coach rather than give answers, so results may differ for other configurations. The source gives no authors, institutions, publication date or grade levels, and it is a working paper rather than a published article.

“The binding constraint appears to be engagement: realizing the promise of AI tutoring will require getting students to use it, not just giving them access.”

— Working paper abstract