StudentBench study finds AI tutoring matches expert humans on GRE gains

StudentBench study finds AI tutoring matches expert humans on GRE gains

Researchers built StudentBench, a suite of AI teaching evaluations paired with a public platform, to test whether large language models can produce learning gains on par with human tutors rather than just advancing raw model capability. The platform has already collected more than 175,000 student-AI messages. Using it, the team ran a study measuring learning gains on Quantitative and Verbal GRE questions across 2,383 human participants, who were assigned to AI tutoring, expert human tutoring, or no tutoring at all. The headline result: AI tutoring was statistically equivalent to expert human tutoring for GRE learning gains (p = .015). Broken down by GRE domain, the best performing AI tutor actually surpassed the human tutor's average results in five of the seven GRE domains tested. In a second, separate study, expert human tutors ran 2,028 pairwise rubric evaluations comparing LLM-generated lesson plans and practice problems against alternatives. Combined, the two studies let the researchers compare AI tutors across five dimensions: lesson planning, practice-problem creation, conversational pedagogy, cost, and engagement. The most striking single finding involved cost: one AI tutor achieved learning gains statistically equivalent to human tutoring (p = .044) while costing 918 times less, USD 0.0052 per percentage point of gain for the AI tutor versus USD 4.81 per percentage point for the human tutor. The researchers also found that within Quantitative GRE sessions, several factors were linked together: faster AI replies correlated with students sending more messages, more messages correlated with more correct practice problems attempted, and more correct practice correlated with larger learning gains, all relationships significant at p < .002. The StudentBench platform itself is freely available at studentbench.org for further research and data collection.

Key facts

  • StudentBench, a new AI teaching evaluation suite and public platform, has collected over 175,000 student-AI messages
  • A study of 2,383 human participants found AI tutoring statistically equivalent to expert human tutoring on GRE learning gains (p = .015)
  • The best AI tutor beat the human tutor's average in five of seven GRE domains tested
  • One AI tutor matched human tutoring outcomes (p = .044) at 918 times lower cost: USD 0.0052 versus USD 4.81 per percentage point gained
  • A separate study used 2,028 pairwise rubric evaluations by expert human tutors to compare AI-generated lesson plans and practice problems

Why it matters

Most AI development effort goes into raw model capability, not into measuring whether that capability actually teaches people anything. StudentBench is built specifically to answer that narrower, more practical question with a controlled study rather than anecdote, and it finds that on a real, high-stakes test like the GRE, AI tutoring is not just plausible but statistically on par with expert human instruction, at a small fraction of the cost.

Who it affects

The direct subjects are the 2,383 human participants who went through GRE tutoring for the study, plus the expert human tutors who both delivered lessons and rated AI-generated lesson plans and practice problems in the 2,028 pairwise evaluations. More broadly it speaks to test-prep companies, education platforms, and anyone building or evaluating AI tutoring products, since the study offers a rare large-scale, controlled comparison against human tutors rather than a vendor's internal benchmark.

How to use it

The StudentBench platform is freely available at studentbench.org, allowing other researchers or developers to run their own AI teaching evaluations or contribute to the growing message dataset. The source text does not name the specific AI models tested or their commercial availability, so there is no product to buy or subscribe to based on this study alone.

How solid is it

The result rests on a fairly large sample for education research, 2,383 participants across AI tutoring, human tutoring, and no-tutoring arms, plus a second study with 2,028 pairwise expert evaluations, and the equivalence claims are backed by specific significance values (p = .015 for the overall equivalence, p = .044 for the single best-cost AI tutor). The source text does not give the authors' names, institutions, publication date, or the specific AI models used, and it does not describe what the no-tutoring control condition actually involved beyond naming it as an arm of the study.

Risks and caveats

The study covers GRE Quantitative and Verbal prep specifically, so the equivalence finding may not generalize to other subjects or teaching contexts. The cost comparison of USD 0.0052 versus USD 4.81 per percentage point applies to one particular AI tutor rather than AI tutoring in general, and the five-of-seven-domains result describes the best performing AI tutor, not every AI tutor tested. Because the source text does not name the models or tutors involved, readers cannot independently verify which specific systems produced the strongest results.

“We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains, and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average.”

— StudentBench study