Two-year study finds banning AI in class leaves students worse off

Thibault Schrepel, a law professor at Vrije Universiteit Amsterdam, spent two years testing how AI access affects student performance in his "Law of AI" course, and while he expected the trained group to lead, the group banned from using AI finished last both years. He randomly split students into three groups working in teams of four or five: one could not use ChatGPT at all, one received ChatGPT-generated revision suggestions embedded in the text with no guidance on how to use them, and one got hands-on training in legal prompt engineering and checking AI output for consistency and accuracy. Every team had 20 minutes to improve a provision of the EU AI Act, graded on substance, clarity, proportionality and innovation, and all students later sat the same multiple-choice test and a take-home exam requiring them to revise another AI Act provision. The experiment ran in 2024 with 66 students and was repeated in 2025 with 164 participants.
The no-AI group mostly made minor wording tweaks and, after ten to fifteen minutes, many subgroups regularly ran out of ideas, an effect Schrepel calls "idea exhaustion," though working without AI did push deeper discussion among team members. The unguided-AI group largely accepted ChatGPT's suggestions without question, students reasoning the wording "sounded better," and every subgroup kept at least one misleading or legally extraneous term from the tool's output; some replaced the terms "shall" and "individual" with AI-suggested alternatives without showing they understood the legal consequences. Only the trained group engaged in genuine back-and-forth with the AI, testing different phrasings and probing the substantive legal questions.
In 2024 the trained group scored well above the other two, especially on the take-home exam, but by 2025 that gap had nearly closed and all three groups performed at roughly the same level. Schrepel attributes this to students already being more familiar with chatbots by then, which left formal training with less to add, though he still argues ethical use and legal responsibility need to be taught. The one result that held in both years: the no-AI group finished last. "Banning AI produces worse average outcomes than allowing it," Schrepel concludes, though he says it remains an open question whether the gains come from structured training or simply from hands-on experience.
The unguided group's showing surprised Schrepel most. He expected the careless errors students made during the in-class exercise to carry into the exam, but the opposite happened: the errors did not repeat, and the group scored slightly higher than the no-AI group, which he attributes to students learning to spot AI's weaknesses through their own use, especially once accuracy carried real consequences. That result overturned his starting assumption that AI only helps with structured training and should otherwise be kept out of the classroom. "I was wrong," he writes.
Schrepel argues department leaders should resist blanket AI bans and give instructors room to experiment, and that universities need to invest in AI skills for faculty, many of whom feel poorly equipped. He also argues the traditional master's thesis needs rethinking, since its core value, the long struggle with research and argumentation, can now be handled in minutes with AI tools; pure literature reviews should no longer count, and programs should require practice-oriented or empirical work where thoughtful AI use is part of the grade. Not every school agrees: UC Berkeley Law has banned AI from nearly all graded work, arguing future lawyers need to build core thinking skills before using AI usefully, a position Schrepel's findings directly contradict.
Schrepel names his own study's limits: a small sample, students enrolled in an AI course who were likely more tech-savvy than average, and no way to verify how much AI students actually used on the unsupervised take-home exam. That last gap matters because other research finds AI's strongest effects show up specifically on unsupervised work. A UC Berkeley study covering more than 500,000 grades found the share of A grades in writing- and programming-heavy courses jumped 13 percentage points after ChatGPT launched, while proctored exams showed no comparable shift. A separate longitudinal study tracking more than 26,000 K-12 students in central China found AI boosted homework grades while exam performance dropped by up to 24 percent, with the damage worst where AI replaced independent thinking rather than supporting it.
Key facts
- A two-year study split law students at Vrije Universiteit Amsterdam into a no-AI group, an unguided-AI group, and a trained-AI group; the no-AI group finished last both years, in 2024 (66 students) and 2025 (164 participants).
- The unguided-AI group accepted ChatGPT suggestions largely without question, and every subgroup kept at least one misleading or legally extraneous term in its output.
- The trained group's 2024 advantage on the take-home exam had nearly disappeared by 2025, which Schrepel attributes to students growing more familiar with chatbots generally.
- Schrepel reversed his starting assumption, writing "I was wrong" after the unguided-AI group's exam scores came in slightly above the no-AI group's.
- UC Berkeley Law bans AI from nearly all graded work on the opposite reasoning, that future lawyers must build core thinking skills first, a stance Schrepel's results directly contradict.
Why it matters
Many universities have responded to generative AI by banning it from coursework on the assumption that unsupervised use erodes learning. This is one of the few controlled, multi-year studies to actually test that assumption against a real alternative, and it finds the ban itself produced the worst outcomes of the three approaches tried, not the safest one.
Who it affects
University departments and instructors deciding AI policy for graded coursework, students in AI-adjacent courses like law, and faculty who Schrepel says often feel poorly equipped to teach with these tools. It also bears directly on schools like UC Berkeley Law that have taken the opposite position and banned AI from nearly all graded work.
How to use it
Schrepel's own recommendation is not to ban AI but to train students in it: teaching legal prompt engineering and how to check AI suggestions for consistency and accuracy produced the strongest results in the first year the study ran. He also argues traditional master's theses should shift away from pure literature review, which AI can now handle in minutes, toward practice-oriented or empirical work that grades thoughtful AI use directly.
How solid is it
The finding held across two independent runs two years apart (66 students in 2024, 164 in 2025), with randomized group assignment and identical exams. Schrepel's own caveats limit how far it generalizes: a small sample from students who chose an AI-focused course and were likely more tech-savvy than average, and no way to verify how much AI students actually used on the unsupervised take-home exam. Two other studies cited alongside it, a UC Berkeley analysis of more than 500,000 grades and a longitudinal study of over 26,000 K-12 students in central China, point the same direction: AI's effects are strongest specifically on unsupervised work, which is exactly the gap this study could not close.
Risks and caveats
Schrepel is explicit that his study cannot say whether the trained group's early advantage came from the structured training itself or simply from more hands-on experience with the tool, and that advantage had nearly vanished by the second year regardless. The China study is a caution against overreading the result: it found AI raised homework grades while exam performance fell by up to 24 percent, worst where AI replaced independent thinking outright, so unrestricted AI access is not shown here to be free of downsides, only better on average than an outright ban in this particular course design.
“I was wrong”
— Thibault Schrepel