AI writing detectors are unreliable, but everyone uses them anyway

Long before ChatGPT, teachers and editors used anti-plagiarism tools like Turnitin to compare a piece of writing against a database of web and scholarly content and flag matching passages. Since ChatGPT, Google Gemini and Microsoft Copilot spread among students, educators have adopted AI detectors just as fast: a Center for Democracy and Technology survey found that 43 percent of sixth to 12th grade teachers in the US regularly used AI detectors between 2024 and 2025. Some universities running Turnitin in their learning management systems found the platform automatically switched on AI detection when that feature launched in 2023.
Unlike the older plagiarism checkers, newer tools such as GPTZero, Pangram and Turnitin's own AI detector do not match text against a database. They use their own AI models to guess whether a passage was AI-written by analyzing wording, rhythm and structure, plus patterns in length and tone, a more subjective process that can get tripped up by writers who speak English as a second language. The vendors nonetheless publish low error rates: Turnitin says it falsely flags less than 1 percent of human writing as AI, Pangram claims a false positive rate of 1 in 10,000, and GPTZero claims a similarly low rate.
The consequences of a wrong flag are already landing on real people. The publisher Minotaur dropped a $2 million book deal with author Jerry Falade last month over suspicion that he used AI, something he denies. Thierry Rignol, a French national, sued Yale after a professor used GPTZero to accuse him of writing part of his final exam with AI, which led to a failing grade and a one year suspension; his lawsuit argues that AI surveillance and detection tools are known to unfairly target non-native English speakers like him. In February, a student at Adelphi University won a lawsuit after a professor made a similar AI accusation against him (the lawsuit does not name the tool used, though Adelphi licenses Turnitin). A 2023 Stanford study found that AI detectors falsely flagged essays by non-native English speakers as AI more often than essays by native speakers, and the tools may also be biased against neurodivergent writers.
Even the vendors hedge their own products. Turnitin maintains its detector 'may not always be accurate' and should not be used alone to take action against a student. Grammarly warns that users 'should never rely on the results of an AI detector alone,' and GPTZero says 'no AI detector can ever truly be 100% perfect.' OpenAI shut down its own AI writing detector in 2023 because of low accuracy.
The accusations have spread beyond the classroom. Last week, in a video broadcast to more than 3.5 million followers across his social channels, Jack Osbourne, Ozzy Osbourne's son, accused journalist and Verge contributor Kat Tenbarge of using AI to write a Rolling Stone article, citing results from a detector called Getsolved as 'proof.' Tenbarge refuted the claim in a video and on her website, but Osbourne has not retracted the accusation or deleted the video, leaving her to deal with the resulting harassment.
Some institutions have pulled back entirely: Yale, Johns Hopkins, Vanderbilt, Georgetown and other universities have disabled or restricted AI detection tools, and MIT warns that 'AI detectors don't work.' Instead, schools are pushing alternatives. The University of Chicago suggests slowing down students' reading, breaking assignments into smaller pieces and requiring reflection on the work. Stanford suggests holding assessments in the classroom. MIT advises letting students disclose AI use on an assignment without penalty.
Detection is also spreading into publishing and social platforms. Substack has built Pangram into its app so users can scan blogs for suspected AI content, and LinkedIn added a 'seems like AI slop' button to posts. The Authors Guild now offers writers 'Human Authored' certification, alongside 'Not by AI' and 'Written by Human' badges. Wikipedia published a guide to help editors spot AI writing, flagging text that 'puffs up' a topic's importance or gives 'superficial analysis of information,' and has banned AI-generated articles outright.
Key facts
- A Center for Democracy and Technology survey found 43 percent of US sixth to 12th grade teachers regularly used AI detectors between 2024 and 2025
- Publisher Minotaur dropped a $2 million book deal with author Jerry Falade over AI-use suspicion, which he denies
- Thierry Rignol sued Yale after a professor used GPTZero to fail and suspend him for a year; his lawsuit says such tools unfairly target non-native English speakers
- Turnitin claims under 1 percent false positives and Pangram claims 1 in 10,000, yet Yale, Johns Hopkins, Vanderbilt and Georgetown have restricted or dropped AI detectors, and MIT says the tools 'don't work'
- Jack Osbourne accused journalist Kat Tenbarge of using AI for a Rolling Stone article, citing the detector Getsolved as proof, to his more than 3.5 million followers
Why it matters
As students adopted ChatGPT, Gemini and Copilot, educators and publishers adopted AI detectors just as fast: 43 percent of US sixth to 12th grade teachers used them regularly between 2024 and 2025. Unlike older plagiarism checkers, which match text against a database, tools like GPTZero, Pangram and Turnitin's AI detector guess based on wording, rhythm and structure, a subjective method with no direct evidence trail. That shift moves the burden of proof from a matched source to an algorithm's opinion, and institutions are relying on it at scale before its limits are widely understood.
Who it affects
Writers and students bear the direct cost. Author Jerry Falade lost a $2 million book deal with Minotaur over AI-use suspicion he denies. Thierry Rignol received a failing grade and a one year suspension from Yale after a professor used GPTZero on his final exam; his lawsuit argues the tools unfairly target non-native English speakers like him. An Adelphi University student won a lawsuit in February after a similar accusation. Journalist Kat Tenbarge was publicly accused by Jack Osbourne, citing the detector Getsolved, in front of his more than 3.5 million followers. A 2023 Stanford study found non-native English speakers' essays were falsely flagged as AI more often than native speakers' essays, and the article notes the tools may also be biased against neurodivergent writers.
How to use it
Universities that keep using detectors are pairing them with alternatives: the University of Chicago suggests slower reading of student work, breaking assignments into smaller pieces and requiring reflection; Stanford suggests holding assessments in the classroom; MIT advises letting students disclose AI use without penalty. Beyond school, the Authors Guild offers a 'Human Authored' certification, and 'Not by AI' and 'Written by Human' badges exist for writers who want to preempt accusations. Wikipedia published a guide for editors to spot AI writing, watching for text that 'puffs up' a topic's importance or gives 'superficial analysis of information,' and has banned AI-generated articles. Substack built Pangram into its app to scan blogs, and LinkedIn added a 'seems like AI slop' button on posts.
How solid is it
The vendors' own numbers sound reassuring: Turnitin claims under 1 percent false positives, Pangram claims 1 in 10,000, and GPTZero claims a similarly low rate. But the same vendors hedge hard. Turnitin says its tool 'may not always be accurate' and should not be used alone against a student. Grammarly says users 'should never rely on the results of an AI detector alone.' GPTZero says 'no AI detector can ever truly be 100% perfect.' MIT states flatly that 'AI detectors don't work.' OpenAI shut down its own AI writing detector in 2023 over low accuracy. A 2023 Stanford study documented the bias against non-native English speakers directly.
Risks and caveats
Yale, Johns Hopkins, Vanderbilt, Georgetown and other universities have already disabled or restricted AI detection tools over the uncertainty. Reputational damage from a false accusation does not come with an easy fix: Kat Tenbarge refuted Jack Osbourne's claim, but he has not retracted the video or deleted it, and she has had to deal with the resulting trolling on her own. The article does not report the outcome of Rignol's lawsuit against Yale, whether Falade's book was ever confirmed or disproven to involve AI, which detector the Adelphi University professor used, or how the Getsolved tool that Osbourne cited actually works.
“AI detectors don't work”
— The Massachusetts Institute of Technology