Pew Research Center finds AI text on a third of web pages since ChatGPT

Pew Research Center finds AI text on a third of web pages since ChatGPT

The Pew Research Center analyzed nearly half a million English-language web pages drawn from the Common Crawl archive, screening them for signs of machine authorship with the AI detection tool Open Pangram. In a sample from July 2026, about 10 percent of all pages examined showed clear signs of AI authorship; restricting the sample to pages published after ChatGPT's launch in late 2022 changed the picture sharply, as more than a third of those newer pages showed signs of AI-generated text, a share that has climbed steadily since the chatbot's debut.

The signal also varied sharply by domain type. About one in ten .com pages showed signs of AI authorship, compared with 4.6 percent of .org pages and about 1 percent each for .edu and .gov pages, making commercial sites roughly ten times more likely to carry AI-written text than educational or government sites.

Pew's analysis also tracked how web writing itself has changed since 2023. Em dashes now appear about twice as often as they did in 2023, and Oxford comma usage is up 63 percent. Words that have become AI-writing hallmarks, including "delve," "interplay," "testament," "pivotal," "landscape," "tapestry," "bolstered," "crucial," "meticulous," and "vibrant," have more than doubled in frequency. Negative parallelisms following the "it's not just X, it's Y" pattern have nearly tripled, though Pew notes they remain rare in absolute terms; a separate study of corporate PR documents, not part of Pew's analysis, found the same phrase had quadrupled since 2022.

A separate study from April 2026 by Imperial College London, the Internet Archive, and Stanford University reached a similar conclusion, estimating that roughly 35 percent of all newly published websites were fully or partly AI-generated. That study also found 33 percent higher semantic similarity between AI-written texts and a much more positive overall tone, though its researchers cautioned that public perception of AI text's harms often outruns what the data actually shows.

The article itself flags a limitation running through all of this research: current detection tools can barely distinguish fully automated text from writing a human drafted and only partly polished with AI. Based on testing hundreds of their own texts, the article's author says such tools can only make a rough call on whether a human or a machine likely wrote something; they cannot reliably say how much AI was involved or at what stage, and they still misfire regularly. The public debate around AI text is growing more polarized, sharpened by discussion of Anthropic's planned watermark for Claude's output, and workplace studies have documented that using AI tools already carries a stigma that cuts both ways.

Key facts

  • Pew Research Center analyzed nearly half a million English-language web pages from the Common Crawl archive using the Open Pangram detection tool.
  • More than a third of pages published since ChatGPT's late-2022 launch show signs of AI-generated text, versus about 10 percent of all pages in a July 2026 sample.
  • Commercial .com domains show AI-authorship signs roughly ten times more often than .edu or .gov domains: about one in ten .com pages versus about 1 percent each for .edu and .gov, and 4.6 percent for .org.
  • Since 2023, em dashes appear about twice as often, Oxford comma usage is up 63 percent, and AI-favorite words like "delve" and "pivotal" have more than doubled in frequency.
  • A separate April 2026 study by Imperial College London, the Internet Archive, and Stanford University found roughly 35 percent of newly published websites fully or partly AI-generated, but the article cautions that detection tools still cannot reliably distinguish fully automated text from lightly AI-assisted writing.

Why it matters

The study puts hard numbers on how much of the web is now AI-assisted or AI-generated, and ties the rise directly to ChatGPT's release in late 2022. Because Pew connects the trend to specific, measurable stylistic markers, such as the jump in em dashes, Oxford commas, and words like "delve" and "pivotal," it gives a more concrete signal than a general impression. An independent April 2026 study from Imperial College London, the Internet Archive, and Stanford University reached a similar conclusion, estimating about 35 percent of newly published websites as fully or partly AI-generated, which lends the trend independent corroboration.

Who it affects

Anyone searching or reading the open web crosses paths with this shift, since commercial .com sites show AI-generated text at roughly ten times the rate of .edu or .gov sites. Corporate communications teams are affected too: a separate study cited in the piece found the "it's not just X, it's Y" construction quadrupled in PR documents since 2022. Workers who use AI writing tools face a related consequence: the piece notes that using AI tools already carries a workplace stigma that cuts both ways, as documented by workplace studies it cites.

How to use it

The study offers a rough, evidence-based way to read the web more skeptically: pages on commercial domains, and text using markers like unusual em-dash frequency, Oxford commas, or clustered words such as "delve," "tapestry," and "meticulous," are statistically more likely to be AI-assisted. Pew's own numbers should still be treated as directional rather than exact, since Open Pangram and similar tools can only flag likely AI involvement, not confirm it, and cannot say whether a page was fully generated or just lightly edited by AI.

How solid is it

The headline numbers come from a large sample, nearly half a million pages pulled from the Common Crawl archive, and are corroborated by an independently conducted study from Imperial College London, the Internet Archive, and Stanford University that put the AI-generated share of new websites at a comparable roughly 35 percent. Both findings rest on automated detectors rather than confirmed ground truth, and the article is explicit that current tools can barely distinguish fully automated text from writing that was only partly AI-assisted, which caps how precise any single percentage can be taken to be.

Risks and caveats

The article stresses that nobody agrees on what "AI text" actually means, since the label covers everything from fully machine-written pages to human drafts lightly polished with AI, and detectors like Open Pangram still misfire regularly and cannot say how much AI was involved or at what stage. The article does not name any individual researcher behind the Pew study, though it is itself bylined to Matthias Bastian, and it does not define the threshold Open Pangram or Pew's analysis used to count a page as showing "signs" of AI authorship. Researchers from Imperial College London, the Internet Archive, and Stanford University separately cautioned that public perception of AI text's negative effects often goes well beyond what their data actually supports, and the growing polarization around the topic, sharpened by debates like Anthropic's planned Claude watermark, leaves little room for the mixed reality of how people actually write with these tools.