AI Observatory finds Anthropic's usage filter excludes half of chats

AI Observatory finds Anthropic's usage filter excludes half of chats

AI companies including Anthropic and OpenAI regularly publish reports on how people use products like Claude and ChatGPT, but they only release the data they choose to show, with no independent way to check it, says Anka Reuel, a computer science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab. "There is no independent source to corroborate it," she says. Reuel co-leads a new project called the AI Observatory, built with Shayne Longpre, a recent PhD graduate from the MIT Media Lab, along with researchers from MIT, Stanford, the Data Provenance Initiative and other institutions. It is a public platform that aggregates and analyzes real AI conversations with popular models such as Claude and Gemini, collected with users' consent through seven existing datasets, so that researchers and policymakers have an independent source of information on how people actually use generative AI. Reuel says highly consequential decisions about AI's benefits and risks are currently being made on the basis of very limited data.

The Observatory aggregated 85,633 conversational turns, meaning a user prompt plus the AI's response, across 24,521 conversations from 5,000 users interacting with 52 different models between 2023 and 2025. The researchers then applied the methodology behind the Anthropic Economic Index, one of the best known and most widely cited sources of AI usage data, which focuses on work and productivity uses of Claude and filters out conversations unrelated to those uses, to their own broader dataset. They found that nearly half of it, 48%, would have been filtered out as non-work-related. Those non-work conversations skewed heavily toward topics that Anthropic's own published breakdown shows far less of: health and relationships made up 44.2% of them versus 31.2% in Anthropic's analysis, adult or illicit topics 7.9% versus 2.1%, harassment and hate 27.5% versus 5.66%, and sexual content 16.7% versus 2.4%. The piece notes a parallel figure from OpenAI's own 2025 report on ChatGPT, which found that only 30% of consumer use was related to work.

David Widder, an assistant professor at the University of Texas at Austin who studies how people interact with AI systems and is not involved with the Observatory, says Anthropic has published separate blog posts on how people use Claude for support or companionship, and even to generate CSAM, but that having a single "bird's-eye-view" analysis rather than that information "sectioned off into a separate report" helps researchers understand real AI use more consistently. Longpre puts the same point more bluntly: "No single company report tells the whole story."

Looking at WildChat, one of the largest and most detailed datasets in its study, the Observatory found conversations grew longer and more elaborate over time, with rising numbers of prompt tokens, response tokens and conversation turns, alongside significantly more small talk, which the researchers read as a sign that AI companionship was increasing even as the assistants' own self-disclosure, admitting to being a chatbot, decreased. At the same time, exchanges the researchers labeled sensitive, meaning ones with potentially harmful or restricted content such as sexual harassment and hate speech, became less frequent over time, which they say might suggest platforms were generally deploying more effective safeguards. Usage also varied sharply by model: people turned to Grok and Gemini more often for information retrieval, with Grok especially popular for news and politics but also, consistent with other research on the platform, where misinformation tended to concentrate (xAI did not respond to a request for comment); people leaned on Anthropic's models for coding, Gemini for social and roleplay uses, and ChatGPT for homework help. Even different versions of the same model differed: people had shorter conversations with ChatGPT when it ran on GPT-3.5, and longer, more iterative ones on GPT-4o, a version the piece notes became known for fostering emotional addiction.

The Observatory's own dataset is far smaller than what the big labs hold internally: Anthropic's latest Economic Index is built on 1 million Claude conversations and OpenAI's ChatGPT usage report on 1.5 million, against the Observatory's 24,521. Because its underlying datasets were collected only from users who consented to share them, the researchers caution that the Observatory's data likely underrepresents sensitive uses that people are less willing to hand over for research, and they say its findings are not indicative of all AI use. An Anthropic representative said the company's published research reflects its research teams' own questions and interests and that it is important to support external independent research; OpenAI did not respond to requests for comment. The Observatory's data will be available to researchers for further analysis, and the team hopes to keep expanding its datasets over time. Reuel's stated ideal is that AI companies would go further and share their own data with independent researchers in ways that protect user privacy; short of that, she says, anyone making decisions based on AI usage data risks "completely operating in the wild and making these really consequential decisions without knowing what's actually happening beyond those company narratives."

Key facts

  • The AI Observatory, co-led by Stanford's Anka Reuel and MIT's Shayne Longpre, aggregated 85,633 conversational turns across 24,521 conversations from 5,000 users and 52 models, drawn from seven existing consented datasets spanning 2023 to 2025.
  • Applying the Anthropic Economic Index's own filtering methodology to that dataset, researchers found 48% of conversations would have been excluded as non-work-related.
  • Non-work conversations skewed heavily toward health and relationships (44.2% versus Anthropic's reported 31.2%), harassment and hate (27.5% versus 5.66%), sexual content (16.7% versus 2.4%), and adult or illicit topics (7.9% versus 2.1%); OpenAI's own 2025 report separately found only 30% of consumer ChatGPT use was work-related.
  • Usage diverged sharply by model and version: people leaned on Anthropic for coding, Gemini for social and roleplay uses, ChatGPT for homework, and Grok and Gemini for information retrieval, with misinformation concentrating on Grok, while ChatGPT conversations ran longer under GPT-4o than under GPT-3.5.
  • The Observatory's sample is a fraction of what the big labs hold internally, 24,521 conversations against Anthropic's 1 million and OpenAI's 1.5 million, and its researchers caution the data likely underrepresents sensitive use and is not indicative of all AI use.

Why it matters

Usage reports from Anthropic and OpenAI are the main public lens on how people actually use chatbots like Claude and ChatGPT, but those reports are built entirely from data the companies choose to publish, with no outside way to check them. The AI Observatory was built to close that gap by aggregating real conversations from independent, consented sources rather than relying on what companies decide to disclose. Its first results suggest the published reports capture only part of the picture: applying the Anthropic Economic Index's own filtering method, built to isolate work and productivity use of Claude, to the Observatory's broader dataset excluded 48% of conversations as non-work. Reuel says highly consequential decisions about AI's benefits and risks are currently being made on the basis of very limited, company-controlled data.

Who it affects

The findings bear most directly on Anthropic and OpenAI, whose usage reports are widely cited as authoritative pictures of how people use AI, and on xAI and Google, whose Grok and Gemini models were also covered by the Observatory's dataset. They matter just as much to the researchers and policymakers who lean on those company reports to judge AI's risks and benefits, since the Observatory found meaningfully higher rates of sensitive material, such as health and relationship talk, harassment and hate, and sexual content, among non-work conversations than Anthropic's own published breakdown shows. Indirectly, the findings concern anyone using these chatbots, since it is their conversations, aggregated with consent through the seven source datasets, that make up the Observatory's evidence base.

How to use it

The AI Observatory is built as a shared resource rather than a product: its aggregated data will be made available to researchers for further analysis, and the team behind it, led by Reuel with researchers from MIT, Stanford, the Data Provenance Initiative and other institutions, plans to keep expanding the datasets it draws on over time. Reuel's stated ideal is that AI companies go further and share their own conversation data directly with independent researchers, in ways that protect user privacy. Short of that, she argues, anyone basing decisions on AI usage data is working without a way to check it against the companies' own narratives.

How solid is it

The Observatory's evidence base is real but small next to what the AI companies hold internally: 85,633 conversational turns across 24,521 conversations from 5,000 users and 52 models, pulled from seven existing datasets, including WildChat, collected with consent between 2023 and 2025. Anthropic's latest Economic Index, by contrast, is built on 1 million Claude conversations, and OpenAI's 2025 usage report on 1.5 million ChatGPT conversations. The Observatory's finding that people used Grok and Gemini most for information retrieval, with misinformation concentrating on Grok, tracks with other independent research on Grok's misinformation problem, which strengthens that particular result. The 48% figure also measures something specific: it is what Anthropic's published filtering method would exclude when run against the Observatory's own broader, multi-model dataset, not a figure Anthropic reported about its own underlying data. Anthropic told the publication its research reflects its teams' own questions and interests and that it supports external independent research; xAI and OpenAI did not respond to requests for comment.

Risks and caveats

Because the underlying datasets were collected only from users who consented to share them, the Observatory's own researchers caution that its data likely underrepresents sensitive use, since people are less willing to hand over their most sensitive conversations for research than an ordinary one, and they say its findings are not indicative of all AI use. The source does not specify, beyond the four named categories, what criteria the Observatory used to classify a conversation as sensitive, which limits how precisely its numbers can be compared with any single company's own definitions. The trends drawn from WildChat, including the growth in small talk, the drop in the assistants' self-disclosure, and the shift between GPT-3.5 and GPT-4o conversation lengths, are also reported only as directional changes, without the underlying figures attached.

“There is no independent source to corroborate it.”

— Anka Reuel, Stanford Trustworthy AI Research (STAIR) Lab, AI Observatory co-lead