Anthropic's bio-weapons filter was down for nearly a year, exposing 133 million chats

Anthropic's bio-weapons filter was down for nearly a year, exposing 133 million chats

Anthropic has disclosed, in a safety report, that its blocking biological classifiers, the filters designed to stop its AI models from being used to extract dangerous knowledge about chemical or biological weapons, were inactive from May 2025 through April 2026, nearly a year during which one of the company's safety controls was simply off.

The gap covered all traffic from external contractors who provide human feedback on Anthropic's models. About 50,000 people ran roughly 133 million chats with the models during that period, all of it passing through without the blocking classifiers applied. According to Anthropic, this contractor pool had been vetted only by the external staffing vendors that supplied them, and those vendors' screening processes were often insufficient.

Anthropic says its internal investigation into the gap found no evidence of actual misuse. The company has since tightened its requirements for contractors, though it has not detailed what those new requirements involve, and the source gives no cause for why the classifiers went dark in the first place or when the gap was discovered.

The disclosure comes from a company whose own chief executive has called AI-assisted development of chemical and biological weapons a bigger threat than cyberattacks, which is the backdrop against which Anthropic runs this kind of filter at all. Separately, and more recently, Anthropic loosened its classifiers on Fable 5 after researchers complained the filters were so aggressive they were blocking legitimate research; the source treats that as a distinct, later adjustment rather than a consequence of the contractor gap.

Key facts

  • Anthropic's blocking biological classifiers, meant to stop its AI models being used to extract dangerous chemical or biological weapons knowledge, were inactive from May 2025 through April 2026.
  • During that gap, about 50,000 external contractors who provide human feedback on Anthropic's models ran roughly 133 million chats without the classifiers applied.
  • Anthropic says the affected contractors had been vetted only by external staffing vendors, whose screening processes were often insufficient.
  • Anthropic's internal investigation found no evidence of actual misuse during the gap, and the company has since tightened its contractor requirements.
  • Separately, Anthropic recently loosened its classifiers on Fable 5 after researchers said the filters were too aggressive and were blocking legitimate research.

Why it matters

Anthropic runs blocking classifiers specifically to stop its models being turned into a tool for extracting dangerous chemical or biological weapons knowledge, a risk the company's own chief executive has called bigger than cyberattacks. That one of these controls sat inactive for nearly a year, across a pool of well over a hundred million chats, is a concrete case of a safety mechanism failing silently rather than being caught in real time. The gap surfaced only because Anthropic chose to disclose it in a safety report, not because it was detected as it happened.

Who it affects

Directly, this affects the roughly 50,000 external contractors who provide human feedback on Anthropic's models, whose combined 133 million chats during the gap ran with no blocking classifier in place. Anthropic says it vetted this pool only through external staffing vendors, and that those vendors' screening was often insufficient, which matters because it defines exactly who had unfiltered access during the period. More broadly, it affects anyone who relies on Anthropic's stated safety controls for chemical and biological weapons knowledge, since the incident shows those controls can lapse for months without being noticed.

How to use it

There is no product here to use; the practical takeaway is for anyone weighing an AI vendor's safety claims. Anthropic's own disclosure shows that a named safety control, the classifiers blocking chemical and biological weapons queries, can sit inactive for close to a year without being caught internally, so a vendor listing a safety measure is not the same as evidence that the measure is running continuously. Anthropic says it has since tightened its contractor vetting requirements, though it has not detailed what those new requirements involve.

How solid is it

This account rests on Anthropic's own safety report, so the core facts, the May 2025 to April 2026 dates, the 50,000 contractors, the 133 million chats, come directly from the company describing its own lapse. The finding of no evidence of actual misuse is from Anthropic's internal investigation rather than an independent audit, and the source gives no detail on how that investigation was run, what caused the classifiers to go dark, or when within the year the gap was actually discovered. Those gaps in the account do not undercut the core disclosure, since it is the company reporting its own failure, but they do limit how much can be verified beyond what Anthropic chose to publish.

Risks and caveats

No evidence of misuse is Anthropic's own finding from its internal investigation, not an independent confirmation that no misuse occurred, and the source gives no cause for why the classifiers were off in the first place, whether that was a bug, a misconfiguration, or a deliberate choice. The gap applied only to traffic from external human-feedback contractors, not to all of Anthropic's users, and the 133 million figure counts chats, not people; the affected population is the roughly 50,000 contractors. The source does not quantify how insufficient the external vendors' screening was, name any vendor, or specify what Anthropic's tightened contractor requirements now consist of. The separate, more recent loosening of classifiers on Fable 5, after researchers said they were blocking legitimate research, is not tied by the source to this contractor gap.