Anthropic announces watermark detection API for Claude text

Anthropic announces watermark detection API for Claude text

Anthropic says it will soon offer a watermark detection API that lets third-party developers plug AI-text detection for Claude's output into their own apps. The watermark itself uses a variant of the SynthID Text method Google DeepMind published in Nature in 2024: it tweaks the randomness source Claude uses when picking words, creating a traceable pattern that Anthropic says has no effect on the content, creativity, or readability of the text. The watermark is more reliable on longer, freely worded text and less reliable on short passages, fact-heavy writing, or code, where there are fewer alternative phrasings to choose from. Pure human corrections will not carry it, but translations will, since Claude chooses every word in that case. Heavy rewriting can strip the watermark out entirely, according to a published Anthropic FAQ. The detector can only flag that Claude was likely involved in producing a text; it cannot say whether Claude wrote the whole thing or just edited it, and it cannot tell a human-written text from one produced by a different AI model. Anthropic is rolling watermarking out worldwide because it says there is currently no technical way to limit the feature to a single region. The push follows the company signing the EU Code of Practice on transparency for AI-generated content in July 2026, alongside roughly 190 other signatories, to comply with the EU AI Act. All Claude models released after August 2, 2025 support watermarking out of the box; older models will get it in the coming months. For files, Anthropic uses the open C2PA standard, which attaches metadata without altering the file itself. This differs from external AI-detection tools like Pangram, which lack access to Anthropic's watermarking keys and instead scan text for telltale AI phrasing patterns; checking for an actual watermark is a different approach that the source says should prove more reliable.

Key facts

  • Anthropic will soon launch a watermark detection API so third-party developers can check whether text was produced by Claude.
  • The watermark is a variant of Google DeepMind's SynthID Text method, published in Nature in 2024, and Anthropic says it does not affect content, creativity, or readability.
  • It works less reliably on short or fact-heavy text and on code, and heavy rewriting can strip it out; pure human edits will not carry it.
  • The rollout is worldwide because there is no technical way to limit it by region, driven by the EU AI Act and Anthropic's July 2026 signing of the EU Code of Practice alongside roughly 190 other signatories.
  • All Claude models released after August 2, 2025 support watermarking out of the box; older models get it in the coming months, and files use the open C2PA metadata standard.

Why it matters

This gives outside parties, not just Anthropic, a way to check whether a piece of text came from Claude. It is built on Google DeepMind's SynthID Text approach, first published in Nature in 2024, which alters the randomness in Claude's word choices to leave a traceable pattern rather than bolting on a separate label. The immediate driver is regulatory: Anthropic is one of roughly 190 signatories of the EU Code of Practice on transparency for AI-generated content, and since it says there is no technical way to confine the watermark to the EU, the feature ships globally rather than region by region.

Who it affects

Third-party developers get an API to add Claude-text detection to their own products. Anyone consuming Claude-generated text, including platforms, publishers, and readers trying to judge provenance, gains a check that did not exist before. EU regulators and the roughly 190 other signatories of the Code of Practice are the direct audience the move is aimed at satisfying.

How to use it

The detection API is not yet live; Anthropic says it is coming soon, without giving a firm date, pricing, or access requirements. All Claude models released after August 2, 2025 carry the watermark automatically; older models are due to gain it in the coming months. For files rather than raw text, Anthropic attaches metadata using the open C2PA standard, which does not alter the file itself.

How solid is it

The report comes from a single article with no named Anthropic spokesperson, drawing on Anthropic's own published FAQ and the earlier DeepMind Nature paper. The mechanism is described in technical detail but not quantified: Anthropic gives no accuracy, false-positive, or detection-confidence figures for the watermark, so its real-world reliability cannot be assessed from what has been published so far.

Risks and caveats

The watermark's own limits are significant. It is less reliable on short passages, fact-heavy writing, and code, where Claude has fewer alternative word choices to embed a pattern in. Text a human corrected word by word will not carry the watermark, while translations will, since Claude picks every word there. Heavy rewriting can strip the watermark out. And even a clean detection only says Claude was likely involved; it cannot show whether Claude wrote the whole text or just edited it, and it cannot distinguish a human author from a different AI model altogether.

“This has no effect on the content, level of creativity, or readability of Claude's text.”

— Anthropic