Pangram CTO argues post-training guardrails make LLM text detectable

Pangram CTO argues post-training guardrails make LLM text detectable

Bradley Emi, chief technology officer of the AI text detector Pangram, argues in a blog post that large language models could in principle write with as much stylistic variety as humans do, but in practice they don't. The reason, in his account, is not a technical ceiling on the models themselves: it is the post-training and safety guardrails layered on top of them.

Systems such as ChatGPT, Claude and Gemini are trained on behavioral rules that stop them from producing dangerous content or certain political statements. According to Emi, that training sharply narrows how varied the models' output can be, an effect he calls mode collapse, and it is precisely that narrowing which makes their text stand out to a detector like Pangram's.

The pattern reverses for text that was never shaped by that guardrail training. So-called base models, the versions of a system that exist before post-training, write with more variety, and Pangram's detector does not flag them, Emi says. He reports the same result for narrowly specialized fine-tunes, such as ones trained only on Hemingway's prose or on a single subreddit's writing, and for outputs that are simply broken or incoherent.

Emi scopes the claim to one exception: it holds only for text that carries no watermark. Watermarked AI text, he says, will likely remain detectable regardless, even from a base model with its greater variety.

Key facts

  • Bradley Emi, CTO of AI text detector Pangram, argues in a blog post that post-training and safety guardrails, not an inherent limit on the models, are what make LLM text detectable.
  • Systems such as ChatGPT, Claude and Gemini learn behavioral rules against dangerous or certain political content, which narrows their expressive range, an effect Emi calls mode collapse.
  • Base models, the versions before post-training, write with more variety and are not flagged by Pangram's detector, Emi says; the same holds for fine-tunes narrowed to one author's or one subreddit's style, and for broken, incoherent output.
  • The exception is watermarked text: per Emi, watermarks will likely remain detectable regardless of a model's stylistic variety.
  • No accuracy figures, sample sizes, or a link and date for the blog post are included in this account.

Why it matters

The claim reframes a debate usually treated as settled: whether AI-written text is detectable at all. Emi's argument is that detectability is not a property of language models as such, it is a side effect of the safety alignment layered on top of them. Because the claim comes from the CTO of a company that sells AI-text detection, the admission that some outputs, base models, narrow fine-tunes, broken text, slip past that detection reads as a disclosure against the company's own commercial interest, which is part of what gives it weight.

Who it affects

Anyone who relies on AI-text detectors, in classrooms, publishing or content moderation, is affected by the gap Emi describes: standard post-trained chatbots such as ChatGPT, Claude and Gemini get flagged, while base models, fine-tunes narrowed to one author's or one community's style, and broken or incoherent text do not. The companies that build and align those chatbots, and Pangram itself as the detector vendor making the disclosure, are the other parties with a direct stake in the claim.

How to use it

There is no product angle to act on here, no pricing or access details are given, but the practical takeaway is a caution: a clean result from an AI-text detector should not be read as proof a text is human written, since Emi's own account leaves base models, narrow fine-tunes and broken text able to pass through undetected. Where the underlying text carries a watermark, per Emi, detection stays reliable regardless of how the text was generated.

How solid is it

The claim rests on one source: a blog post by Pangram's own CTO, without a link, publication date, or supporting evidence included in this account. No accuracy figures, sample sizes or error rates for Pangram's detector are given here, the specific behavioral rules and censored political statements are not named, and no mechanism is offered for why watermarks would keep working despite a base model's greater variety. It is one company insider's characterization, not an independently verified study, and this account gives no way to check it further.

Risks and caveats

The account gives no numbers to weigh the risk against: how often a base model, a narrow fine-tune, or broken output evades detection in practice, or how available such models are outside research and detection labs, is not stated. The claim that watermarks will likely always work is presented as an expectation, not a proven result, with no mechanism given for why a base model's greater stylistic variety would fail to defeat a watermark. And the claim comes from the head of a company whose product is the very detector under discussion, which does not make it wrong, but is a reason to weigh it as one insider's account rather than as settled fact.