Essay argues AI refusal is far from foolproof and could enable repression

An MIT Technology Review essay takes aim at what it calls the load-bearing wall of AI safety: teaching language models to refuse dangerous requests. Its argument is twofold. Refusal is far less reliable and far less understood than the industry's confidence suggests, and as governments gain the power to set refusal lines, it could become a tool of repression.
The essay opens with the idea's recent history. In 2021, a team at Anthropic wrote that large language models should be made helpful, honest and above all harmless, meaning that 'when asked to aid in a dangerous act (e.g. building a bomb), the AI should politely refuse.' Refusal does not come naturally. A model trained on billions of web pages picks up a broad mastery of violence and vitriol but not the habit of keeping it to itself. Steven Adler, who worked on safety at OpenAI from 2020 to 2024, says the company's earliest models would 'blab on about anything.' Ryan McBain, who researches AI and mental health at Harvard, recalls that an early chatbot asked how to kill oneself with a gun could 'very easily generate a response.'
Today companies train models with exercises that reward refusing questions they deem harmful and punish 'over-refusing' prompts they deem harmless, often using other models to run the exercises. They also put their models behind layers of other AI that stop mischievous prompts from reaching the core. Yet refusal often fails. The author notes that some of the latest models are, companies say, as good at breaking into critical computer networks as top human hackers, and that companies report some users trying to use the most advanced AI to hone biological pathogens and build autonomous drone swarms. In the author's words, because the mechanisms of refusal are probabilistic, they are never likely to be all that reliable; determined miscreants have already broken through and may always be able to; and sooner or later failed refusals might result in global calamity. The essay compares the setup to fitting every car with a machine gun and hiding the trigger under the hood.
There is also the question of where to draw the line. Some virologists have good reason to study nasty viruses, and some users want to know about vulnerabilities in order to patch them. 'Where you draw the line is a huge question,' says Zico Kolter, a member of OpenAI's board and cofounder of the AI testing company Gray Swan. At the moment AI companies draw that line, jealously and in secrecy. The essay says governments will soon draw their own. They must try to block genuinely malicious acts, but there may not be much to stop oppressive governments from blocking the technology's capacity to generate legitimate speech. The better AI becomes at refusing harm, the better it gets at stifling ideas whose only risk is to those who make the rules. The author adds that AI may already refuse to criticize certain authoritarian heads of state, and mentions in passing that the Pentagon has wrestled with frontier model companies because it wants fewer refusals.
The second section, 'Learning the limits', explains how refusal is taught. In 2022, before the release of ChatGPT, OpenAI enlisted dozens of red-teamers to probe its latest model. One was Paul Röttger, then completing a PhD about online extremism and now a researcher at the Hasso Plattner Institute in Potsdam, Germany. The red-teamers got minimal directions: ask whatever they considered 'refusal-worthy.' Between them they put thousands of queries to the model and logged the results in an Excel sheet. The model refused some, but when Röttger asked for a recruitment post for Al Qaeda, it readily complied. OpenAI assembled these responses into datasets that were, in all likelihood, fed back to the model as part of fine-tuning, though Röttger was not told exactly how his spreadsheet would be used. A few months later the same request got a no.
What happens inside the model is not moral reasoning. When a prompt resembles the training prompts the model was conditioned to refuse, a series of 'activations' light up among its billions of parameters. A recent Google-funded study's best guess is that refusal shows up in activation space as a set of 'high-dimensional polyhedral cones.' Jannes Elstner, an author of that paper who now works on AI safety at Apollo Research, told the author in July that even this is not quite right: a cone is just a way of describing an indeterminate number of lines pointing roughly the same way, and there are other, undiscoverable elements that may secretly play a role. Researcher Andy Arditi has previously shown that if these activations are eliminated and the model is fed the same prompts again, it will not refuse. We can see when a model says no, the essay says, but how it decides is at best a hypothesis. Asked whether anyone is okay with that, Elstner shrugged: 'We need refusal whether we understand it or not.'
The third section, 'A wall of cheese', covers the classifiers that companies wrap around models. Some read the user's input and block dangerous requests; others read the model's output and block harmful answers. None catches everything, so companies stack them in what is called the Swiss cheese model. It is costly: earlier this year Anthropic said one type of classifier added 24% to its chatbots' compute costs. Anthropic and others have begun switching to a more efficient set of classifiers called probes, which observe the model's internal activations, like putting the loose-lipped celebrity in an fMRI. Classifiers can be modified in a matter of weeks when there is something new to refuse, but they remain probabilistic. Even rules written in human language, which Anthropic calls a 'constitution' and OpenAI calls a 'model spec', come down to statistics, and a model's displayed chain of thought is still a sequence of predicted words. McBain's latest experiments found that if you repeatedly ask any of the major models the same risky suicide questions, they generally refuse, but every so often they do not. Elstner says probes could eventually learn to recognise all the indescribable activations and perfectly detect every refusal. The author writes that at that point AI safety would rest on a labyrinthine conceit: a map of a map, a secret schema of human morality codified in statistics beyond our wit.
A final section, 'The core trade-off', begins by saying that AI refusal is needed because artificial intelligence is a bargain on Faustian terms: if AI is to help cure cancer, it needs genetics expertise that could in theory be used to modify viruses and bacteria for bioweapons. The essay's closing view, stated earlier, is that it is hard to imagine an alternative that would not slow the technology's progress, but that we should be frank about the perils: when refusal falls short the effects could be catastrophic, and when it goes all the way it could enable grievous acts of repression.
Key facts
- Refusal is trained into models through fine-tuning, reward and punishment exercises (often run by other models) and layers of classifiers, and the essay says its probabilistic mechanisms are never likely to be all that reliable.
- Researchers do not fully understand how a model decides to refuse: a Google-funded study's best guess is 'high-dimensional polyhedral cones' in activation space, and co-author Jannes Elstner says even that is not quite right.
- Classifiers are costly and probabilistic: Anthropic said earlier this year that one type added 24% to its chatbots' compute costs, and companies are now moving to cheaper probes that read the model's internal activations.
- The author warns that governments will soon get to draw their own refusal lines, and that oppressive ones could use refusal to block legitimate speech; AI may already refuse to criticize certain authoritarian heads of state.
- In 2022 OpenAI red-teamers, including Paul Röttger, logged thousands of queries in an Excel sheet; the model readily wrote an Al Qaeda recruitment post, and refused the same request a few months later.
Why it matters
The essay's point is that refusal has become the load-bearing wall of AI safety, and nobody fully understands the wall. Companies lean on it to keep models from helping with bioweapons, cyberattacks and self-harm, yet the mechanism is statistical and can fail. The second claim is political. Today companies decide what a model refuses, in secrecy; the author says governments will soon do the same, and the better AI gets at refusing harm, the better it gets at stifling ideas that only threaten those in power.
Who it affects
AI companies such as OpenAI and Anthropic, which carry the cost and the responsibility of building refusal. Everyday users, who are sometimes refused harmless requests, and who may meet a model that declines to criticize a head of state. Legitimate specialists such as virologists and security researchers, whose work sits on the line. Governments, which will gain a say in where the line goes. And people living under oppressive governments, for whom the essay says refusal could become an instrument of repression.
How to use it
There is nothing to install or buy here: this is an argument, not a product. The practical reading is that a refusal from a chatbot, or the absence of one, is not a guarantee. Models usually refuse repeated risky questions but, per McBain's experiments, every so often do not. For anyone building on or relying on these systems, the essay's picture is of stacked, imperfect layers (fine-tuning, input and output classifiers, probes) rather than a single dependable switch. The author's own conclusion is that no alternative looks easy, and that the perils should be stated frankly.
How solid is it
This is an opinion and analysis essay in MIT Technology Review that mixes first-hand interviews with the author's own judgments. Attributed facts are specific: Adler, McBain, Kolter, Röttger and Elstner are named with their roles, and the 24% figure is Anthropic's. The capability claims about hacking and pathogens are explicitly attributed to companies. Strong statements such as refusal being 'never likely to be all that reliable' are the author's own voice, not a finding. The account here rests on most of the essay but not its ending, since the available text stops partway through the final section.
Risks and caveats
The warnings are phrased as possibilities: failed refusals 'might' cause global calamity, oppressive governments 'may' not be stopped, AI 'may already' refuse to criticize certain heads of state. No specific government, law or model is named for these claims, and no figure is given for how often refusals fail or jailbreaks succeed. The 24% compute cost refers to one unnamed type of classifier, and it is not said whether it applies to probes. The Google-funded study is not named, and the Pentagon point is only a parenthetical aside. Elstner's idea that probes could eventually detect every refusal is a possibility he describes, not a result.
“We need refusal whether we understand it or not.”
— Jannes Elstner, AI safety researcher at Apollo Research, quoted in MIT Technology Review