Refusal steering in four open-weight LLMs traces to sparse components
Activation steering changes a large language model's behavior by intervening on its internal activations. The paper starts from the observation that the mechanistic basis of these interventions remains poorly understood, and it takes refusal as the test case.
The authors decompose refusal steering into component-level interventions across four open-weight models. The goal is to identify the sparse subsets of attention and MLP components whose steering is enough to reproduce the full behavioral effect.
They report two layers of sparsity. First, refusal directions concentrate in sparse component mechanisms that make up 28 to 48% of upstream components (the range runs across the four models). Steering only those components retains 88 to 101% of the steering effectiveness.
Second, within these component mechanisms, effective steering concentrates further in approximately 50% of residual stream dimensions. Steering just those dimensions retains 85 to 98% of the component-mechanism baseline, which is a different reference point from full steering. The authors describe this as consistent with a privileged basis structure.
In the authors' words, sparsity operates at two levels: which components are steered, and which dimensions within those components carry the signal. Their conclusion is that refusal is not diffusely encoded across a transformer but assembled by a structured, identifiable mechanism, which they present as a foundation for understanding how refusal behaviors are represented and steered.
For reproducibility, they release all code and raw experimental results at github.com/wang-research-lab/Refusal_Mechanisms.
Key facts
- The study decomposes refusal steering into component-level interventions across four open-weight models.
- Refusal directions concentrate in sparse attention and MLP component mechanisms making up 28 to 48% of upstream components, retaining 88 to 101% of steering effectiveness.
- Within those mechanisms, effective steering concentrates further in approximately 50% of residual stream dimensions, retaining 85 to 98% of the component-mechanism baseline.
- The authors say this is consistent with a privileged basis structure and conclude that refusal is assembled by a structured, identifiable mechanism rather than diffusely encoded.
- All code and raw experimental results are released on GitHub.
Why it matters
Activation steering is a way to change what a model does by editing its internal activations, yet the paper says the mechanistic basis of these interventions remains poorly understood. This work narrows that gap for one behavior, refusal. The central finding is that steering does not need the whole network: a sparse set of attention and MLP components, 28 to 48% of upstream components depending on the model, reproduces 88 to 101% of the steering effect. Inside those components, about half of the residual stream dimensions carry most of the signal. The authors read this as evidence that refusal is assembled by an identifiable mechanism, not spread diffusely through the transformer.
Who it affects
Mainly researchers working on mechanistic interpretability and activation steering, who get a component-level and dimension-level map of where refusal steering acts in four open-weight models. Anyone who studies how refusal behaviors are represented in transformers has a concrete starting point and the released code to build on.
How to use it
The authors release all code and raw experimental results at https://github.com/wang-research-lab/Refusal_Mechanisms. Researchers can use the repository to reproduce the component decomposition and to inspect the raw results behind the reported ranges.
How solid is it
The claims come from the authors' own experiments on four open-weight models, and the release of code and raw results lets others check them. The reported figures are ranges across models: 28 to 48% of components, 88 to 101% of steering effectiveness, about 50% of dimensions, 85 to 98% of the component-mechanism baseline. The authors say the dimension result is consistent with a privileged basis structure, which is weaker than saying it proves one. The source gives no venue or peer-review status, and the four models are not named.
Risks and caveats
The retained-effectiveness figure reaches 101%, above 100% for some cases, and the source does not explain this. It also does not say how steering effectiveness is measured, or which models sit at the low or high end of each range. The 85 to 98% figure is measured against the component-mechanism baseline, not against full steering, so the two percentages should not be multiplied or compared directly. No claim is made about safety implications, jailbreaks, or removing refusal, so none should be read into the results.