Paper splits causal influence in language models into capacity, responsiveness and alignment

Researchers often locate latent structures in the activation space of language models, because doing so is central to understanding and controlling model behavior. But the structures they find differ a great deal in how much causal influence they have. The paper asks what makes a structure actionable.

Its answer is to cast causal influence as a product of three factors, which the authors show empirically to be interpretable and distinct constraints. Capacity measures the sensitivity of the model's output to movement along the structure. Responsiveness captures how promotable the concept is given the current context. Alignment reflects how well the structure aligns with the context-specific representation of the concept.

The authors tested this across 4 LM families and 50 concepts. They observe that causal effectiveness requires all three factors to be high. Low capacity reduces causal effectiveness by 84%, and low responsiveness reduces it by 95%. Low alignment can do worse than reduce it: it can reverse the effect, suppressing concept expression instead of promoting it.

The paper also finds that causality is context-dependent rather than an intrinsic property of the structure. Causally effective directions form a low-dimensional subspace, and that subspace varies across contexts.

Building on this, the authors restrict the training of linear probes to that subspace and call the result causal probes. Across models, causal probes achieve a 17%-118% improvement in steering, with only a 3% reduction in concept detection.

Key facts

  • The paper casts causal influence of a latent structure in a language model as a product of three factors: capacity, responsiveness and alignment.
  • Tested across 4 LM families and 50 concepts, causal effectiveness required all three factors to be high; low capacity cut it by 84% and low responsiveness by 95%.
  • Low alignment can reverse the effect and suppress concept expression.
  • Causality is context-dependent: causally effective directions form a low-dimensional subspace that varies across contexts.
  • Causal probes, linear probes trained only within that subspace, improve steering by 17%-118% across models at a cost of only 3% in concept detection.

Why it matters

Finding a direction in a model's activation space that corresponds to a concept does not mean you can use it to change the model's behavior. The paper tackles that gap. It offers a three-part account of why some localized structures steer a model and others do not: how sensitive the output is to movement along the structure (capacity), how promotable the concept is in the current context (responsiveness), and how well the structure matches the context-specific representation of the concept (alignment). Its finding that causality depends on context, not on the structure alone, challenges the idea that a located direction has a fixed causal role.

Who it affects

The work speaks to researchers who localize latent structures in language models in order to understand or control them, especially those who steer model outputs by moving along a direction and those who use linear probes to detect concepts. The abstract does not name the 4 LM families or the 50 concepts, so the exact scope of the models covered is not stated.

How to use it

The practical takeaway in the abstract is the method itself: restrict the training of linear probes to the low-dimensional subspace of causally effective directions, which gives causal probes. Because that subspace varies across contexts, the paper's findings point to a context-aware approach rather than one fixed direction per concept. The abstract mentions no code release, dataset or submission date, so anyone wanting to apply the method would need the full paper for the details.

How solid is it

The evidence is an empirical study across 4 LM families and 50 concepts, and the claims here come from the paper's abstract. The figures are the authors' own: 84% and 95% reductions in causal effectiveness for low capacity and low responsiveness, and a 17%-118% steering improvement with a 3% reduction in concept detection. The abstract does not say which model gets the 17% and which the 118%, does not name the baseline probes causal probes are compared against, and does not specify the steering or detection metrics. The baseline for the percentages is also not specified. The paper is an arXiv preprint.

Risks and caveats

The gains vary widely across models, from 17% to 118%, so the benefit of causal probes may differ a lot from one model to the next. The 3% reduction in concept detection is a trade-off the authors report as small, and the abstract gives no further detail on how it was measured. Low alignment can reverse a structure's effect, so a direction that looks useful in one context can suppress the concept it is meant to promote. The abstract makes no claim about practical deployment or safety applications.

“we find that causality is context-dependent rather than an intrinsic property of the structure”

— From the paper's abstract