NEEDLE removes LLM backdoors without retraining, reaching 0% attack success on code injection

Backdoor attacks can be planted in large language models during training. The model then behaves badly whenever a trigger shows up in the input. The paper starts from a problem with existing defences: they try to remove the backdoor, but they inadvertently shift the model's output distribution on benign prompts, which can degrade both performance and safety.
The authors propose NEEDLE, a training-free method for targeted backdoor removal. It works once a trigger has been identified. From activation vectors, the method estimates two things: a backdoor direction and a refusal subspace. It then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations. In other words, the backdoor is edited out of the weights, and the part of the model that handles refusals is protected from the edit.
NEEDLE requires neither a clean reference model nor the original poisoned training data. Evaluation covers multiple model families and attack types. According to the abstract, NEEDLE achieves the lowest mean Attack Success Rate (ASR) among the evaluated defences, including 0% on challenging code injection attacks. It also results in the lowest KL divergence and minimal changes in capability and safety. The abstract gives no other numeric results.
Key facts
- NEEDLE is a training-free method for targeted backdoor removal in LLMs, applied once a trigger has been identified.
- It estimates a backdoor direction and a refusal subspace from activation vectors, then applies sequential weight orthogonalisation.
- It needs neither a clean reference model nor the original poisoned training data.
- The authors report the lowest mean ASR among the evaluated defences, including 0% on challenging code injection attacks.
- They also report the lowest KL divergence and minimal changes in capability and safety.
Why it matters
Existing backdoor defences for LLMs can remove the backdoor yet shift the model's output on benign prompts, which can hurt performance and safety. NEEDLE is built to avoid that side effect by keeping refusal-related representations unchanged while it suppresses the backdoor. The reported lowest KL divergence points to the same goal: leaving the model's normal behaviour as close to untouched as possible.
Who it affects
The work is aimed at anyone who runs a language model that may have been backdoored during training. The method does not need a clean reference model or the original poisoned training data, so it is framed for cases where neither is at hand.
How to use it
The method applies once a trigger has been identified. It then estimates a backdoor direction and a refusal subspace through activation vectors and applies sequential weight orthogonalisation to the model. No retraining is involved, since the method is training-free.
How solid is it
The evidence here is the paper's own abstract. The authors report that evaluation spans multiple model families and attack types, and that NEEDLE has the lowest mean ASR among the evaluated defences, including 0% on challenging code injection attacks. Only that one figure is given as a number, so the size of the gap to other defences cannot be judged from this text.
Risks and caveats
The method starts from a known trigger, so it addresses removal rather than detection. The abstract's results are the authors' own claims, and the specific model families, attack types other than code injection and the compared defences are not named in it. The 0% figure applies to code injection attacks; the other results are described only as lowest mean ASR, lowest KL divergence and minimal changes in capability and safety.