Dust pretrains transformers without backprop, authors report
Samip Dahal, Bishwas Mandal, Serdar Gülbahar and Akshay Vegesna have published a 2026 research write-up on qlabs.sh describing Dust, a way to pretrain transformer language models without backpropagation. They call it the first zeroth-order method that is competitive with backprop at this task.
Dust is a zeroth-order optimization algorithm. It perturbs activations, rewards each perturbation by how much it lowers the loss, and averages the reward-weighted perturbations over a population to estimate the gradient. In practice, the authors add Gaussian noise to the output of each linear layer, independently at every token, run a forward pass, and reward each token's noise by the change in loss at that token. The reward-weighted noise, averaged over draws, is the estimated error at the layer's output, and its outer product with the layer's input gives the weight gradient. Attention internals get a variant: they are credited through the estimated error at the attention output over current and future tokens, rather than through the tokens' loss directly. Every weight is trained this way except the 2L residual mixing scalars, which are trained by ordinary weight-space evolution strategies (ES).
The central idea is what the authors call a virtual population. Weight-space ES, including EGGROLL (Sarkar et al., 2025), a state-of-the-art method that uses low-rank weight perturbations, evaluates one population member per forward pass, and each member needs its own perturbed copy of the weights. Dust instead treats each token as a member. A sequence in a transformer has a few thousand tokens, so one forward pass evaluates a few thousand members per sequence instead of one. A member is created by adding noise to a hidden state, which is cheap. The authors say that on a modern transformer a single forward pass evaluates a population at least three orders of magnitude larger than weight-space ES. Perturbing activations rather than weights is known as node perturbation, an older technique the paper builds on.
The TL;DR lists four headline claims. First, Dust approximates backprop closely at large population, meaning substantially more compute, and in multiple settings even exceeds it; the authors say this hints that in a compute-rich regime backprop might be surpassed. Second, from 1M tokens up, Dust is on the order of 10^3 to 10^4 times more efficient than a transformer implementation of EGGROLL, based on the authors' extrapolations. Third, contrary to the widespread belief that zeroth-order methods do not scale to large networks, larger models turned out to be more population-efficient: a 243M-parameter model outperforms a model 120 times smaller at most population sizes. The authors read this as a new view of overparameterization, a larger search space with potentially better geometry. Fourth, Dust's gradient estimates align better with backprop's as population grows, and stay well aligned at every scale tested, up to 1B tokens.
The introduction frames the work through Sutton's bitter lesson: general methods that scale with compute eventually win, as AlphaGo Zero's pure self-play eventually overtook a version bootstrapped on human data. The authors suggest that differentiability and backprop may be good inductive biases when compute is scarce but limit what works when compute is plentiful. They also argue that activations are a more interesting search space than weights, since interpretability work has shown reasoning lives in activations, and cite their earlier paper on decoupling search and learning in neural net training.
The authors are explicit about scope. The goal is to lay foundations for a search-based credit assignment algorithm that is competitive with backprop on a hard task, pretraining transformers. They do not attempt to make it compute-efficient enough to replace backprop today. They also do not train the new kinds of networks it makes accessible, such as nets with an external program in the loop or transformers looped over many steps that backpropagation through time struggles to train; both are left to future work.
Key facts
- Dust perturbs activations with Gaussian noise independently at every token, so each token acts as a population member and one forward pass evaluates a few thousand members per sequence.
- The authors claim it is the first zeroth-order method competitive with backprop for pretraining transformer language models, and say it exceeds backprop in multiple settings at large population.
- From 1M tokens up, Dust is on the order of 10^3 to 10^4 times more efficient than a transformer implementation of EGGROLL, based on the authors' extrapolations.
- Larger models were more population-efficient: a 243M-parameter model beats a model 120 times smaller at most population sizes.
- The authors say they do not try to make Dust compute-efficient enough to replace backprop today.
Why it matters
Backprop is the only credit assignment algorithm capable of training modern neural nets, in the authors' words, and architectures, optimizers and hardware have grown up around it. Dust is pitched as a search-based alternative that leans on brute-force compute instead. The most surprising claim is about scale: zeroth-order methods are widely believed not to scale to large networks, yet the authors find larger models are more population-efficient, not less. They also report that Dust's gradient estimates align better with backprop's as population grows, and the authors hedge that this only hints at a compute-rich regime where backprop might be surpassed.
Who it affects
The immediate audience is researchers working on alternatives to backprop, evolution strategies and zeroth-order optimization. The authors also point to network types that backprop handles poorly, such as nets with an external program in the loop and transformers looped over many steps, as future work. For anyone training models today, the authors say plainly that this is not a replacement for backprop yet.
How to use it
The source gives a method description rather than a usage guide. The recipe: add Gaussian noise to each linear layer's output at every token, run a forward pass, reward each token's noise by the change in loss at that token, and average over draws to estimate the error at the layer's output. The outer product of that estimate with the layer's input is the weight gradient. Attention internals use a variant, and the 2L residual mixing scalars are still trained with ordinary weight-space ES. The authors list avoiding interference between perturbed modules among the implementation details that complete the algorithm.
How solid is it
This is the authors' own write-up. The efficiency gain of 10^3 to 10^4 times over a transformer implementation of EGGROLL is explicitly based on extrapolations, and is an order-of-magnitude figure. The alignment of gradient estimates with backprop was tested at scales up to 1B tokens. The account here covers the summary, introduction and the opening of the method section; the experiments are not covered, so the 'competitive' and 'exceeds backprop' claims appear without benchmark figures. No peer review, independent replication or external reaction is mentioned.
Risks and caveats
The authors state that Dust is not compute-efficient enough to replace backprop today, and that the larger-than-backprop results come at large population, which means substantially more compute. The claim of being first is the authors' own. The suggestion that backprop might be surpassed is phrased as a hint, not a result. The 10^3 to 10^4 efficiency figure rests on extrapolation rather than direct measurement at that scale. The method is also not purely activation-space: the 2L residual mixing scalars still need ordinary weight-space ES. Networks with external programs or looped transformers are untested.
“Strikingly, we find larger models are more population-efficient, not less”
— Dahal, Mandal, Gülbahar and Vegesna, Dust write-up on qlabs.sh