For a linear layer $y_t = W x_t$ we jitter its output at all tokens, $y_t \to y_t + \sigma a_t$ with $a_t \sim \mathcal{N}(0, I)$ and $\sigma$ the noise scale, and run the forward pass. At each token $s$ we calculate the centered loss reduction $c_s = \tilde{\ell}_s - \ell_s$, where $\ell_s$ is the perturbed loss and $\tilde{\ell}_s$ is the mean perturbed loss at that token across draws evaluated together in the same batched forward. The reward of the jitter at token $t$ is the loss reduction at $t$ and, with a decay $\gamma$, the loss reductions at the tokens after it, which the jitter also reaches through attention,
With $\gamma = 0$ a jitter is rewarded by its own token alone. We leave it to the tuning of Section 2.3 to decide which layers see the future tokens. One independent jitter of all tokens is a draw, and a population is $K$ draws. Averaged over the draws, the reward-weighted noise
is the estimated error at the layer’s output, and its outer product with the layer’s input, which the forward pass already computed, summed over tokens, is the weight gradient,
Backprop forms the same outer product with the same input. The only difference is that it gets the output error from the chain rule and we get it from the population. For an embedding layer $x_t$ is one-hot, so the outer product is a scatter-add of $\hat g_t$ into the token’s row.
With unlimited population the estimator above is the whole method, and every perturbation could go into one forward pass. Each draw’s reward-weighted noise is the gradient plus an error with no preferred direction. Over draws the gradient adds up linearly while the errors add up as the square root, so their ratio falls as the population grows and the interference vanishes in the limit.
At a population we can afford, the main cost is interference, i.e. we jitter many layers at many tokens in the same forward pass, so the loss change that rewards one token’s noise also picks up the effect of every other perturbation in that pass. We reduce such interference in three ways. First, different layer types are jittered in separate forward passes, each with its own noise scale, and each block gets its own passes. These passes are cheaper than full forwards because the clean forward is cached and a draw for block $l$ only reruns blocks $l$ onward. Second, the attention internals (query, key, value, gate, value embedding) are jittered separately. Token losses barely register their jitters, so they are rewarded through the attention output instead, as described below. Third, the language modeling head is jittered directly on the cached logits, re-evaluating only the cross-entropy and only a slab of the vocabulary per draw, which costs a small fraction of a forward pass and lets the head run a much larger population.
For the attention internals we recompute only their block’s attention output with the jitter, from the cached clean activations. We score the jitters using the alignment with the estimated attention output gradients,
where $\Delta o_s$ is the change the jitter makes in the attention output at token $s$ and $\hat g_s$ is that output’s estimated error from Equation 2. The reward is Equation 1 with these scores, computed per head, in place of the loss reductions.
The hyperparameters, the noise scale of each layer type, the credit decay of the attention internals and the share of the population each layer gets, can be tuned in two ways. One is a grid search that trains with each setting on a small token budget and keeps the setting that lowers the loss most. It is reliable but expensive. The other is a grid search that maximizes the cosine between our estimate and the backprop gradient on a single batch, which needs no training at all. A larger cosine on one batch doesn’t always lower the loss after training, so the cosine picks candidates and training decides. Either way the tuning is mostly a one-time cost, since the settings it finds largely generalize across token budgets and population sizes, with one exception. At the largest population at 10M and 20M tokens, a slower credit decay and a shift of draws toward the attention side still pay (Appendix F). Hence, this search recovers general principles of the method rather than settings for one run. As expected, all layers require no credits from future tokens except the keys, values, gates and value embeddings, which are read by the later tokens that attend to them and get a $\gamma$ close to one.
We train GPT-style transformers on FineWeb with a 4096-token BPE tokenizer, a batch of 16k tokens (8 sequences of 2048 tokens), one epoch, and SGD with momentum at a constant learning rate. The base model has 8 layers and width 512. Every method gets the same protocol and three seeds per cell, and is tuned separately at every token budget and population. Dust and backprop share one momentum and learning-rate grid; we implement EGGROLL with the same transformer architecture (named EGGROLL-Transformer) and it is tuned over its own grid of step size, momentum, noise scale and fitness shaping. Validation and test are held-out sets of 544 sequences each. We report the test loss at best validation checkpoint.
We count population in draws for Dust and in forward passes of the batch for EGGROLL (Sarkar et al., 2025Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. arXiv preprint arXiv:2511.16652, 2025.), our weight-space baseline. A draw is one jitter of activations across all tokens of selected layers, rewarded with the token losses, and the population $K$ is the number of draws per update. A draw is slightly cheaper than a forward pass, because the clean forward is cached and a draw that jitters block $l$ reruns only the blocks from $l$ on. $K$ leaves out the draws of the head and the attention internals, which are a small fraction of the update’s FLOPs (Appendix E). Taken together, from a population of 256 up Dust uses less compute than EGGROLL at the same population, so the comparison is lenient to EGGROLL.
5.0 6.0 7.0 8.0 Validation loss
10M tokens 0 0.25 0.5 0.75 1 Training fraction
Dust
64
1k
16k EGGROLL- Transformer
64
1k
16k
Backprop
Figure 1: Validation loss over training at 1M and 10M tokens for Dust (violet), EGGROLL at the same populations (gray) and backprop (black, dashed). The x-axis is the fraction of the token budget consumed. Shade is population, 64, 1k and 16k, darker for larger. Curves are means over three seeds.
7.2 7.4 7.6 7.8 8.0 Test loss
100k tokens 64 256 1k 4k 16k Population
4.0 4.5 5.0 5.5 6.0 6.5 7.0 7.5
1M tokens 64 256 1k 4k 16k Population
4.0 4.5 5.0 5.5 6.0 6.5 7.0 7.5 Test loss
10M tokens 64 256 1k 4k 16k Population
4.0 4.5 5.0 5.5 6.0 6.5 7.0 7.5
20M tokens 64 256 1k 4k 16k Population
∞ 4.43
fit 1.53 P−0.145 + 4.43
limit as P → ∞
95% interval of the limit
EGGROLL-Transformer
Dust
Backprop
Figure 2: Test loss at the validation-selected checkpoint against population, one panel per token budget, for Dust (violet) and EGGROLL-Transformer (gray). Points are the cells of Table 1, the dotted black line is backprop tuned at the same budget, and the dashed curve is a power law fit through Dust’s five cells. The 100k panel has its own y range; the other three share one. At 20M the dash-dotted violet line is the fit’s limit, 4.43, an extrapolation from a ladder that is still falling, and the bar at $\infty$ is its 95% bootstrap interval, with the seeds resampled at each population.
We sweep the token budget from 100k to 20M against populations from 64 to 16k, with backprop tuned on the same grid at every budget (Table 1, Figure 2). At 100k and 1M tokens Dust ends below backprop, from a few hundred draws at 100k and from a thousand at 1M. At 10M and 20M tokens, the gap with backprop shrinks with population. At 10M the ladder has flattened and its fitted limit lands just above backprop. At 20M the ladder is still falling at 16k draws. Its power law fit puts the limit at 4.431 (95% interval 3.89 to 4.58), below backprop’s 4.633, but with the ladder still falling the fit is loosely constrained, so we read it as evidence that the gap keeps closing with population rather than as a measured limit.
Weight-space ES is far less efficient. With 256 times the population, EGGROLL at 16k still does not reach Dust at 64 draws. It comes within 0.02 at 100k tokens and stays 0.4 to 0.6 above at 1M, 10M and 20M (Figure 2). EGGROLL’s ladder is still falling steeply at 16k, so it would keep improving with more population, but continuing the ladder puts what it needs to match Dust’s smallest population at several thousand to about $10^4$ times that population (Appendix D). Figure 1 also shows the gap during training; Dust at 1k and 16k draws stays close to backprop’s curve throughout, while every EGGROLL curve falls behind early and flattens well above Dust at 64 draws.
EGGROLL-Transformer
Dust 5.248 Backprop 5.361
Figure 3: Test loss against population at 1M tokens with Adam, for Dust (violet), EGGROLL (gray) and backprop (black, dotted). The violet curve is a power law fit in population over the populations run, and the dash-dotted violet line is its limit, 5.248, below backprop. The bar at $\infty$ is the limit’s 95% bootstrap interval, 5.17 to 5.30, with the seeds resampled at each population. Means over five seeds for Dust and backprop and three for EGGROLL; the Adam runs of Dust give the token embedding more draws, about $1.4\times$ the listed population in total.
While we focus primarily on SGD for the rest of the paper, we repeat the 1M token ladder with Adam on all three methods (Figure 3), re-tuning backprop’s learning rates, Dust’s own hyperparameters and EGGROLL’s step size, momentum and fitness shaping at every population. Interestingly, EGGROLL gains almost nothing from Adam and its tuned Adam ladder lands within 0.01 of its SGD ladder in Table 1 at every population. However, Adam improves both Dust and backprop and leaves the ladder’s shape similar to before. Dust closes on backprop with a large population and the interval on its limit sits below backprop (Figure 3). So Dust’s estimate already works with modern optimizers, even though modern optimizers were optimized for backprop gradients. We suspect that coevolution of optimizers with Dust could lead to further gains and leave this to future work.
256 5.486
5.362
5.358
5.419
1k 5.265
5.161
5.158
5.214
4k 5.189
5.095
5.053
5.124
16k 5.171
5.065
5.036
5.086
Backprop 5.180
5.066
5.015
5.048
Figure 4: Four model sizes at a fixed 10M tokens, 2M to 243M parameters (L2/d128, L4/d256, L8/d512, L16/d1024), test loss against population. The table has the same cells and backprop per size, with the best size at each population in bold.
The conventional view is that zeroth-order methods cannot train large networks (Lillicrap et al., 2020Timothy P. Lillicrap, Adam Santoro, Luke Marris, Colin J. Akerman, and Geoffrey Hinton. Backpropagation and the brain. Nature Reviews Neuroscience, 21 (6): 335–346, 2020. doi: 10.1038/s41583-020-0277-3.). A forward pass returns a single scalar, so the variance of the gradient estimate grows with the number of perturbed dimensions, and with it the population needed for a useful update (Nesterov and Spokoiny, 2017Yurii Nesterov and Vladimir Spokoiny. Random gradient-free minimization of convex functions. Foundations of Computational Mathematics, 17 (2): 527–566, 2017. doi: 10.1007/s10208-015-9296-2.; Werfel et al., 2003Justin Werfel, Xiaohui Xie, and H. Sebastian Seung. Learning curves for stochastic gradient descent in linear feedforward networks. In Advances in Neural Information Processing Systems, volume 16, 2003.). We test this directly by training four sizes, 2M, 7M, 38M and 243M parameters (a $120\times$ range in parameter count), at a fixed 10M tokens across the population sweep (Figure 4; see Appendix F for tuning details). We make two striking observations that challenge conventional wisdom.
The right way to think about model size is therefore as the size and the geometry of the search space. A larger model has a larger space to search over, which is what lets it put a large population to use, and potentially a better-conditioned loss landscape geometry, which could be why the search is more effective even at small populations.
0.00 0.25 0.50 0.75 1.00
Dust, 10M tokens 64 1k 16k 128k
Dust, 100M tokens 64 1k 16k 128k
Dust, 1B tokens 64 1k 16k 128k
Population
MLP avg
Q
K
V
head
Figure 5: Cosine between the estimated and the backprop gradient on the same batch, per layer type (mean over layers; MLP averages the input and output projections), against population, for Dust on backprop-trained checkpoints at 10M, 100M and 1B tokens and for EGGROLL at 10M tokens at the same compute budget (left, note the axis). For Dust the lines are the fitted law $\cos(K) = c_{\max}/\sqrt{1 + c/K}$ through measurements at populations of 64 to 128k; for EGGROLL they connect the measurements. The dotted vertical line is the 16k population the ladders train at. Initialization and EGGROLL at 100M tokens are in Appendix B.
We measure the cosine between Dust’s estimate and the backprop gradient on the same batch, per layer type and per layer, on backprop-trained checkpoints spanning two orders of magnitude in tokens, i.e. from 10M to 1B tokens, with the population growing from 64 to 128k forward passes per step (Figure 5). The cosine rises with population for every layer type at every stage of training, and a two-parameter law
fits each layer type with RMSE below 0.06, with $c_{\max}$ the ceiling and $c$ the population at which the layer reaches $c_{\max}/\sqrt{2}$. Useful gradients emerge from the large population alone, with nothing about the chain rule built in, just light tuning of the hyperparameters to maximize the cosine similarity as described in Section 2.3. For the 100M token checkpoint, the fits to layer type mean cosine similarities have RMSEs of $0.0032$–$0.0367$ (see Appendix G for details). Dust’s gradients approach backprop’s in cosine, but how close they get varies by layer type and by layer. EGGROLL’s estimate, measured the same way at the same number of forward passes, is growing with population but is still below 0.05 in cosine at 128k for every type but the head, which might explain why training on it went nowhere in Section 3.2.
Importantly, the cosines hold up across most layers as the token count grows, which is encouraging for scaling. Section 3 shows the population required to match and exceed backprop grows with tokens. However, the cosine stays flat across two orders of magnitude in tokens at large populations, which means that at a large enough population it’s possible that this requirement no longer grows with more tokens. Furthermore, the fact that Dust’s gradients approach backprop’s but don’t converge to them exactly is in fact a good property. The estimate points in a similar direction without being backprop’s gradient, which leads to a different optimization trajectory, and in our experiments that trajectory can even be better than backprop’s (Section 3.2).
Since Rumelhart et al. (1986)David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams. Learning representations by back-propagating errors. Nature, 323: 533–536, 1986., backprop has been the algorithm that trains neural nets, and the architectures, optimizers and hardware we have were all built around it. As compute becomes more abundant, we think much better alternatives are possible. We introduced Dust, an algorithm that drastically improves upon existing ES algorithms and approximates backprop closely at pretraining transformers, even exceeding it with large amounts of computation.
There are many interesting open questions. The first is whether, and how, Dust can find better directions than backprop’s first-order gradient by implicitly exploring the loss landscape, picking up higher-order curvature that pulls the search toward flat regions. We have hints that it can, since at large populations it sometimes ends up below backprop, but the mechanism is not clear. The second is that Dust opens up the search space over architectures, since it does not need the network to be end-to-end differentiable, and it may do better where backprop is known to struggle, like recurrent or looped computation trained by backpropagation through time. The third is compute efficiency, which was not the focus of this paper. We must have orders of magnitude more compute efficiency before Dust becomes a practical alternative to backprop at current levels of compute.
Early work on zeroth-order pretraining of language models (Allaire et al., 2025Nathan Allaire, Mahsa Ghazvini Nejad, Sébastien Le Digabel, and Vahid Partovi Nia. Zeroth order optimization for pretraining language models. In Proceedings of ICPRAM, pages 113–121, 2025. doi: 10.5220/0013261100003905. URL https://doi.org/10.5220/0013261100003905.) examined the difficulty of training transformers from scratch with weight perturbations. Their later method, KronZO (Allaire et al., 2026Nathan Allaire, Sébastien Le Digabel, Dominique Orban, and Vahid Partovi Nia. Zeroth-order Kronecker optimization for pretraining language models. SN Computer Science, 7, 2026. URL https://www.gerad.ca/en/papers/G-2025-44. Article 162.), uses compact perturbations with Kronecker structure and selective directional updates to improve pretraining while reducing memory use. EGGROLL (Sarkar et al., 2025Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. arXiv preprint arXiv:2511.16652, 2025.) makes large populations of weight perturbations efficient on GPUs through low-rank structure. Dust instead searches over activations and uses rewards for each token to extract more credit from each forward pass.
For fine tuning, MeZO (Malladi et al., 2023Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D. Lee, Danqi Chen, and Sanjeev Arora. Fine-tuning language models with just forward passes. In Advances in Neural Information Processing Systems, volume 36, 2023.) showed that language models can be adapted with forward passes alone and memory use close to inference. Evolution Strategies at Scale (Qiu et al., 2026Xin Qiu, Yulu Gan, Conor F. Hayes, Qiyao Liang, Yinggan Xu, Roberto Dailey, Elliot Meyerson, Babak Hodjat, and Risto Miikkulainen. Evolution strategies at scale: LLM fine-tuning beyond reinforcement learning. In Proceedings of the International Conference on Machine Learning, 2026. URL https://arxiv.org/abs/2509.24372. arXiv:2509.24372.) demonstrates fine tuning of all parameters with ES in language models with billions of parameters. Neural Thickets (Gan and Isola, 2026Yulu Gan and Phillip Isola. Neural thickets: Diverse task experts are dense around pretrained weights. arXiv preprint arXiv:2603.12228, 2026. URL https://arxiv.org/abs/2603.12228.) finds useful task experts by randomly perturbing pretrained weights, selecting the best candidates and ensembling their predictions. These results show how much search can achieve around a pretrained model. Our experiments address learning the representations themselves through pretraining from scratch.
Our activation perturbations build on node perturbation (Werfel et al., 2003Justin Werfel, Xiaohui Xie, and H. Sebastian Seung. Learning curves for stochastic gradient descent in linear feedforward networks. In Advances in Neural Information Processing Systems, volume 16, 2003.). GEMINI (Le Cun et al., 1988Yann Le Cun, Conrad C. Galland, and Geoffrey E. Hinton. GEMINI: Gradient estimation through matrix inversion after noise injection. In Advances in Neural Information Processing Systems, volume 1, pages 141–148, 1988. URL https://papers.neurips.cc/paper_files/paper/1988/file/a0a080f42e6f13b3a2df133f073095dd-Paper.pdf.) injected noise into the first hidden layer and recovered layerwise gradient estimates through iterative matrix inversion. Zoop (Hu et al., 2025Xixi Hu, Bo Liu, Qiang Liu, Xiaocong Du, Bhargav Bhushanam, Louis Feng, Chengyue Gong, and Kaizhao Liang. Zoop it! Efficient zero-order optimization with output perturbation. In ICML Workshop on Tiny Titans: The next wave of On-Device Learning for Foundation Models, 2025. URL https://openreview.net/forum?id=Tc8vFyRhPO.) uses output perturbations for language model fine tuning, converting estimated output gradients into parameter updates with local derivatives. Scaling Forward Gradient with Local Losses (Ren et al., 2023Mengye Ren, Simon Kornblith, Renjie Liao, and Geoffrey Hinton. Scaling forward gradient with local losses. In International Conference on Learning Representations, 2023.) combines activation perturbations and local losses with forward mode automatic differentiation to reduce estimator variance. Forward gradients with multiple tangents (Flügel et al., 2025Katharina Flügel, Daniel Coquelin, Marie Weiel, Charlotte Debus, Achim Streit, and Markus Götz. Beyond backpropagation: Optimization with multi-tangent forward gradients. In International Joint Conference on Neural Networks, 2025. URL https://arxiv.org/pdf/2410.17764v2.) also use forward mode differentiation, combining multiple directional derivatives through orthogonal projection to improve gradient estimates.
A separate line of work replaces the global backward pass with local learning dynamics: Sakana AI’s PC-ALM (Seely and Gould, 2026Jeffrey Seely and Julian Gould. Augmented Lagrangian predictive coding. arXiv preprint arXiv:2605.31022, 2026. URL https://arxiv.org/abs/2605.31022.) propagates credit through local predictive coding dynamics and Lagrange multipliers. It uses local derivatives and is evaluated on image classification tasks. Dust estimates credit from forward perturbations during transformer pretraining, combining rewards for each token with local targets for attention outputs.
Nathan Allaire, Mahsa Ghazvini Nejad, Sébastien Le Digabel, and Vahid Partovi Nia. Zeroth order optimization for pretraining language models. In Proceedings of ICPRAM, pages 113–121, 2025. doi: 10.5220/0013261100003905. URL https://doi.org/10.5220/0013261100003905.
Nathan Allaire, Sébastien Le Digabel, Dominique Orban, and Vahid Partovi Nia. Zeroth-order Kronecker optimization for pretraining language models. SN Computer Science, 7, 2026. URL https://www.gerad.ca/en/papers/G-2025-44. Article 162.
Katharina Flügel, Daniel Coquelin, Marie Weiel, Charlotte Debus, Achim Streit, and Markus Götz. Beyond backpropagation: Optimization with multi-tangent forward gradients. In International Joint Conference on Neural Networks, 2025. URL https://arxiv.org/pdf/2410.17764v2.
Yulu Gan and Phillip Isola. Neural thickets: Diverse task experts are dense around pretrained weights. arXiv preprint arXiv:2603.12228, 2026. URL https://arxiv.org/abs/2603.12228.
Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, and Jack Lindsey. Verbalizable representations form a global workspace in language models. arXiv preprint arXiv:2607.15495, 2026.
Xixi Hu, Bo Liu, Qiang Liu, Xiaocong Du, Bhargav Bhushanam, Louis Feng, Chengyue Gong, and Kaizhao Liang. Zoop it! Efficient zero-order optimization with output perturbation. In ICML Workshop on Tiny Titans: The next wave of On-Device Learning for Foundation Models, 2025. URL https://openreview.net/forum?id=Tc8vFyRhPO.
Yann Le Cun, Conrad C. Galland, and Geoffrey E. Hinton. GEMINI: Gradient estimation through matrix inversion after noise injection. In Advances in Neural Information Processing Systems, volume 1, pages 141–148, 1988. URL https://papers.neurips.cc/paper_files/paper/1988/file/a0a080f42e6f13b3a2df133f073095dd-Paper.pdf.
Timothy P. Lillicrap, Adam Santoro, Luke Marris, Colin J. Akerman, and Geoffrey Hinton. Backpropagation and the brain. Nature Reviews Neuroscience, 21 (6): 335–346, 2020. doi: 10.1038/s41583-020-0277-3.
Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, et al. On the biology of a large language model. Transformer Circuits Thread, 2025.
Shengchao Liu, Dimitris Papailiopoulos, and Dimitris Achlioptas. Bad global minima exist and SGD can reach them. In Advances in Neural Information Processing Systems, volume 33, 2020.
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D. Lee, Danqi Chen, and Sanjeev Arora. Fine-tuning language models with just forward passes. In Advances in Neural Information Processing Systems, volume 36, 2023.