Dust: Training Transformers Without Backprop

Original: Dust: Pretraining Transformers Without Backpropagation

Why This Matters

If zeroth-order methods scale, the field's dependency on differentiable architectures weakens significantly.

Q Labs researchers published Dust, the first zeroth-order method competitive with backpropagation for pretraining transformer LMs. It perturbs activations per token, running virtual population members in one forward pass, and scales to 243M parameters.

Q Labs has released Dust, a research paper introducing what the team claims is the first zeroth-order optimization method that can match backpropagation when pretraining transformer language models. Instead of computing gradients via a backward pass, Dust uses node perturbation — independently perturbing activations at every token position so that each token acts as a virtual population member. One forward pass evaluates the entire population in parallel, eliminating the need for differentiability.

The efficiency gains over weight-space evolution strategies are substantial. The team estimates Dust is 10³ to 10⁴ times more efficient than EGGROLL, a state-of-the-art evolutionary strategy method, based on extrapolations from 1M tokens upward.

The counterintuitive finding: larger models benefit more, not less, from this approach. A 243M-parameter model outperformed a model roughly 120× smaller at most population sizes tested — directly challenging the widely-held assumption that zeroth-order methods fail to scale to large networks. Gradient estimates from Dust also align more closely with backprop as population size grows, and remain well-aligned up to 1B tokens. The authors frame this as a compute-scaling argument: in a compute-rich regime, brute-force search methods may eventually surpass backprop-based training, drawing a parallel to how AlphaGo Zero's self-play eventually overtook human-data bootstrapping.

Source

qlabs.sh — Read original →