Training Transformer Language Models Without Backpropagation
TL;DR
- The largest models trained from scratch without backpropagation at any level. Every parameter update follows the Equilibrium Propagation (EP) rule: no layer, block, or output head is trained with a backprop rule. The models are standard transformer LMs (OLMo2-style blocks, 32k vocabulary, FineWeb-Edu); the trained network runs ordinary forward inference.
- Matches backprop within a 4-5% perplexity band at 72M parameters, against a backprop twin trained on identical data, steps, and optimizer: with multi-seed controls on both sides.
- First transformer language model in the EP family, at 2.15× the size of the largest prior EP-family result (VGG10 ImageNet classifier, ~63M params): models trained up to 135M parameters on 2.7B tokens.
- Over 5× cheaper per training step than the closest EP-family work: one nudged phase with 3 relaxation sweeps per step, versus two phases with 10 iterations each (20 total) in the prior state of the art.
- Verified gradient fidelity: the EP update maintains cosine ≈0.99 to the true backprop gradient throughout training: measured, at scale. Larger non-backprop transformers in the literature train most parameters with local backprop inside blocks; ours use none.
- Hardware-ready by measurement, not assumption: 8-bit quantization shows no EP-specific penalty against an equally-quantized backprop twin; under injected analog faults (1% forward noise, 10% error-channel noise, device-tolerance mismatch) the EP estimator tracks the faulted network at cosine ≈0.97: learning co-adapts to the hardware.
Abstract
Analog and physical accelerators promise order-of-magnitude energy savings for training, but they cannot run backpropagation natively: exact gradients on a physical substrate require per-step digitization or an adjoint copy of the hardware. Equilibrium Propagation extracts gradients from the physics itself: two relaxations and local reads: and is the only member of its family with a gradient-equivalence guarantee that we verify directly at scale. What the field has lacked is scale and rigor: EP results stopped at mid-size vision models, without matched controls. We train standard transformer language models end-to-end with EP: 72M parameters within 4-5% perplexity of a matched backprop twin, models up to 135M, with backprop twins, multi-seed discipline, and hardware-relevant ablations (quantization, analog faults, nudge operating windows, energy accounting) at every stage. Scaling a physical learning rule also surfaces new science: we identified a width-scaling loss in the EP gradient invisible to per-step alignment metrics, built an instrument that measures it in 90 minutes per candidate recipe, mapped its dose-response law, and demonstrated an estimator-side treatment that recovers 97% of it without touching the model or the cost budget.
Headline results
| Model | Data | EP (val CE) | BP twin | Gap |
|---|---|---|---|---|
| 72M transformer LM | FineWeb-Edu, 1.4B tok | 3.33 | 3.29 | +4-5% ppl |
| 135M transformer LM | FineWeb-Edu, 2.7B tok | trained end-to-end, zero instability events; scaling analysis below | ||
Twin discipline: identical architecture, tokenizer, data order, optimizer, steps, and evaluation; multi-seed on both sides (BP n=3, band ±0.006; EP n=2).
| Training cost, EP family | This work | Closest EP work (VGG10, ImageNet) |
|---|---|---|
| Nudged phases per step | 1 | 2 |
| Relaxation iterations per step | 3 | 20 |
| Model class and scale | transformer LM, 72M and 135M | convolutional classifier, ~63M |
The scaling science
Scaling a physical learning rule surfaces phenomena backprop never meets. Between widths 512 and 768 we identified a width-scaling loss in the EP gradient: localized to the top half of the network, invisible to every per-step alignment metric, and traced to response components that finite nudge displacement under-reaches. We built a screening instrument that measures this leak in 90 minutes per candidate recipe, mapped its dose-response law (logarithmic across two decades of displacement amplification), and demonstrated a pure estimator-side treatment that closes 97% of it: no change to the model, its inference path, or the cost budget. The same instruments provide the go/no-go protocol for each next rung of the ladder.
Hardware line
- Measured energy projection for an integrated weight-stationary realization: 0.21-0.63 pJ/MAC (SPICE-measured analog core + datasheet periphery), against a 0.3-1 pJ/MAC digital INT8 system envelope.
- Single-column analog prototype: SPICE-modeled, discrete multiplying-DAC parts list: kept at the “hardware someone can actually build” level.
- Nudge-amplitude operating windows and their evolution over training are mapped: the dynamic-range spec an analog implementation must meet.
Roadmap
Staged scaling with matched BP controls and hardware-relevant ablations at every rung: a 150M-600M ladder (does the gap grow or shrink with scale: measured, not assumed), then 1B-3B; each stage gated on the previous stage’s loss, alignment, and throughput numbers. In parallel: the algorithm→regime map across the activity-difference family (contrastive / coupled-learning arms on the same harness), and a bounded single-column hardware feasibility study.
BibTeX
@misc{ept2026,
title={Training Transformer Language Models Without Backpropagation},
author={Hao, Yuren and Wan, Xiang and Gladstone, Alexi and Liu, Zeyi and Zhai, ChengXiang},
year={2026},
note={Project page}
}