Training Transformer Language Models Without Backpropagation
Standard LLMs, trained end-to-end by Equilibrium Propagation — with matched backprop controls at every scale, and every design choice made under analog-hardware constraints.
Author list forthcoming · University of Illinois Urbana-Champaign
Paper — in preparationCode — on requestTraining curves — on request
TL;DR
The largest models ever trained from scratch without backpropagation at any level.
Every parameter update follows the Equilibrium Propagation (EP) rule; no layer, block, or output head
is trained with a backprop rule. Our models are standard transformer LMs (OLMo2-style blocks,
32k-token vocabulary, FineWeb-Edu) — the trained network runs ordinary forward inference.
Matches backprop within a 4–5% perplexity band at 72M parameters, against a
backprop twin trained on identical data, steps, and optimizer — with multi-seed controls on both sides.
First transformer language model in the EP family, and 2.15× the size of the
largest prior EP-family result (VGG10 image classifier, ∼63M params, Kerjan–Høier–Scellier 2026):
models trained up to 135M parameters on 2.7B tokens.
2× wall-clock overhead versus backprop — where the closest prior EP work pays
6.7–12×. One nudged phase and 3 relaxation sweeps per step versus their two phases × 10
iterations: 6.7× fewer relaxation iterations per step.
Verified gradient fidelity, not alignment by luck: the EP update maintains cosine
≈0.99 to the true backprop gradient throughout training — measured, at scale. Prior non-backprop
transformer results at larger sizes train most parameters with local backprop inside blocks;
ours use none.
Hardware-ready by measurement, not assumption: 8-bit weight/compute quantization shows
no EP-specific penalty against an equally-quantized backprop twin; under injected analog faults
(1% forward noise, 10% error-channel noise, device-tolerance mismatch) the EP estimator tracks the
faulted network at cosine ≈0.97 — learning co-adapts to the hardware.
Why this matters
Analog and physical accelerators promise order-of-magnitude energy savings for training, but they cannot
run backpropagation natively: exact gradients on a physical substrate require either per-step digitization
or a second, adjoint copy of the hardware. Equilibrium Propagation extracts gradients from the physics
itself — two relaxations and local reads — and is the only member of its family with a
gradient-equivalence guarantee we can verify at scale. The missing evidence has always been scale and
rigor: EP results stopped at mid-size vision models, without matched controls. This project supplies both,
on the model class that matters commercially: language models.
The claim we defend: a standard transformer LM can be trained to backprop-class quality with a
physics-compatible learning rule, at 2× backprop wall-clock in simulation — and the constraints that
matter for analog hardware (quantization, noise, nudge operating windows, energy) are measured quantities
in our stack, not assumptions.
Headline results
Model
Data
EP (val CE)
BP twin
Gap
72M transformer LM
FineWeb-Edu, 1.4B tokens
3.33
3.29
+4–5% ppl
135M transformer LM
FineWeb-Edu, 2.7B tokens
trained end-to-end, zero
instability events; scaling analysis below
Twin discipline: identical architecture, tokenizer, data order, optimizer (Muon hybrid),
steps, and evaluation; multi-seed on both sides (BP n=3, band ±0.006; EP n=2). Comparison row uses the
sealed flagship pair.
Cost vs backprop (wall-clock)
This work
Closest EP work (ImageNet VGG10)
Training overhead
2.0×
6.7× (single-sided) / 12× (centered)
Relaxation iterations / step
3 (one phase)
20 (two phases × 10)
The scaling science
Scaling a physical learning rule surfaces phenomena backprop never meets. Between 72M (width 512) and
135M (width 768) we identified a width-scaling loss in the EP gradient — localized to the top half of
the network, invisible to every per-step alignment metric, and traced to response components that finite
nudge displacement under-reaches. We built a screening instrument that measures this leak in 90 minutes
per candidate recipe, mapped its dose–response law, and demonstrated a pure estimator-side treatment
that recovers 97% of it without touching the model, its inference path, or the cost budget. Full-schedule
validation of the treated recipe is running now; the same instruments give the go/no-go protocol for each
next rung of the ladder (300M → 1B).
Hardware line
Measured energy projection for an integrated weight-stationary realization: 0.21–0.63 pJ/MAC
(SPICE-measured analog core + datasheet periphery), against a 0.3–1 pJ/MAC digital INT8 system envelope.
Single-column analog prototype: SPICE-modeled, discrete multiplying-DAC parts list — deliberately
kept at the “hardware someone can actually build” level.
Nudge-amplitude operating windows and their evolution over training are mapped — the dynamic-range
spec an analog implementation must meet.
Roadmap
Staged scaling with matched BP controls and hardware-relevant ablations at every rung: 150M–600M
scaling ladder (does the gap grow or shrink with scale — measured, not assumed), then 1B–3B.
Each stage is gated on the previous stage’s loss, alignment, and throughput numbers. In parallel: the
algorithm→regime map across the activity-difference family (contrastive / coupled-learning arms on the
same harness), and the single-column hardware feasibility study.
BibTeX
@misc{ept2026,
title = {Training Transformer Language Models Without Backpropagation},
author = {(author list forthcoming)},
year = {2026},
note = {Project page}
}