From e2845963b369c54dbe69c9bf473a2cb211e4b9e0 Mon Sep 17 00:00:00 2001 From: YurenHao0426 Date: Wed, 29 Jul 2026 19:08:44 -0500 Subject: No dash punctuation; overhead comparison EP-family only Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn --- index.html | 54 ++++++++++++++++++++++++++++-------------------------- 1 file changed, 28 insertions(+), 26 deletions(-) diff --git a/index.html b/index.html index 82ac0b6..9c2210c 100644 --- a/index.html +++ b/index.html @@ -154,12 +154,12 @@ - Paper — in preparation + Paper: in preparation - Code — on request + Code: on request @@ -184,23 +184,24 @@
  • The largest models trained from scratch without backpropagation at any level. - Every parameter update follows the Equilibrium Propagation (EP) rule — no layer, block, or + Every parameter update follows the Equilibrium Propagation (EP) rule: no layer, block, or output head is trained with a backprop rule. The models are standard transformer LMs (OLMo2-style blocks, 32k vocabulary, FineWeb-Edu); the trained network runs ordinary forward inference.
  • -
  • Matches backprop within a 4–5% perplexity band at 72M parameters, against a backprop - twin trained on identical data, steps, and optimizer — with multi-seed controls on both sides.
  • +
  • Matches backprop within a 4-5% perplexity band at 72M parameters, against a backprop + twin trained on identical data, steps, and optimizer: with multi-seed controls on both sides.
  • First transformer language model in the EP family, at 2.15× the size of the largest prior EP-family result (VGG10 ImageNet classifier, ~63M params): models trained up to 135M parameters on 2.7B tokens.
  • -
  • 2× wall-clock overhead versus backprop — where the closest prior EP work pays - 6.7–12×: one nudged phase × 3 relaxation sweeps per step versus two phases × 10 iterations.
  • +
  • Over 5× cheaper per training step than the closest EP-family work: one nudged phase + with 3 relaxation sweeps per step, versus two phases with 10 iterations each (20 total) in the + prior state of the art.
  • Verified gradient fidelity: the EP update maintains cosine ≈0.99 to the true backprop - gradient throughout training — measured, at scale. Larger non-backprop transformers in the + gradient throughout training: measured, at scale. Larger non-backprop transformers in the literature train most parameters with local backprop inside blocks; ours use none.
  • Hardware-ready by measurement, not assumption: 8-bit quantization shows no EP-specific penalty against an equally-quantized backprop twin; under injected analog faults (1% forward noise, 10% error-channel noise, device-tolerance mismatch) the EP estimator tracks the faulted network at - cosine ≈0.97 — learning co-adapts to the hardware.
  • + cosine ≈0.97: learning co-adapts to the hardware.
@@ -220,15 +221,15 @@ Analog and physical accelerators promise order-of-magnitude energy savings for training, but they cannot run backpropagation natively: exact gradients on a physical substrate require per-step digitization or an adjoint copy of the hardware. Equilibrium Propagation extracts gradients from the - physics itself — two relaxations and local reads — and is the only member of its family with a + physics itself: two relaxations and local reads: and is the only member of its family with a gradient-equivalence guarantee that we verify directly at scale. What the field has lacked is scale and rigor: EP results stopped at mid-size vision models, without matched controls. We train standard - transformer language models end-to-end with EP — 72M parameters within 4–5% perplexity of a - matched backprop twin, models up to 135M — at 2× backprop wall-clock, with backprop twins, + transformer language models end-to-end with EP: 72M parameters within 4-5% perplexity of a + matched backprop twin, models up to 135M, with backprop twins, multi-seed discipline, and hardware-relevant ablations (quantization, analog faults, nudge operating windows, energy accounting) at every stage. Scaling a physical learning rule also surfaces new science: we identified a width-scaling loss in the EP gradient invisible to per-step alignment metrics, built an - instrument that measures it in 90 minutes per candidate recipe, mapped its dose–response law, and + instrument that measures it in 90 minutes per candidate recipe, mapped its dose-response law, and demonstrated an estimator-side treatment that recovers 97% of it without touching the model or the cost budget.

@@ -250,7 +251,7 @@ - + @@ -258,10 +259,11 @@

Twin discipline: identical architecture, tokenizer, data order, optimizer, steps, and evaluation; multi-seed on both sides (BP n=3, band ±0.006; EP n=2).

ModelDataEP (val CE)BP twinGap
72M transformer LMFineWeb-Edu, 1.4B tok3.333.29+4–5% ppl
72M transformer LMFineWeb-Edu, 1.4B tok3.333.29+4-5% ppl
135M transformer LMFineWeb-Edu, 2.7B toktrained end-to-end, zero instability events; scaling analysis below
- + - - + + +
Cost vs backpropThis workClosest EP work (VGG10, ImageNet)
Training cost, EP familyThis workClosest EP work (VGG10, ImageNet)
Wall-clock overhead2.0×6.7× (single-sided) / 12× (centered)
Relaxation iterations / step3 (one phase)20 (two phases × 10)
Nudged phases per step12
Relaxation iterations per step320
Model class and scaletransformer LM, 72M and 135Mconvolutional classifier, ~63M
@@ -277,11 +279,11 @@

The scaling science

Scaling a physical learning rule surfaces phenomena backprop never meets. Between widths 512 and 768 - we identified a width-scaling loss in the EP gradient — localized to the top half of the network, + we identified a width-scaling loss in the EP gradient: localized to the top half of the network, invisible to every per-step alignment metric, and traced to response components that finite nudge displacement under-reaches. We built a screening instrument that measures this leak in 90 minutes per - candidate recipe, mapped its dose–response law (logarithmic across two decades of displacement - amplification), and demonstrated a pure estimator-side treatment that closes 97% of it — no change to + candidate recipe, mapped its dose-response law (logarithmic across two decades of displacement + amplification), and demonstrated a pure estimator-side treatment that closes 97% of it: no change to the model, its inference path, or the cost budget. The same instruments provide the go/no-go protocol for each next rung of the ladder.

@@ -289,18 +291,18 @@
  • Measured energy projection for an integrated weight-stationary realization: - 0.21–0.63 pJ/MAC (SPICE-measured analog core + datasheet periphery), against a - 0.3–1 pJ/MAC digital INT8 system envelope.
  • -
  • Single-column analog prototype: SPICE-modeled, discrete multiplying-DAC parts list — kept at the + 0.21-0.63 pJ/MAC (SPICE-measured analog core + datasheet periphery), against a + 0.3-1 pJ/MAC digital INT8 system envelope.
  • +
  • Single-column analog prototype: SPICE-modeled, discrete multiplying-DAC parts list: kept at the “hardware someone can actually build” level.
  • -
  • Nudge-amplitude operating windows and their evolution over training are mapped — the +
  • Nudge-amplitude operating windows and their evolution over training are mapped: the dynamic-range spec an analog implementation must meet.

Roadmap

-

Staged scaling with matched BP controls and hardware-relevant ablations at every rung: a 150M–600M - ladder (does the gap grow or shrink with scale — measured, not assumed), then 1B–3B; each stage +

Staged scaling with matched BP controls and hardware-relevant ablations at every rung: a 150M-600M + ladder (does the gap grow or shrink with scale: measured, not assumed), then 1B-3B; each stage gated on the previous stage’s loss, alignment, and throughput numbers. In parallel: the algorithm→regime map across the activity-difference family (contrastive / coupled-learning arms on the same harness), and a bounded single-column hardware feasibility study.

-- cgit v1.2.3