Twin discipline: identical architecture, tokenizer, data order,
optimizer, steps, and evaluation; multi-seed on both sides (BP n=3, band ±0.006; EP n=2).
Scaling a physical learning rule surfaces phenomena backprop never meets. Between widths 512 and 768
- we identified a width-scaling loss in the EP gradient — localized to the top half of the network,
+ we identified a width-scaling loss in the EP gradient: localized to the top half of the network,
invisible to every per-step alignment metric, and traced to response components that finite nudge
displacement under-reaches. We built a screening instrument that measures this leak in 90 minutes per
- candidate recipe, mapped its dose–response law (logarithmic across two decades of displacement
- amplification), and demonstrated a pure estimator-side treatment that closes 97% of it — no change to
+ candidate recipe, mapped its dose-response law (logarithmic across two decades of displacement
+ amplification), and demonstrated a pure estimator-side treatment that closes 97% of it: no change to
the model, its inference path, or the cost budget. The same instruments provide the go/no-go protocol
for each next rung of the ladder.
@@ -289,18 +291,18 @@
-
Staged scaling with matched BP controls and hardware-relevant ablations at every rung: a 150M–600M
- ladder (does the gap grow or shrink with scale — measured, not assumed), then 1B–3B; each stage
+
Staged scaling with matched BP controls and hardware-relevant ablations at every rung: a 150M-600M
+ ladder (does the gap grow or shrink with scale: measured, not assumed), then 1B-3B; each stage
gated on the previous stage’s loss, alignment, and throughput numbers. In parallel: the
algorithm→regime map across the activity-difference family (contrastive / coupled-learning arms on
the same harness), and a bounded single-column hardware feasibility study.
--
cgit v1.2.3