From 2ac145ed97ea81d0606c7c5b9fa71fd3e7336296 Mon Sep 17 00:00:00 2001 From: YurenHao0426 Date: Wed, 29 Jul 2026 20:43:00 -0500 Subject: Abstract: same doctrine Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn --- index.html | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) (limited to 'index.html') diff --git a/index.html b/index.html index b15f65f..334895b 100644 --- a/index.html +++ b/index.html @@ -228,10 +228,10 @@ matched backprop twin, models up to 135M, with backprop twins, multi-seed discipline, and hardware-relevant ablations (quantization, analog faults, nudge operating windows, energy accounting) at every stage. Scaling a physical learning rule also surfaces new science: - we identified a width-scaling loss in the EP gradient invisible to per-step alignment metrics, built an - instrument that measures it in 90 minutes per candidate recipe, mapped its dose-response law, and - demonstrated an estimator-side treatment that recovers 97% of it without touching the model or the - cost budget. + the frozen recipe fails one width step up, with sharp measurable structure. Locating the mechanism, + fitting a transfer law, and validating it blind at the next width is the current phase; the + instruments for that campaign (a 90-minute estimator-quality screen and a stability probe) are built + and calibrated.

-- cgit v1.2.3