diff options
| author | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-29 20:43:00 -0500 |
|---|---|---|
| committer | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-29 20:43:00 -0500 |
| commit | 2ac145ed97ea81d0606c7c5b9fa71fd3e7336296 (patch) | |
| tree | 9f5caf107070581332361fa6e96e065048ccb655 | |
| parent | 47ab7f9777852b5fd70e0dad90557057a3f42c13 (diff) | |
Abstract: same doctrine
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
| -rw-r--r-- | index.html | 8 |
1 files changed, 4 insertions, 4 deletions
@@ -228,10 +228,10 @@ matched backprop twin, models up to 135M, with backprop twins, multi-seed discipline, and hardware-relevant ablations (quantization, analog faults, nudge operating windows, energy accounting) at every stage. Scaling a physical learning rule also surfaces new science: - we identified a width-scaling loss in the EP gradient invisible to per-step alignment metrics, built an - instrument that measures it in 90 minutes per candidate recipe, mapped its dose-response law, and - demonstrated an estimator-side treatment that recovers 97% of it without touching the model or the - cost budget. + the frozen recipe fails one width step up, with sharp measurable structure. Locating the mechanism, + fitting a transfer law, and validating it blind at the next width is the current phase; the + instruments for that campaign (a 90-minute estimator-quality screen and a stability probe) are built + and calibrated. </p> </div> </div> |
