diff options
| author | Yuren Hao <yurenh2@illinois.edu> | 2026-07-10 15:08:24 -0500 |
|---|---|---|
| committer | Yuren Hao <yurenh2@illinois.edu> | 2026-07-10 15:08:24 -0500 |
| commit | 35f9d6e8fea9cfbd4fc235ec00135564f93b7465 (patch) | |
| tree | 1cad18e33f2b89d09f8a64809aedd264c1d9fa42 | |
| parent | 5ad46a697bbc2fa23ef2f202b3898edac7ac9d1b (diff) | |
RESULT 8 final: Muon generic (+0.13 both columns; BP+Muon 1.7020 vs EP+Muon 1.7254); Muon default from Stage-2; 3v3 fill launched
| -rw-r--r-- | docs/campaign/CASCADE_ABLATION_PLAN.md | 10 |
1 files changed, 10 insertions, 0 deletions
diff --git a/docs/campaign/CASCADE_ABLATION_PLAN.md b/docs/campaign/CASCADE_ABLATION_PLAN.md index a182507..9cbefda 100644 --- a/docs/campaign/CASCADE_ABLATION_PLAN.md +++ b/docs/campaign/CASCADE_ABLATION_PLAN.md @@ -420,3 +420,13 @@ for digital-standardness, is also the EP stability fix — "EP as configuration WATCH ITEM: cos drifting slowly (0.9947@6k -> 0.9892@14k), beta already at floor. Watcher re-armed with cos<0.985 trigger; if it keeps sliding by ~30k, try beta_floor 5e-4 or accept (grad quality still fine at 0.989). Remaining epoch ETA ~6h. + +### RESULT 8-FINAL (2026-07-10 13:3x): Muon attribution = GENERIC (helps both columns ~0.13 CE). +BP+Muon s1 **1.7020** vs BP+AdamW 1.8335; EP+Muon (1.7316/1.7191, n=2 mean 1.7254) vs EP+AdamW 1.8643. +Muon's advantage TRANSFERS to EP gradients at full magnitude — not an EP-specific synergy, the known +small/mid-scale Muon-beats-AdamW result, now demonstrated on backprop-free training. **Muon = default +optimizer for BOTH columns from Stage-2 (FineWeb-Edu) onward.** Muon-column EP-BP gap +0.023 ~ AdamW +column's +0.031 (consistent slight BP-lean on OLMo2, noise-band edge, on the watch list). +HW-narrative guard: Muon's Newton-Schulz is matrix-matrix (analog-dead) but the optimizer lives +DIGITAL-side per standing doctrine — GPU-pretraining Muon does NOT conflict with the factored-Adam +analog training story. Filling to 3v3 (BP+Muon s2/s3, EP+Muon s3) for the seal. |
