diff options
Diffstat (limited to 'docs')
| -rw-r--r-- | docs/campaign/CASCADE_ABLATION_PLAN.md | 10 |
1 files changed, 10 insertions, 0 deletions
diff --git a/docs/campaign/CASCADE_ABLATION_PLAN.md b/docs/campaign/CASCADE_ABLATION_PLAN.md index a182507..9cbefda 100644 --- a/docs/campaign/CASCADE_ABLATION_PLAN.md +++ b/docs/campaign/CASCADE_ABLATION_PLAN.md @@ -420,3 +420,13 @@ for digital-standardness, is also the EP stability fix — "EP as configuration WATCH ITEM: cos drifting slowly (0.9947@6k -> 0.9892@14k), beta already at floor. Watcher re-armed with cos<0.985 trigger; if it keeps sliding by ~30k, try beta_floor 5e-4 or accept (grad quality still fine at 0.989). Remaining epoch ETA ~6h. + +### RESULT 8-FINAL (2026-07-10 13:3x): Muon attribution = GENERIC (helps both columns ~0.13 CE). +BP+Muon s1 **1.7020** vs BP+AdamW 1.8335; EP+Muon (1.7316/1.7191, n=2 mean 1.7254) vs EP+AdamW 1.8643. +Muon's advantage TRANSFERS to EP gradients at full magnitude — not an EP-specific synergy, the known +small/mid-scale Muon-beats-AdamW result, now demonstrated on backprop-free training. **Muon = default +optimizer for BOTH columns from Stage-2 (FineWeb-Edu) onward.** Muon-column EP-BP gap +0.023 ~ AdamW +column's +0.031 (consistent slight BP-lean on OLMo2, noise-band edge, on the watch list). +HW-narrative guard: Muon's Newton-Schulz is matrix-matrix (analog-dead) but the optimizer lives +DIGITAL-side per standing doctrine — GPU-pretraining Muon does NOT conflict with the factored-Adam +analog training story. Filling to 3v3 (BP+Muon s2/s3, EP+Muon s3) for the seal. |
