From 35f9d6e8fea9cfbd4fc235ec00135564f93b7465 Mon Sep 17 00:00:00 2001 From: Yuren Hao Date: Fri, 10 Jul 2026 15:08:24 -0500 Subject: RESULT 8 final: Muon generic (+0.13 both columns; BP+Muon 1.7020 vs EP+Muon 1.7254); Muon default from Stage-2; 3v3 fill launched --- docs/campaign/CASCADE_ABLATION_PLAN.md | 10 ++++++++++ 1 file changed, 10 insertions(+) (limited to 'docs/campaign/CASCADE_ABLATION_PLAN.md') diff --git a/docs/campaign/CASCADE_ABLATION_PLAN.md b/docs/campaign/CASCADE_ABLATION_PLAN.md index a182507..9cbefda 100644 --- a/docs/campaign/CASCADE_ABLATION_PLAN.md +++ b/docs/campaign/CASCADE_ABLATION_PLAN.md @@ -420,3 +420,13 @@ for digital-standardness, is also the EP stability fix — "EP as configuration WATCH ITEM: cos drifting slowly (0.9947@6k -> 0.9892@14k), beta already at floor. Watcher re-armed with cos<0.985 trigger; if it keeps sliding by ~30k, try beta_floor 5e-4 or accept (grad quality still fine at 0.989). Remaining epoch ETA ~6h. + +### RESULT 8-FINAL (2026-07-10 13:3x): Muon attribution = GENERIC (helps both columns ~0.13 CE). +BP+Muon s1 **1.7020** vs BP+AdamW 1.8335; EP+Muon (1.7316/1.7191, n=2 mean 1.7254) vs EP+AdamW 1.8643. +Muon's advantage TRANSFERS to EP gradients at full magnitude — not an EP-specific synergy, the known +small/mid-scale Muon-beats-AdamW result, now demonstrated on backprop-free training. **Muon = default +optimizer for BOTH columns from Stage-2 (FineWeb-Edu) onward.** Muon-column EP-BP gap +0.023 ~ AdamW +column's +0.031 (consistent slight BP-lean on OLMo2, noise-band edge, on the watch list). +HW-narrative guard: Muon's Newton-Schulz is matrix-matrix (analog-dead) but the optimizer lives +DIGITAL-side per standing doctrine — GPU-pretraining Muon does NOT conflict with the factored-Adam +analog training story. Filling to 3v3 (BP+Muon s2/s3, EP+Muon s3) for the seal. -- cgit v1.2.3