summaryrefslogtreecommitdiff
path: root/docs/campaign
diff options
context:
space:
mode:
Diffstat (limited to 'docs/campaign')
-rw-r--r--docs/campaign/CASCADE_ABLATION_PLAN.md10
1 files changed, 10 insertions, 0 deletions
diff --git a/docs/campaign/CASCADE_ABLATION_PLAN.md b/docs/campaign/CASCADE_ABLATION_PLAN.md
index a182507..9cbefda 100644
--- a/docs/campaign/CASCADE_ABLATION_PLAN.md
+++ b/docs/campaign/CASCADE_ABLATION_PLAN.md
@@ -420,3 +420,13 @@ for digital-standardness, is also the EP stability fix — "EP as configuration
WATCH ITEM: cos drifting slowly (0.9947@6k -> 0.9892@14k), beta already at floor. Watcher re-armed
with cos<0.985 trigger; if it keeps sliding by ~30k, try beta_floor 5e-4 or accept (grad quality still
fine at 0.989). Remaining epoch ETA ~6h.
+
+### RESULT 8-FINAL (2026-07-10 13:3x): Muon attribution = GENERIC (helps both columns ~0.13 CE).
+BP+Muon s1 **1.7020** vs BP+AdamW 1.8335; EP+Muon (1.7316/1.7191, n=2 mean 1.7254) vs EP+AdamW 1.8643.
+Muon's advantage TRANSFERS to EP gradients at full magnitude — not an EP-specific synergy, the known
+small/mid-scale Muon-beats-AdamW result, now demonstrated on backprop-free training. **Muon = default
+optimizer for BOTH columns from Stage-2 (FineWeb-Edu) onward.** Muon-column EP-BP gap +0.023 ~ AdamW
+column's +0.031 (consistent slight BP-lean on OLMo2, noise-band edge, on the watch list).
+HW-narrative guard: Muon's Newton-Schulz is matrix-matrix (analog-dead) but the optimizer lives
+DIGITAL-side per standing doctrine — GPU-pretraining Muon does NOT conflict with the factored-Adam
+analog training story. Filling to 3v3 (BP+Muon s2/s3, EP+Muon s3) for the seal.