From 129aee2fd1fb344c1980ec687d4bff3f9e295734 Mon Sep 17 00:00:00 2001 From: Yuren Hao Date: Fri, 10 Jul 2026 03:05:42 -0500 Subject: RESULT 3: K4 depth-parity SEALED EP-favorable (EP 1.8907 <= BP 1.9192 @ L12xC512); full-epoch generation run auto-launched (6.9h) --- docs/campaign/CASCADE_ABLATION_PLAN.md | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/docs/campaign/CASCADE_ABLATION_PLAN.md b/docs/campaign/CASCADE_ABLATION_PLAN.md index a4fab80..279b2fe 100644 --- a/docs/campaign/CASCADE_ABLATION_PLAN.md +++ b/docs/campaign/CASCADE_ABLATION_PLAN.md @@ -217,3 +217,17 @@ E-tier: after C1 lands (tolerance scripts port directly) Naming: `casc_*` runs, wandb project **ept-cascade**. Gates report mean over ≥4 batches. In-flight single-block arms (rescv2, govfloor, fastfull/fastpair, gov_s11-14) continue untouched — they carry the dynamics paper + the two-stage-recipe science; D1 takes over the DEMO role only. + +### RESULT 3 (2026-07-10 03:03): K4 DEPTH-PARITY SEALED (EP-favorable) + full-epoch launched. +- **beta-floor 3e-4 EP 3-seed: 1.9005 / 1.9125 / 1.8591 = MEAN 1.8907** vs BP 1.9169/1.9194/1.9214 + (1.9192). **EP <= BP at L12xC512 (real GPT-small shape)** -- all 3 EP seeds below the best BP seed, + cos pinned 1.0000 throughout, zero skips. The L12 depth-tax is FULLY removed by the beta-floor; K4 + closes EP-favorable. (Un-floored control was 2.00 + unstable/non-reproducible -- see RESULT 2.) +- Headline now: "standard L12 transformer, no backprop, equilibrium-EP with beta-floor = BP quality + (slightly better) at matched tuning, real GPT-small shape." +- **FULL-EPOCH run LAUNCHED (user directive, auto-launched on verdict):** epoch_ep_bf3e4 -- 58,800 + steps = 1 epoch over TinyStories-BPE (361M tokens), beta_floor 3e-4 + --cosine (new flag), warmup + 500, save_every 5000. Running 2.376 it/s solo on GPU1 -> ~6.9 h. This is the "neng kan" generation + demo (task #15). BP twin epoch DEFERRED (no free GPU; parity already sealed so it is nice-to-have). +- Next: generation samples at checkpoints; BP-twin epoch when a GPU frees; then scale-up corpus + decision (FineWeb-Edu vs OLMo2/Dolma) for the larger model. -- cgit v1.2.3