summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
-rw-r--r--docs/campaign/CASCADE_ABLATION_PLAN.md21
1 files changed, 21 insertions, 0 deletions
diff --git a/docs/campaign/CASCADE_ABLATION_PLAN.md b/docs/campaign/CASCADE_ABLATION_PLAN.md
index c29d476..931bee7 100644
--- a/docs/campaign/CASCADE_ABLATION_PLAN.md
+++ b/docs/campaign/CASCADE_ABLATION_PLAN.md
@@ -578,6 +578,27 @@ power law (dose ~ width^k -> soft wall, budget k for 270M) vs width THRESHOLD ne
0.076 gap hides a small closable leak (lk512 arms answer directly). Chained behind fw135m_dg128
on all three GPUs; ~12h of ladder after the validation completes.
+### RESULT 74 (2026-07-29): PHASE M LAUNCH PLAN (user-approved) — the 2x-halving ladder +
+successive-blind-extrapolation protocol; the deliverable is the EXTRAPOLATION-HORIZON CURVE.
+Criterion (user): everything serves transfer to 1B & 8B; experience that cannot survive scale
+jumps has zero inductive value. Meta-diagnosis from the wreckage: every broken constant was
+DIMENSIONAL (beta, res-gate, amp certificate, dgain=128); everything that transferred was
+STRUCTURAL (K-saturation, bsign stability, free-phase=forward) or a MEASUREMENT PROTOCOL
+(rho probe, leak instrument). Transferable assets can only take three forms: dimensionless
+laws, structural invariants, measurement protocols. A standing DIMENSIONAL AUDIT table will
+classify every recipe element.
+LADDER (frozen 72M recipe, no per-size tuning, class-1 data): C128/11M, C192/18M, C320/36M
+(new, EP+BP twins, n=2 each, Chinchilla 20 tok/param) + existing C512/72M, C768/135M = five
+points spanning 12x params. Queued behind dg128 on all 3 GPUs (~9h).
+PROTOCOL (strengthened): fit any law on the SMALLEST points only; blind-predict upward rung
+by rung (C320 -> C512 -> C768 -> C1152); the x128 dose at C768 is a mandatory PREDICTION
+target. Log prediction error vs extrapolation factor = the horizon curve. Horizon >= 3x ->
+1B (1.8x beyond C1152) is within demonstrated range; horizon < 1.5x -> the honest verdict is
+"EP-transformer experience does not compress into transferable laws", which is decisive for
+the hardware thesis too (a chip is a frozen recipe).
+After trainings: M1 displacement spectroscopy + M2 per-nonlinearity linearization battery +
+M3 optimizer-step differentials at every width; PREREG freeze before any validation run.
+
### RESULT 73 (2026-07-29): GOVERNING DOCTRINE (user ruling) — NO TRANSFERABLE SCALING
KNOWLEDGE EXISTS YET; positive scaling claims SUSPENDED; results reclassified into three
strict classes; next phase = mechanism identification with a pre-registered transfer law and