# Cascade-EP ablation program — standard multi-layer LLM, EP only in training **Date opened:** 2026-07-09 · **Trigger:** user directive — product form = standard L-layer transformer (plain-forward inference); the looped/weight-tied block is demoted to physics testbed. **Bridge:** layered energy E = Σ_l ½‖z_l − f_l(z_{l−1})‖² over DISTINCT standard blocks. Free equilibrium == the standard forward pass (E=0) ⟹ inference is a normal LLM forward. Training = two-phase (±β·CE at the top), relax states to nudged equilibria, ∇θ = (1/2β)[∂E/∂θ|₊ − ∂E/∂θ|₋]. Lineage: predictive-coding≈BP theorem family (Whittington-Bogacz 17; Song+ 20 / Z-IL), EP two-phase readout. **First gate (2026-07-09):** `cascade_probe.py` L3 C128 random init → cos(cascEP, BP) **0.9968** (blocks 0.9975/0.9980/0.9992, |EP|/|BP| 0.80–0.91). ## The five claims we are buying evidence for - **K1 exactness-on-trajectory** — the two-phase gradient matches BP not just at init but along a real training trajectory (weights with grown Jacobians stiffen the relaxation). - **K2 cost** — the nudged relaxation can be engineered to a small multiple of a BP step (scheme × K frontier), and the *physical* (Jacobi/parallel) scheme is not hopeless (analog story). - **K3 training parity** — full training closes to BP final CE at equal arch/steps (the money claim). - **K4 depth scaling** — no depth penalty vs BP at matched params (signal attenuation under control). - **K5 analog price** — per-block Jᵀ feedback, dynamic noise, quantization: the tolerance ledger ports from the looped-block program; PAR wall applies per block. Honest cost framing: on GPU cascade-EP is strictly MORE expensive per step than BP (K relax sweeps, each ≈ one fwd+state-vjp). The value is: standard-form deployment + local rules + analog trainability. The looped-EP precedent multiplier was ~230× BP; the K-frontier decides whether cascade beats that. --- ## STATUS 2026-07-11: K1+K2+K3 SEALED; D-tier in flight - K1 exactness: cos 0.9998-1.0000 on-trajectory + BP-free formally audited (test_bp_free.py in repo). - K2 cost: exact mode ~3.6x BP (v7); Sol audit says remaining eager headroom 5-10% (v8 queued). - K3 quality: **matched-tuning PARITY n=3** (EP-exact 2.0500±0.015 vs BP 2.0530±0.004 @ C256 L6, lr 1e-3 both). Arc: fake-win (lr artifact) -> fake-tax (v7 dedups) -> parity. Fast mode = documented -4%CE/+20%speed dial. A0.4: TF32 free, bf16 production-only (cos 0.9427). - D1a (L12xC512 45M): BP s1/s2 SEALED 1.9169/1.9194 (H8, lr1e-3, tok_init0.02, 4000 steps, adamw). - E-tier: next in queue (softmax pathology / error-channel SNR / write pricing) -> Demo-0 spec sheet. ## D1a AUTOPSY + K-LADDER DIAGNOSTIC (2026-07-09 night) **>>> CORRECTION (2026-07-10 02:xx): the "parent-death" below was a MISDIAGNOSIS. <<<** The original D1a arms did NOT die -- they completed normally. When I checked at ~23:44 they were ALIVE at step 3200 on GPUs 0/3 (both at 100%); my /proc scan was mangled by a zsh eval wrapper so I misread "no casc alive", and GPU1 being free (11 MiB) fooled me (the runs were on 0/3, not 1). d1_ep_s1.log is continuous 0->4000 at steady 0.679 it/s (3200->4000 = 19.6 min, matches its 00:06 mtime). **Original D1a finals: d1_ep_s1 1.9745 / s2 2.0013 / s3 2.1188 (fixed K3; s2 blew at step 4000 skips=9, s3 blew hard skips=23 governor ramped K->7); d1_ep_muon 2.7515 (Muon-on-EP, cos collapsed 0.82).** The d1b experiments I launched (thinking the originals died) ran on the GENUINELY-FREE GPU1, so no competition -- and they independently isolated the real mechanism + fix (below), which is the bigger prize. Net: no harm, wrong death-story, and we now have BOTH the original un-floored 3-seed AND the beta-floor fix. KEY read of the original 3-seed: un-floored K3 is HIGH-VARIANCE near the SNR cliff -- s1 got lucky and stayed stable (1.9745, closest to BP), s2/s3 blew up late. Same-seed non-determinism (fb+autograd reductions) means the un-floored estimator is not even reproducible near the cliff. That is the strongest argument FOR the beta-floor (which pins cos=1.0000, stable, reproducible). **What I ORIGINALLY (wrongly) concluded:** the 4 D1a arms all died at wall-clock 23:36, mid-run, at a step boundary with NO traceback and NO DONE marker -> classic PARENT-DEATH (launched inline, not nohup'd; the launching shell/session terminated and took them down). No OOM in journalctl/dmesg. NOT a training failure. **Lesson (re)applied: every relaunch is nohup + 0.9942 (@2800) -> 0.9897 (@3200) as beta_t adapted DOWN 3e-3 -> 1.9e-5. - EP s3: cos fell to 0.9834 AND the quality gate started SKIPPING steps (skips=4). - vs BP s1/s2 which finished clean at 1.917. So at step ~3200 EP is ~0.10-0.13 CE above BP and the curve is stalling while cos degrades -- the DEPTH-ATTENUATION / estimator-SNR prediction (B6/K4). **Mechanism hypothesis:** K=3 fb message-passing rounds were tuned at L6xC256; the deeper L12 nudged equilibrium under-converges, and as beta_t shrinks (nudge -> tiny) the two-phase difference becomes a small signal against fixed relaxation error -> cos erodes -> gradient quality drops late in training. **Diagnostic launched (local GPU1, nohup, seed 1, full 4000 steps, H8 lr1e-3 tok_init0.02 beta3e-3):** - `d1b_ep_K3_s1` (K=3 control, honest 4000-step reproduction) - `d1b_ep_K8_s1` (K=8 = kmax, strongest relaxation -- does more convergence hold cos~1 and close CE?) - `d1_bp_s3` relaunch (completes the 3-seed BP reference). **Decision rule:** if K8 holds cos>=0.999 through step 4000 and reaches ~BP CE -> gap was under-convergence, fix = scale K with depth, then relaunch full 3-seed at min-sufficient K for the K4 verdict. If K8 does NOT close it -> genuine estimator depth-tax; next arm = beta-floor (needs a code flag) and/or lambda_l per-layer energy weighting (B4). Follow-on (not yet launched): Muon-on-EP arm. ### RESULT 1 (2026-07-10 00:40): K REFUTED as the lever; BP 3-seed sealed. - BP 3-seed reference SEALED: 1.9169 / 1.9194 / 1.9214 = **1.9192 +/- 0.0019** (L12 C512 H8). - **cos is K-INVARIANT.** K3 and K8 track to 4 decimals through step 1200 (both 1.0->0.9997->0.9991) AND give identical val CE at every matched step (900: 2.545 vs 2.548; 1100: 2.399 vs 2.404). More relaxation rounds do NOTHING -> the cos erosion is NOT fb under-convergence. K8 killed (redundant). - **Real mechanism = finite-beta SNR collapse.** beta_t = beta0*bscale*(SIG0/sig)^2 collapses ~120x (3e-3 -> 2.5e-5) as sig_tok grows 1.6->17.8. The estimator computes E/(NBT*beta_t) from residuals (z-o) that are O(beta_t*sig) ~ 4e-4 obtained by subtracting two O(17) states -> catastrophic cancellation as beta shrinks AND sig grows. Both worsen with depth. cos erodes 1.0 -> 0.997 (@2000) -> 0.98 (@2800 in the dead run). This is a beta-SCHEDULE problem, not a relaxation-depth problem. - **Fix under test:** added `--beta_floor` / `--beta_fixed` flags. Launched paired arms seed 1 (control = K3 floor=0, still running): `d1b_ep_bf1e4_s1` (floor 1e-4), `d1b_ep_bf3e4_s1` (floor 3e-4). Decision rule: if floored cos stays high through step 2000-2800 and CE drops toward BP 1.919 -> beta-floor is the depth fix; pick min-sufficient floor, run 3-seed K4 verdict. Watch drift guard at the higher floor (larger nudge). If floors DON'T help -> escalate to double-sided estimator (cancels O(beta) Taylor bias, allows large beta, 2x cost) or lambda_l energy weighting. ### RESULT 2 (2026-07-10 01:26): beta-floor CONFIRMED as the depth fix. Paired seed-1 sweep, cos in the erosion zone (where control collapses): | arm | cos @2000..4000 | best CE | skips | |---|---|---|---| | K3 control (floor 0) | 0.977 -> 0.944 -> **0.896@4000** | 2.0009 | **17** | | bf1e4 (floor 1e-4) | 0.9996 (nearly flat) | 2.174@2000 (desc) | 0 | | bf3e4 (floor 3e-4) | **1.0000 flat** | 2.161@2000 (desc) | 0 | - Flooring beta_t ELIMINATES the erosion: bf3e4 holds cos=1.0000 exactly where the un-floored control collapses to 0.896 w/ 17 skips. Higher floor monotonically better CE at matched steps (3e-4 < 1e-4 < control). 3e-4 already achieves perfect cos + zero drift/skips -> the operating point (higher can only add Taylor bias). The un-floored control still banked best 2.0009 (from ~step 3200 before the late collapse), so beta-floor's CE win over 2.0009 is the depth-tax recovery. - **K4 verdict LAUNCHED:** d1b_ep_bf3e4_s1/s2/s3 (floor 3e-4) 4000 steps vs BP 1.9169/1.9194/1.9214 (1.9192). If EP 3-seed ~ 1.919 -> **K4 depth-parity SEALED at L12xC512 (real GPT-small shape)** -> green-light D1b long-run demo (the "neng kan" gate) + hardware outreach. Poller baqcm84j4 armed. - FIX SHIPPED to trainer: `--beta_floor` is the depth knob. Recommend it becomes default-on (e.g. 3e-4) for L>=12; harmless at L6 (schedule never drops that low there). NOTE for the paper: this is a clean "EP as configuration microscope" second instance -- depth exposes a finite-beta SNR floor that BP (exact grad, scale-robust) never sees; the floor is the physical-relaxation analog of gradient precision. Muon-on-EP arm still pending after the verdict. ## STATUS 2026-07-09 (same day): Tier 0 CLOSED GREEN via the zil scheme; C1 running - **Naive relaxation FAILS at depth** (the B1-lite sweep): jacobi K=40·L → cos 0.82 (L6) / 0.67 (L12) / 0.53 (L24), shrink dying 0.41→0.28; gsf/gsr with small-η+momentum no better; β-insensitive (0.01/0.03/0.1 identical) ⟹ binding error = RELAXATION INCOMPLETENESS, not Taylor bias. warp2.0 catastrophic (cos 0.11) under naive descent. - **Two implementation traps found:** (1) NBT-normalized energy made γ=1 actually γ=1/128; (2) plain γ=1 reverse sweep WITHOUT interleaved reads contaminates e_l with J_l·δ_{l−1} (same β-order as the signal) — final-state readout is directionally ruined (cos 0.30@L6). - **The fix = zil scheme (interleaved reverse sweep):** update z_l (γ=1, SUM units) then read θ_l IMMEDIATELY (e_l = −β·δ_l exact at the feedforward point; δ-recursion has NO linearization error). Single phase, β cancels exactly. **Results: cos = 1.0000 at L=6/12/24; io gate 0.9999; warp2.0 → 1.0000; real-trajectory ckpts (casc_bp6 s0→s4000) → 0.9998–1.0000. A0.1/A0.2/A0.3 all green.** Honest framing: zil is numerically BP restructured as per-layer local two-factor energy reads (no global backward graph); the EQUILIBRIUM mode (jacobi/CG to convergence) remains the physically-meaningful EP column — priced expensive by the sweep, CG/preconditioning is the B2 job, and it is the analog-hardware rung (E-tier). - **C1 (zil) ran and is RETIRED with zil itself:** casc_ep6 best 3.3236 vs BP twin 2.9746 (gap 0.35 — single-sided zil top-read carries an O(β) shift on the readout term; moot now). **USER DIRECTIVE (2026-07-09 night): zil is NOT the route — it is BP in disguise; the project stays on TRUE EP = equilibrium-mode two-phase relaxation.** zil survives only as (a) a diagnostic upper bound, (b) optionally a numerical STATE-INIT trick for GPU simulation (`--init_sweep`: readout still taken at the relaxed equilibrium = clean EP semantics; hardware needs no init trick — physics settles). **Critical path = B2: make the equilibrium solver cheap** (Adam-on-states / init-sweep warm start / GS-multi-sweep / λ_l preconditioning), then rerun C1 in equilibrium mode. ## Tier 0 — gate hardening (probe-scale, hours, no training) → K1 | ID | question | design | decision rule | |---|---|---|---| | A0.1 | does cos survive depth? | cos vs L ∈ {3,6,12,24}, C128, Jacobi K auto-scaled; ≥4 batches | cos ≥ 0.98 at L12 or B1 must fix it | | A0.2 | does cos survive training? | BP-train C256 L6 4k steps saving every 500 (`casc_bp_train.py`); gate at every ckpt; ALSO record required-K to reach res-tol | cos ≥ 0.97 at all ckpts; K growth ≤ 3× init→4k | | A0.3 | full-θ gate | include emb/pos/readout(tied) grads in the gate | all groups ≥ 0.97 | | A0.4 | precision | fp32 vs TF32 vs bf16 on the two-phase difference | pick cheapest safe mode (looped-EP lesson: TF32 killed relaxation — re-test here) | ## Tier 1 — relaxation engineering (the cost frontier) → K2 | ID | axis | arms | metric | |---|---|---|---| | B1 | scheme × K | Jacobi (physical, parallel) vs Gauss-Seidel fwd vs GS reverse (algorithmic; Z-IL limit) × K ∈ {12,25,50,100,200,400} at L6 & L12 | K needed for cos ≥ 0.98; wall-clock multiple vs one BP step | | B2 | state optimizer | GD vs +momentum vs Adam-on-states; η sweep | same | | B3 | nudge β | {0.003,0.01,0.03,0.1,0.3} × one-sided vs two-sided | cos, shrinkage |EP|/|BP|, required K | | B4 | energy weighting | raw ℓ₂ vs per-layer precision λ_l=1/RMS² vs LN-in-energy | per-block shrinkage PROFILE (fix the 0.80→0.91 depth attenuation) + relax conditioning | | B5 | stopping | fixed-K vs relax-to-tol | natural K distribution | | B6 | **depth attenuation / estimator SNR profile** | measure per-block error amplitude ‖e_l‖ and per-block cos vs depth, as f(L, β, K) | the estimator-precision law: how fast does the deep-layer signal die, and which knob (β, K, λ_l weighting) restores it | B1 is the single most consequential experiment in the program: if GS-reverse needs K≈L (Z-IL limit) we have a ~BP-cost algorithmic mode for GPU pretraining, and the Jacobi column is the honest analog-hardware price. Report all three columns — they are different products. **Dynamics-vs-estimator tradeoff (user insight, 2026-07-09):** the cascade is dynamically SIMPLER — the free phase is EXACT (a plain forward; no res/T1/fixed-point error, no Hopf, no collapse), so **C-tier default arms run with NO regularizers at all** (jr/resreg don't exist here; stability regs return only if evidence demands). The difficulty MOVES to the estimator: the two-phase difference must resolve per-layer error signals that ATTENUATE with depth (visible at L=3 already: shrink 0.80 bottom vs 0.91 top), finite-β Taylor bias and finite-K relaxation bias hit the deepest blocks first, and the difference-of-O(1)-quantities structure makes precision (A0.4, fp32-vs-TF32) bind harder than in looped-EP. B6 is the dedicated measurement; λ_l weighting (B4), β/K scheduling (B3/B1) and per-block rebalance (C5) are the candidate antidotes. ## Tier 2 — small full-training ablations (C256 L6 T256 TinyStories, 8–16k steps) → K3 | ID | arm | vs | |---|---|---| | C1 | **money run**: cascade-EP (B-tier winner) ×2–3 seeds | BP twin, same arch/data/AdamW/steps — target gap ≤ 0.05 CE | | C2 | K budget: {K*, 2K*, 4K*} | CE-vs-cost curve (training may need less relax than the gate does — looped-EP precedent: t2sel 40 trains, 80 gates) | | C3 | one-sided β (half cost) | two-sided | | C4 | AdamW | SGDM (shrinkage sensitivity — does 0.8–0.9 amplitude matter under Adam's rescaling?) | | C5 | shrinkage compensation: none | per-block grad-norm rebalance to BP profile (one-time calibration) | | C6 | B4-winner energy weighting | raw | Placement: 1080 farm **after a Pascal canary** (cascade-EP is a new workload class; the Pascal pathology ban was derived on looped-EP+regs — do a 800-step canary + cross-env fingerprint first). C256 L6 fits 8 GB (~19M params, ~2-3 GB act). ## Tier 3 — depth/scale rungs (Delta A40 chains) → K4 | ID | design | |---|---| | D1 | **north-star demo re-target**: L12 C512 (≈45M, a real GPT-small shape) cascade-EP vs BP twin — replaces the single-block 33M rung as the flagship demo (task #15) | | D2 | depth ladder at fixed params: L6/C724 vs L12/C512 vs L24/C362 — depth penalty vs BP? | | D3 | T 256→512 sanity (relax cost tracks attention; expect no surprise) | ## Tier 4 — analog/hardware arms (port the tolerance machinery) → K5 | ID | design | |---|---| | E1 | Jacobi + per-sweep dynamic noise: does the fnoise ≥1e-3 cliff reappear in cascade relaxation? | | E2 | Jᵀ ablation: replace J_lᵀe with fixed random Bᵀ (feedback-alignment) / PAR projection — the per-block analog-feasibility tax; FA classically works on shallow stacks, test at L6 | | E3 | static tolerance: wq8/wq6 weights inside relax | ## Sequencing & fleet ``` now: A0.1 + A0.3 + B1-lite (shared local GPU, ~1h) + casc_bp_train ckpt producer (107 free 1080) gate ok → B1 full / B2 / B3 / B4 (local A6000s as arms free; each = minutes-hours) → Pascal canary → C-tier fan-out on 1080 farm (6 arms × 1-2 days) → D1 chains on Delta A40 (queue behind current five lines) E-tier: after C1 lands (tolerance scripts port directly) ``` Naming: `casc_*` runs, wandb project **ept-cascade**. Gates report mean over ≥4 batches. In-flight single-block arms (rescv2, govfloor, fastfull/fastpair, gov_s11-14) continue untouched — they carry the dynamics paper + the two-stage-recipe science; D1 takes over the DEMO role only. ### RESULT 3 (2026-07-10 03:03): K4 DEPTH-PARITY SEALED (EP-favorable) + full-epoch launched. - **beta-floor 3e-4 EP 3-seed: 1.9005 / 1.9125 / 1.8591 = MEAN 1.8907** vs BP 1.9169/1.9194/1.9214 (1.9192). **EP <= BP at L12xC512 (real GPT-small shape)** -- all 3 EP seeds below the best BP seed, cos pinned 1.0000 throughout, zero skips. The L12 depth-tax is FULLY removed by the beta-floor; K4 closes EP-favorable. (Un-floored control was 2.00 + unstable/non-reproducible -- see RESULT 2.) - Headline now: "standard L12 transformer, no backprop, equilibrium-EP with beta-floor = BP quality (slightly better) at matched tuning, real GPT-small shape." - **FULL-EPOCH run LAUNCHED (user directive, auto-launched on verdict):** epoch_ep_bf3e4 -- 58,800 steps = 1 epoch over TinyStories-BPE (361M tokens), beta_floor 3e-4 + --cosine (new flag), warmup 500, save_every 5000. Running 2.376 it/s solo on GPU1 -> ~6.9 h. This is the "neng kan" generation demo (task #15). BP twin epoch DEFERRED (no free GPU; parity already sealed so it is nice-to-have). - Next: generation samples at checkpoints; BP-twin epoch when a GPU frees; then scale-up corpus decision (FineWeb-Edu vs OLMo2/Dolma) for the larger model. ## ROADMAP PIVOT (2026-07-10 03:2x, user directive): QK-norm inserted; staged scale-up. **User: cancel the full epoch (done — killed epoch_ep_bf3e4); insert a QK-norm version after the current 3-seed; then stages TinyStories-full-epoch -> FineWeb-Edu -> OLMo2.** **Why QK-norm:** RMS-normalize q,k per head before the scores (OLMo2/Llama-style). It BOUNDS the attention logits, attacking the SAME root cause as the beta-floor (sig_tok growth -> logit blowup -> finite-beta SNR collapse) but structurally. Analog-friendly (my analysis): it's divisive normalization (mature analog/neuromorphic primitive), its Jacobian is symmetric (does NOT worsen the PAR/non-reciprocity wall), it's feedforward (no digital root-finder / no adjoint), and it REUSES the softmax current-normalization circuitry (reuse doctrine, no tapeout). Bonus analog wins: bounds the input range of the analog softmax exp device; reduces sig-growth so relaxation is more robust. Analog-preferred alternative to A/B in E-tier: tanh logit soft-cap (tanh is a native analog transfer function -- possibly cheaper than the norm's square-sum+divide). **Code:** nn.MultiheadAttention replaced by explicit CausalSelfAttn (SDPA-backed, fast) in BOTH trainers; `--qk_norm` flag (RMS-norm over head_dim w/ learnable per-dim gain). Smoke: EP+qk_norm cos=1.0000, 40.06M preserved, 2.49 it/s, SDPA works in the fb backward (fb is first-order, no double-backward needed). Also added `--cosine` (warmup->cosine to 0.1x lr) for the long runs. **QK-norm validation matrix (8 runs, L12 C512, 4000 steps, launched on GPU1):** - qk_bp_s1/s2/s3 = BP + qk_norm (new reference with the new block) - qk_ep_bf_s1/s2/s3 = EP + qk_norm + beta_floor 3e-4 (PARITY test vs qk_bp) - qk_ep_nf_s1/s2 = EP + qk_norm, NO beta_floor (ANALOG test: does qk_norm ALONE hold cos, letting us DROP the beta-floor? un-floored non-qk collapsed to cos 0.896 by step 4000 -- see RESULT 2). Decision: (1) qk_ep_bf ~ qk_bp => parity preserved with qk_norm. (2) if qk_ep_nf ALSO holds cos~1 and matches => qk_norm supersedes the beta-floor (fewer knobs, cleaner analog story). Watcher qk_watch.sh fires at the early analog read (nf step 2500) or all-done. **STAGED SCALE-UP (after qk_norm validates):** Stage 1: TinyStories FULL EPOCH (58,800 steps, 361M tok) with the validated qk_norm recipe + cosine -> the "neng kan" generation demo (task #15). Stage 2: FineWeb-Edu (real corpus, 32-50k tokenizer, ~150-300M params) -- best small-LM quality. Stage 3: OLMo2 / Dolma recipe -- fully-open reproducible baseline for the paper/collaborators. EP scaling knobs carried forward: beta_floor (or qk_norm if it supersedes), possibly double-sided nudge at larger scale (cancels O(beta) Taylor bias). $20k/run (Rain) ~ few-B tokens/run. ### RESULT 4 (2026-07-10 06:16): QK-norm validated — parity holds; beta-floor still needed; Stage 1 launched. - **Parity with QK-norm (EP-favorable again):** BP+qknorm 1.9253/1.8753/1.9192 = 1.9066; EP+qknorm+beta_floor 1.8588/1.9176/1.8841 = **1.8868 <= BP**. QK-norm preserves EP=BP parity at L12. - **ANALOG ANSWER: QK-norm does NOT replace the beta-floor** (they are complementary). EP+qknorm WITHOUT the floor still erodes cos (1.0 -> 0.946 by step 3200) and lands ~0.09 worse CE (2.02 vs 1.89). Milder than the old non-QK collapse (0.896) but not fixed. **Why: sig_tok still grows to 21.5 even with QK-norm** -- QK-norm normalizes q,k INSIDE attention (bounds the attention LOGITS) but does NOT bound the residual/embedding scale that drives beta_t = beta0*(sig0/sig)^2. So beta_t still collapses -> estimator SNR still needs the floor. QK-norm's payoff is (a) attention logit-bounding (analog softmax device range), (b) scale robustness (logit growth is worse in bigger/deeper models), (c) it is standard OLMo2/Llama -> good for the scale-up. Recipe = **qk_norm + beta_floor together**. - **STAGE 1 LAUNCHED (user directive):** stage1_ep_qkbf -- TinyStories full epoch (58,800 steps, 361M tok), qk_norm + beta_floor 3e-4 + cosine, warmup 500, 2.4 it/s solo -> ~6.8 h. The "neng kan" generation demo. Watcher fires at step 10000 (first generation-worthy ckpt) / done / death. Then Stage 2 (FineWeb-Edu) -> Stage 3 (OLMo2). ### RESULT 5 (2026-07-10 07:3x): Stage-1 epoch BLEW UP @step 12100 — root-cause diagnosis (sig story REFUTED). The qk_norm+beta_floor+cosine epoch was healthy to ~11400 (best val 1.6669) then blew up (val 1.67->7.6, gn pre-clip 0.5->53) and oscillated in a degraded regime. **My first guess (sig_tok growth -> SNR collapse -> add final_ln) was WRONG, refuted by its own telemetry:** - sig rose only +8% (29.8@10000 -> 32.2@12000) then PLATEAUED; it was already ~30 at step 10000 when everything was healthy. An 8% change cannot cause a catastrophic transition. - cos was FINE (0.9935) until step 11900; the cos drop is a CONSEQUENCE of the blowup, not the cause. - grad-clip is ALREADY present (clip 1.0); gn=53 is pre-clip telemetry. Not a magnitude-spike issue. **LEADING INDICATOR = skips (drift-guard rejections = nudged fb relaxation drift>0.5 = CONVERGENCE FAILURE).** skips accelerate from ~step 11000 (4->13 by 11400) BEFORE gn (11700), cos (12000), val (12100). **Diagnosis: a CONTRACTIVITY BIFURCATION in the nudged fb relaxation** -- as training sharpens the operator (block Jacobians grow), an increasing fraction of batches have a non-contractive nudged iteration -> skipped -> gradient bias -> a marginally-converged batch emits a bad step -> over the edge. **This is the cascade analog of the looped-EP Hopf wall** (non-conservative attention loses contractivity as CE drops -- documented in ep-c512-residual-defense-fix). 4000-step runs never saw it (operator not sharp enough yet; edge ~step 11400). Right fix = CONTRACTIVITY control (resreg/jacreg or geta<1 damping), NOT final_ln. **CONFIRMATORY A/B/C (resume from ckpt-10000, pre-bifurcation, beta floored 3e-4 via --sig0 1.6):** A=control (K3,lr1e-3) -> should reproduce skip-climb+blowup; B=K8 (does more fb rounds hold skips? = marginal-contractivity test); C=lr3e-4 (slower sharpening -> delayed edge? = driver test). Code added: --resume, --sig0, --final_ln, --qk_norm(CausalSelfAttn/SDPA). Watcher diag_watch.sh armed. ## AUDIT (2026-07-10, model switch): re-review of the day's conclusions. Corrections + added controls. **What SURVIVES audit:** RESULT 1 (K-invariance data is solid; K plumbed, paid wall-clock, identical cos/CE); RESULT 2 (beta-floor effect is decisive and mechanistic: floored arms pin cos, unfloored collapses); the 4k-horizon numbers themselves; the blowup telemetry read (skips lead gn lead cos lead val); the D1a "no-death" correction; Delta cancellation scope. **CORRECTIONS from audit:** 1. **Muon verdict RETRACTED as confounded.** d1_ep_muon (2.7515, cos 0.82) ran in the ORIGINAL D1a batch, i.e. WITHOUT beta_floor — its cos collapse mirrors the unfloored control (0.896). "Naive Muon-on-EP fails" is NOT established; needs a re-run with beta_floor before any conclusion. 2. **Parity claims toned down.** n=3 with best-of-noisy-val (6-batch val, min over ~500 evals -> selection bias ~0.02-0.03, applied to both arms) means "EP 1.8907 vs BP 1.9192" is PARITY with an EP-leaning point estimate, not "EP beats BP". (RESULT 3's all-3-EP-below-all-3-BP is p~=0.05 rank evidence — suggestive, not sealed.) Same for RESULT 4 (EP s2 1.9176 > BP best 1.8753). 3. **"Depth-tax FULLY removed" was premature** — true only at the 4k-step horizon; the epoch blowup at ~11.4k shows a second, longer-horizon wall. Claim scoped accordingly. 4. **"skips = relaxation non-convergence" is UNVERIFIED.** The skips counter conflates the drift-guard and the gn-EMA-guard; drift telemetry is stale-on-reject (GOV['drift'] not updated on drift-reject) while gn telemetry does update on gn-reject. Guard-split counters (skd/skg) now added to the log line for all future runs. The contractivity-bifurcation story remains the leading HYPOTHESIS, not a finding. 5. **A/B/C lacked the decisive control: a BP arm.** If BP-from-the-same-ckpt ALSO blows up, the blowup is a CONFIG instability (tied readout + NO final LayerNorm + sig~30 logits is genuinely nonstandard — every real GPT has final-LN; final_ln then likely IS the fix, via bounded logits/curvature, even though the sig->beta-SNR mechanism was refuted), and EP is exonerated. If BP sails through while A blows, the bifurcation is EP-specific -> jacreg/damped-fb. **diag_D_bp launched** (BP + --resume added to casc_bp_train, same ckpt-10000, qk_norm, lr 1e-3). 6. **Resume confounds now on record:** optimizer state is NOT in the ckpt (fresh Adam moments — sig jumped 29.8->35.6 within 300 steps of resume, visibly faster drift than the original run) and the data-order RNG restarts from the step-0 stream. So arm A can only reproduce the blowup STATISTICALLY, not at step 12100; if ALL arms blow immediately after resume, suspect the Adam-cold-start artifact rather than the original mechanism. 7. **Arm B (K8) is weakly informative by design:** for a genuinely divergent nudged iteration, MORE rounds = MORE drift, so both "K8 helps" and "K8 hurts" fit the story. The causal weight is on C (lr, sharpening-rate driver) and D (BP, EP-specificity). 8. Process fixes: watcher was not harness-tracked (user caught it — now all watchers via tracked bg tasks); zsh $VAR word-splitting cost two launch retries (all launches now via bash scripts). ### RESULT 6 (2026-07-10 09:35): WALL-2 DIAGNOSED — marginal under-convergence, EP-specific; kretry fix shipped; OLMo2 matrix launched. A/B/C/D verdict (resume from pre-bifurcation ckpt-10000, beta floored): | arm | skips @ window | note | |---|---|---| | A ctl (K3, lr1e-3) | **16, accelerating** (val wobble 2.00@12400) | leading indicator REPRODUCES | | B K8 | **2** | rejections nearly eliminated | | C lr3e-4 | **1**, best 1.5163 (best of all) | never touches the edge | | D BP (same ckpt/config/lr) | clean through 12750 | **EP-specific confirmed** | **Mechanism (two walls, two levers — revises "K refuted"):** - Wall-1 (~2-4k): cos erosion = finite-beta SNR -> beta-floor (K genuinely irrelevant there). - Wall-2 (~11k+): operator sharpens -> a growing fraction of batches sit at the CONTRACTIVITY EDGE of the nudged fb relaxation and under-converge at K3 -> drift-guard rejections climb -> gradient bias + occasional marginal escapes -> blowup. K8 CONVERGES those batches (16 -> 2 rejections) => marginal under-convergence, NOT hard divergence. lr modulates when the edge arrives (C: skips~1 and better CE). BP has no relaxation -> no wall-2 (D clean). Original 12100 didn't literally replay in A (fresh Adam + different data order — the recorded confounds) but the leading indicator did. **FIX SHIPPED: `--kretry N`** — on drift-reject, RETRY the batch once with N fb rounds (B proved K8 converges them) instead of dropping it. Converts biased skips into converged gradients; costs extra rounds ONLY on marginal batches (~0.1-1% of steps). Telemetry: skips=(d/g/r). **OLMo2 4k matrix LAUNCHED** (ol_bp_s1-3 + ol_ep_s1-3, wd 0.1, EP: beta_floor 3e-4 + kretry 8; twin step-0 losses bitwise-identical per seed). Watcher auto-computes parity and — if EP mean within 0.05 of BP — AUTO-LAUNCHES the Stage-1 OLMo2 TinyStories epoch (stage1_ol_ep, 58.8k steps, kretry armed). OLMo2's bounded-per-branch signals may also shift wall-2 later; kretry is the belt-and-suspenders. ### RESULT 6-ADDENDUM (2026-07-10 10:5x): B(K8) ALSO BLEW at matched step — K delays, does NOT prevent. diag_B_k8 @12600: train 3.37 / val 3.63 (best 1.6755 pre-blowup), gn 18.6, skips 2->14. So wall-2 is NOT merely marginal under-convergence: the nudged fb iteration becomes GENUINELY DIVERGENT for a growing batch fraction as the operator sharpens (true contractivity crossing — the cascade Hopf wall). More rounds converge the marginal shell only; once past the edge no K helps. **kretry DEMOTED from fix to mitigation** (still right for sporadic healthy-regime rejections). Note also: B blew with only 14 total rejections => most bad gradients passed UNDER the drift-0.5 threshold (loose guard + gn-EMA poisoning during degradation). Surviving facts: C (lr 3e-4) clean at 12600 (skips=1) -> sharpening RATE is the driver; D (BP) clean -> EP-specific. **Defense ranking now: (1) OLMo2 arch (different operator: bounded branches + QK-norm; diagnostics were all on the OLD arch) -> (2) lr channel (lower peak / faster decay through the mid-training danger window) -> (3) true contractivity control (damped-fb gamma<1 / cascade-jacreg) if OLMo2 still hits the wall.** Stage-1 OLMo2 epoch (parity-gated autolaunch) is the live test; watch skips=(d/g/r) through the 10-14k window. ### RESULT 7 (2026-07-10 12:0x): OLMo2 4k PARITY — gate PASSED; arch worth ~0.07-0.10 CE to BOTH; epoch auto-launched. - **BP+OLMo2: 1.8294/1.8333/1.8378 = 1.8335 (±0.004)** | **EP+OLMo2: 1.8294/1.8818/1.8817 = 1.8643** | gap +0.031 -> PASS (<=0.05) -> stage1_ol_ep AUTO-LAUNCHED (58.8k steps, qk+floor+kretry+cosine+wd). - OLMo2 improved BOTH columns vs old arch at 4k (BP 1.9066->1.8335; EP 1.8868->1.8643) — the arch upgrade pays for itself immediately. - HONEST READ: EP s1 == BP s1 to 4 decimals (1.8294, twin init); but EP s2/s3 trail their BP twins by ~0.045. Mean gap +0.031 is WITHIN the best-of-noisy-val metric band (~0.02-0.03, per audit), so: parity within noise, point estimate now slightly BP-leaning (was EP-leaning on old arch). Watch, not act: candidate causes = beta_t schedule now keyed to untied W_out sigma; norm-after changing fb conditioning (canary cos 0.9991 vs 1.0000). If the epoch shows a real gap, revisit. - Pascal canaries GREEN (EP cos 0.9991 flat, 0 skips, 0.64 it/s; BP 2.0 it/s) -> farm UNBANNED for cascade: 2x BP-twin epochs (stage1_ol_bp_s1/s2, ~8h) + Muon-with-floor retest (ol_ep_muon_s1) now running on timan107 GPUs 6/2/7. NOTE farm-vs-local init differs (torch 2.3.1 vs 2.10 CUDA RNG) — config-matched anchors, not init-twins. - Wall watch armed on the epoch: report at step 14000 (past the old 11.4k wall) with skips=(d/g/r). ### RESULT 8-PRELIM (2026-07-10 12:4x): Muon retraction CLOSED — with beta-floor, EP+Muon WINS big (n=1). ol_ep_muon_s1 (OLMo2 + beta_floor + kretry + Muon, Pascal GPU7): **1.7316**, cos 0.9992, ZERO skips. vs same-config-seed AdamW columns: BP 1.8294 / EP 1.8294 -> **-0.098 CE** (3-5x the metric noise band). The original "naive Muon-on-EP fails (2.7515)" was ENTIRELY the missing beta-floor (audit correction vindicated). Muon's known small/mid-scale advantage over AdamW TRANSFERS to EP gradients. Controls launched: ol_bp_muon_s1 (the fair Muon-column comparison) + ol_ep_muon_s2 (seed robustness). If BP+Muon lands ~1.73 too -> Muon helps both equally (parity preserved, recipe upgraded for BOTH columns). If BP+Muon ~1.83 -> EP-specific synergy (bigger story, needs replication before claiming). Interim: BP-twin epochs healthy at ~13.7k (best ~1.60 — already past the old-arch EP wall step); local EP epoch at 4.8k, best 1.8148, zero skips, 1.97 it/s. ### RESULT 9 (2026-07-10 13:1x): WALL-2 ELIMINATED BY ARCHITECTURE — epoch cleared 10-14k with ZERO guard events. stage1_ol_ep cleared the wall window (through step 14300): **skips=0 (d0/g0/r0) THE ENTIRE RUN** — not one drift rejection, not one gn rejection, kretry never fired. (Old arch: 32 skips by 13000, blowup at 12100; K8 variant blew by 12600.) gn calm (~0.5), best val 1.6702 and descending at 1.97 it/s. **OLMo2's bounded operator (norm-after-sublayer + QK-norm) stays contractive where the old block went divergent — defense #1 closed the case; mitigations (kretry) unused.** The architecture change, made for digital-standardness, is also the EP stability fix — "EP as configuration microscope" ends as "modern standard config is EP-compatible out of the box." WATCH ITEM: cos drifting slowly (0.9947@6k -> 0.9892@14k), beta already at floor. Watcher re-armed with cos<0.985 trigger; if it keeps sliding by ~30k, try beta_floor 5e-4 or accept (grad quality still fine at 0.989). Remaining epoch ETA ~6h. ### RESULT 8-FINAL (2026-07-10 13:3x): Muon attribution = GENERIC (helps both columns ~0.13 CE). BP+Muon s1 **1.7020** vs BP+AdamW 1.8335; EP+Muon (1.7316/1.7191, n=2 mean 1.7254) vs EP+AdamW 1.8643. Muon's advantage TRANSFERS to EP gradients at full magnitude — not an EP-specific synergy, the known small/mid-scale Muon-beats-AdamW result, now demonstrated on backprop-free training. **Muon = default optimizer for BOTH columns from Stage-2 (FineWeb-Edu) onward.** Muon-column EP-BP gap +0.023 ~ AdamW column's +0.031 (consistent slight BP-lean on OLMo2, noise-band edge, on the watch list). HW-narrative guard: Muon's Newton-Schulz is matrix-matrix (analog-dead) but the optimizer lives DIGITAL-side per standing doctrine — GPU-pretraining Muon does NOT conflict with the factored-Adam analog training story. Filling to 3v3 (BP+Muon s2/s3, EP+Muon s3) for the seal. ### WATCH (2026-07-10 14:0x): late-epoch cos erosion = intrinsic late-training SNR decline. Decision: let it run. cos 0.9894@17k -> 0.9840@22k (-0.0011/1k), ZERO skips, gn calm, val still improving (1.6353). sigma plateaued (~32) and beta at floor => ratio stable => NOT the sigma-growth wall-1. Mechanism: true gradient magnitude shrinks as CE approaches optimum while the estimator noise floor stays constant -> SNR falls with the signal. Extrapolates to cos~0.94 by 58.8k. DECISION: no mid-flight surgery (resume reintroduces Adam/data confounds; cosine-LR shrinks late steps anyway). The BP epoch twins ARE the measurement: EP final within ~0.03 of BP -> erosion harmless; 0.1 behind -> quantified problem with a ready dial (late beta_floor schedule, e.g. 5e-4 past 20k). Pre-validation probe queued: when the farm frees, run ckpt-25000 + floor 5e-4 x 2k steps. Watcher re-armed at cos<0.96. ### RESULT 8-SEALED + beta-floor dose-response (2026-07-10 16:1x). **Muon 3v3 SEALED: BP+Muon 1.7020/1.7137/1.7136 = 1.7098 | EP+Muon 1.7316/1.7191/1.7435 = 1.7314.** Muon default for both columns from Stage-2. The +0.02-0.03 BP-lean now CONSISTENT across two optimizer columns (6v6) -> upgraded from noise to "probably real small effect"; primary suspect = late-training SNR (see below), because it is beta-liftable: **beta_floor dose-response @ckpt-25000 (same weights/batch): 3e-4 -> cos 0.984 | 5e-4 -> 0.9889 | 1e-3 -> 0.9940.** Raising the floor lifts cos exactly as the SNR mechanism predicts, with drift=0.000 and zero skips at 1e-3 (larger nudge does NOT destabilize the OLMo2 relaxation). 2k-step traces harvesting (auto-kill at 27k). RECIPE UPDATE for Stage-2 (and the next epoch): late beta_floor schedule — floor 3e-4 early, ramp to ~1e-3 in the back half (or floor ∝ 1/grad-norm). This likely also closes the +0.02-0.03 gap. ### Dose-response SUSTAINED (25.6k-27k, 2k-step parallel traces): floor 3e-4 ~0.975 (accelerating down, -0.0025/1k) | 5e-4 ~0.983 | 1e-3 ~0.990 flat, zero instability. Late-SNR mechanism + fix both confirmed in-training. `--bf_late/--bf_late_at` flags shipped. NEXT-RUN RECIPE (post-epoch): OLMo2 + Muon + beta_floor 3e-4 + bf_late 1e-3 @ ~20k + kretry 8 — expected to hold cos>=0.99 end-to-end and likely close the +0.02-0.03 column gap. ### PLAN UPDATE (2026-07-10 17:2x): cos crossed 0.96 (0.9544@39.3k, accelerating) — flagship stays UNTOUCHED (the control measuring erosion damage vs BP twins); PARALLEL bf1e3 continuation launched from ckpt-40000 on the farm (floor 1e-3 for the remaining 18.8k steps). Endpoint comparison becomes a clean quad: EP-control(3e-4) / EP-floor-lift(1e-3 from 40k) / BP-s1 / BP-s2 — quantifies BOTH the erosion damage AND the fix's recovery in one shot. ### RESULT 10 (2026-07-10 19:0x): EPOCH ENDPOINTS + "NENG KAN" GATE PASSED + stage1b (improved recipe) launched. **Epoch endpoints (58,800 steps / 361M tokens, OLMo2, AdamW):** | arm | best val CE | |---|---| | BP s1 / s2 | **1.2750 / 1.2509** | | EP (floor 3e-4 fixed) | **1.4802** (zero guard events end-to-end) | | EP bf1e3-cont (floor->1e-3 @40k) | 1.4835@46k, running to 58.8k | **EP-BP gap at epoch scale = +0.22** (was +0.03 at 4k): the late-SNR cos erosion (1.0 -> ~0.92-0.95) is a REAL, horizon-growing CE cost with fixed floor 3e-4. Mechanism + dial both established (dose-response); the improved recipe is designed to close this. **GENERATION GATE ("neng kan") PASSED:** casc_gen.py (new; plain-forward standard-LLM inference) from EP s55000: coherent multi-paragraph TinyStories — named characters, balanced-quote dialogue, cause-effect, emotional arc (minor charm-defects vs BP's tighter coherence, consistent with +0.22). **A 42.75M standard 12-layer transformer trained end-to-end WITHOUT backprop tells coherent stories; inference is a plain forward pass.** Task #15 demo artifact exists. **stage1b launched (the improved-recipe head-to-head):** stage1b_ep_muon (local GPU1: Muon + floor 3e-4 + bf_late 1e-3@15k + kretry + cosine[now also on Muon via build_hybrid total_steps]) vs stage1b_bp_muon (farm GPU6: Muon + cosine). Expectation: EP ~1.25-1.35 (Muon -0.13 and erosion fix ~-0.1+), BP+Muon anchor moves too. ~8h both. ### QUEUE (user, 2026-07-10): double-sided nudge — implement AFTER stage1b endpoint. The 0.22 diagnosis: EP's one extra constraint = the gradient is a DIFFERENTIAL MEASUREMENT (SNR ∝ β|g|/(ε·σ)) vs BP's analytic adjoint. Escalation ladder: stage1b ramp (running) → double-sided ±β (kills O(β) Taylor bias, unlocks ~10× β for SNR, 2× nudge cost; A/B at 25k-ckpt 2k-step probe when implemented) → fp64 E-accumulation / readout averaging. Analog note: this constraint IS the hardware constraint (ε = device noise); β-scheduling learned here = chip ops manual; hardware bonus = nudge amplitude free under multiplicative noise (r-indifference). ### RESULT 11 (2026-07-11): stage1b SEALED (gap 0.22->0.050); beta ceiling not reached; K exonerated on OLMo2; bf16 naive-cast dead. - **stage1b endpoints: EP+Muon+floor-ramp 1.2808 | BP+Muon 1.2311 -> epoch gap 0.050** (fixed-floor was +0.22). EP now beats the old BP-AdamW epoch (1.2509/1.2750). Intervention-timing quad complete: fixed-floor 1.4802 / lift@40k 1.4479 / full ramp 1.2808 -- monotone earlier-is-better dose curve. - **Gap probes @s45000 (2k-step sustained traces):** control cos 0.9947 | b2e3 0.9968 | **b3e3 0.9974 (deficit halved, zero drift/skips)** | K5 0.9951 ~= control -> **K-invariance now proven on BOTH architectures; the residual deficit is beta-liftable, not relaxation-depth.** sigma(W_out)=80 by 45k: without the floor beta_t would be ~1e-6 -- the floor carries the entire late phase. CE-endpoint test launched: stage1b_f3e3cont (s45000 -> 58.8k at floor 3e-3, farm). If it closes to <=0.03, next-flagship recipe = ramp ...->3e-3@~35k; else double-sided (queued) takes the residual. - **bf16 gate: naive full-cast FAILS at any beta.** floor 3e-4 -> cos 0.33; 3e-3 -> 0.67; 1e-2 -> 0.65 (no longer SNR-limited: bf16 rounding distorts the nudged equilibrium itself; beta cannot compensate). Speed was 2.1x (5.2 it/s). VERDICT: cost baseline stays TF32 (validated); the x0.5 lever requires proper mixed precision (bf16 weights/matmuls + fp32 states/accumulation, autocast-style) -- queued as engineering upside, NOT in the Ben cost baseline. Wall-1 physics predicted all of this (SNR ∝ beta/eps; bf16 eps ~8000x fp32): the fp32/bf16/analog-noise beta-epsilon scaling story now has a second measured point. ## STANDING DIRECTIVE (user, 2026-07-11): LOOPED LINE ABANDONED. The looped/weight-tied single-block line is retired as a research direction. Default everywhere: cascade (tied, PCN-form energy over distinct standard blocks) is THE line. The looped record survives ONLY as historical evidence inside the dynamics paper (Hopf phenomenology, dips, governor, eig audits — valid data, past tense). Consequences: no new looped runs; looped-specific queue items closed (adaptive-eps integrator, Pascal five-arm reg triage, S1-S3 looped ladder); report v3 sections 7.1/8 to be reframed past-tense on the Overleaf pass ("a companion system we studied", not "our companion product"); AsymEP machinery = dynamics-paper subject matter, not the training recipe. ### RESULT 12 (2026-07-12): f3e3cont NEGATIVE — the residual 0.050 gap is NOT late-beta-SNR-limited. stage1b_f3e3cont (floor 3e-3 from s45000): **1.2883** vs stage1b 1.2808 (floor 1e-3) — no gain (cos 0.994->0.997 bought nothing in CE). beta lever exhausted at 1e-3. Residual-gap suspects, in order: (a) single-sided O(beta) Taylor bias sustained over 59k steps -> **next lever = RANDOM-SIGN beta** (flip sign per batch; single-phase cost; averages away the systematic first-order bias; validated competitive at full ImageNet by Kerjan-Hoier-Scellier) then centered (2x nudge) if needed; (b) fb K=3 finite-relaxation bias; (c) Muon x gradient-noise interaction; (d) ~0.02-0.03 of the 0.05 is metric-noise band. Recipe note: random-sign is a one-line trainer change (sign of beta_t per step). ### RESULT 13 (2026-07-12): bsign (random-sign beta) NEUTRAL at 4k — 42M gap-chasing has hit the noise floor. THREAD CLOSED. bsign 3-seed: 1.7031/1.7562/1.7552 (mean 1.7382) vs single-sided 1.7314 vs BP+Muon 1.7098. The bias reduction is cancelled by injected update-direction variance at this horizon (seed spread now dominates: s1 alone beat the BP mean). Ledger of the residual-0.05 epoch gap after three probes: NOT late-beta-SNR (f3e3cont), NOT K (K5 probe), NOT first-order sign bias at short horizon (bsign). Remaining mass: ~0.02-0.03 metric-noise band + small unattributed accumulation. **Decision: stop polishing 42M.** Carry `--bsign_rand` and a future centered mode as Stage-2 A/B flags; the gap question re-opens at 300M/real-corpus where it means something. Effort pivots to: (1) Stage-2 data pipeline (FineWeb-Edu + 32k tokenizer), (2) E-tier tolerance suite on the idle farm (hardware track / UIUC outreach feed). ### RESULT 14 (2026-07-12): E-TIER WAVE-1 — full analog-fault tolerance ledger at stage1b s55000. `etier_probe.py`, farm GPUs 2/3/7 (shards A/B/C), stage1b_ep_muon_s55000.pt (clean valCE 1.2678), B=8 eval batches; metrics = faulted valCE (Δ vs clean), cos(EP_faulted, BP_faulted) [self-consistency of the learning signal under fault], cos(EP_faulted, BP_clean) [direction vs the ideal update]. | fault (component) | mild | medium | severe | verdict | |---|---|---|---|---| | wq — weight quant (crossbar #3) | 8b: +0.004 / 0.956 | 6b: +0.051 / 0.812 | 4b: +2.23 / 0.05 | **8b FREE, 6b marginal, 4b dead → ≥7b effective is the binding spec** | | fnoise — fwd additive state noise (softmax/relax #5) | 1e-3: 0.000 / 0.975 | 3e-3: 0.000 / 0.974 | 1e-2: +0.001 / 0.970 | **FREE at 1% — looped-era 1e-3 cliff does NOT transfer to cascade** | | divmis — divisive-norm mismatch (#4/#7) | 1%: 0.000 / 0.971 | 3%: +0.003 / 0.948 | 10%: +0.042 / 0.785 | 3% (routine matching) FREE; 10% marginal | | rope — phase error rad (#2) | 0.01: 0.000 / 0.973 | 0.03: +0.001 / 0.967 | 0.1: +0.012 / 0.923 | 0.03 rad FREE; ~2° I/Q accuracy suffices | | gilbert — gate gain error (#6) | 1%: 0.000 / 0.974 | 3%: 0.000 / 0.971 | 10%: +0.006 / 0.945 | **FREE at 10%** — translinear practice is comfortably inside | | fbnoise — nudge/error-channel noise | 1e-2: 0.969 / 0.975 | **1e-1: 0.951 / 0.957** | 3e-1: 0.764 / 0.768 | **10% relative noise on the ERROR CHANNEL is FREE** (cos 0.95) — the r-indifference/large-nudge gift, now measured on cascade | Reading: (a) the only hard constraint is crossbar weight precision (≥7b effective — inside standard SRAM-CIM capability; 6b rescue = wave-2 quant-aware co-training); (b) everything dynamic — forward noise 1%, error-channel noise 10%, gate/divider/phase mismatch at routine device tolerances — is FREE at this scale. cos(EP,BP_faulted) stays ~0.97 under every non-fatal fault: the EP estimator tracks whatever network the faults define, i.e. learning co-adapts to the fault (the analog-training thesis in one number). CAVEAT: static probes at a trained checkpoint (eval CE + one-step gradient direction), not training-under-fault; wave-2 = co-training with faults injected from step 0 (expectation from the literature and from (c): tolerances IMPROVE). Feeds COMPONENT_HW_MAP.md (per-row status updated) + UIUC outreach dossier. ### QUEUED (2026-07-29, user design): WIDTH MINI-LADDER — scale DOWN to extrapolate UP. New FineWeb rungs below the existing pair: C256 (~27M, 0.53B tok) + C384 (~46M, 0.93B tok), L12/hd64, Chinchilla-matched, EP(bsign)+BP twins; plus the leak instrument AT C512 (existing fw72m_bsign_s190000: dgain dose {1,4,16} + allbp bar). Four-point deliverables: gap(width) and dose*(width). Key discriminator: C512->C768 dose jumped x1 -> x128 — smooth power law (dose ~ width^k -> soft wall, budget k for 270M) vs width THRESHOLD near C600 (leak absent below -> different physics, sharper story). Also quantifies whether the 72M 0.076 gap hides a small closable leak (lk512 arms answer directly). Chained behind fw135m_dg128 on all three GPUs; ~12h of ladder after the validation completes. ### RESULT 74 (2026-07-29): PHASE M LAUNCH PLAN (user-approved) — the 2x-halving ladder + successive-blind-extrapolation protocol; the deliverable is the EXTRAPOLATION-HORIZON CURVE. Criterion (user): everything serves transfer to 1B & 8B; experience that cannot survive scale jumps has zero inductive value. Meta-diagnosis from the wreckage: every broken constant was DIMENSIONAL (beta, res-gate, amp certificate, dgain=128); everything that transferred was STRUCTURAL (K-saturation, bsign stability, free-phase=forward) or a MEASUREMENT PROTOCOL (rho probe, leak instrument). Transferable assets can only take three forms: dimensionless laws, structural invariants, measurement protocols. A standing DIMENSIONAL AUDIT table will classify every recipe element. LADDER (frozen 72M recipe, no per-size tuning, class-1 data): C128/11M, C192/18M, C320/36M (new, EP+BP twins, n=2 each, Chinchilla 20 tok/param) + existing C512/72M, C768/135M = five points spanning 12x params. Queued behind dg128 on all 3 GPUs (~9h). PROTOCOL (strengthened): fit any law on the SMALLEST points only; blind-predict upward rung by rung (C320 -> C512 -> C768 -> C1152); the x128 dose at C768 is a mandatory PREDICTION target. Log prediction error vs extrapolation factor = the horizon curve. Horizon >= 3x -> 1B (1.8x beyond C1152) is within demonstrated range; horizon < 1.5x -> the honest verdict is "EP-transformer experience does not compress into transferable laws", which is decisive for the hardware thesis too (a chip is a frozen recipe). After trainings: M1 displacement spectroscopy + M2 per-nonlinearity linearization battery + M3 optimizer-step differentials at every width; PREREG freeze before any validation run. ### RESULT 73 (2026-07-29): GOVERNING DOCTRINE (user ruling) — NO TRANSFERABLE SCALING KNOWLEDGE EXISTS YET; positive scaling claims SUSPENDED; results reclassified into three strict classes; next phase = mechanism identification with a pre-registered transfer law and a BLIND holdout. All prior "progress" on the C768 leak is symptom-and-intervention, not root cause: amp (demoted), head throttle (retracted), softmax saturation (refuted) — a chain of per-scale patches with no unifying mechanism = an invalid scaling experiment by definition. THE THREE CLASSES (every past and future number must declare its class): 1. FROZEN-RECIPE CURVE (valid, negative): the fixed recipe fails C512->C768 (gap 0.047 -> 0.294); small-width frozen gaps already trend with width. This is the one established fact. 2. POST-HOC ORACLE ENVELOPE (repairability only, ZERO extrapolation power): per-width tuned results incl. every leak-battery arm and dgain dosing. dg128 at cos 0.25-0.28 is NOT the original estimator made precise — it is a DIFFERENT update field; if it works it is a new algorithm candidate and may not silently rejoin the original method's scaling curve. fw135m_dg128 (running, left to complete per user) is EXPLORATORY regardless of outcome — also a 250k-resume, so at best it proves rescue-ability of one state. 3. TRANSFER CURVE (the only thing called scaling): a control law with no explicit size dependence, fitted at small widths, hitting an UNSEEN width with ZERO tuning. THE LAW MUST SIMULTANEOUSLY EXPLAIN: top-half localization; the 4-6x C768/C512 leak-rate factor; leak ~ 1/beta; the x128 displacement requirement; immunity to bsign/centered/fp32/ K8/head-LR; complete closure under BP substitution. PROTOCOL (frozen before any validation run): C256/C384/C512/C768 for mechanism ID only — per-layer normalized displacement RMS(d_l)/RMS(z_l), Jacobian/curvature at each nonlinearity family, true optimizer-step differences; LOCAL LINEARIZATION battery per nonlinearity class (softmax vs SwiGLU vs RMSNorm — substitute frozen local linearization in the read, see which substitution kills the leak = which family carries the width-growing term); fit the dimensionless law (e.g. gain from RMS(d)/RMS(z) x local curvature scale — never a hand-written per-width constant); C1152/270M = blind holdout, formula/thresholds/cost/error-tolerance frozen BEFORE launch, from-scratch, n>=2 seeds; any unpredicted phenomenon at the holdout = verdict "does not scale", no patch-and-reconnect. Project page reworded same day: 135M row states the frozen-recipe failure plainly; scaling section = negative result + protocol; "97% treatment" claim removed from the abstract. ### RESULT 72 (2026-07-29): THE CURVE GOES ALL THE WAY — dg128 closes 97% (tail within 0.0015 of the all-BP bar); the physics leak is FULLY REDEEMABLE by displacement amplification. FULL-DOSE VALIDATION LAUNCHED. Battery 8: dg64 92%, dg128 97% (log-law held the whole way: 34/50/63/75/92/97 at x4..x128), dgrand64 64% (spread-spectrum loses to fixed dose — the value lives in the high octaves; log-uniform wastes steps low). dg128 mechanics: r18/6k mild, no skips, drift 0.009; gn 5.07 (amplified-read apparent norm — clip + msign both immune); cos 0.28 with BP-equivalent CE — the direction-metric coffin nailed shut. STANDING QUESTION (the honest one for the ladder): does the required dose scale with width? x128 at C768 vs ~x1 at C512 — if dose ~ width^k the wall is soft but real; if it saturates, this is a cure. 270M will answer. Production note: dgain is estimator-side only — model, inference, cost, hardware story all untouched; hardware realization = per-layer nudge amplifier gain (standard analog practice). IN FLIGHT: fw135m_dg128 (bsign + dgain_top 128, s250000 -> 440k, 3-GPU, ~22h). Bars: untreated 3.3839 | BP 3.0902. Success = endpoint ~3.15-3.18, gap back in the 72M band. ### RESULT 71 (2026-07-29): DOSE CURVE — logarithmic recovery (34/50/63% at x4/8/16), DISPLACEMENT-EQUIVALENCE LAW triple-confirmed, cos retired as a metric; TREATMENT VALIDATION LAUNCHED (bsign + dg8top, full remaining segment). Battery 7: dg8top closes 50% (cos 0.63), dg4-x-beta6e-3 closes 50% (cos 0.84, tail IDENTICAL 3.5439=3.5439 — route-independence again), dg16top 63% (cos 0.47, CE still improving). Laws: (1) CE depends only on the total top-half displacement product, not the route (three exact equalities now); (2) recovery ~ +13-16% per dose DOUBLING = thresholds spread over decades, log-slow to full closure, practical zone 50-70%; (3) anchor contamination (cos down to 0.47) is benign noise vs the systematic threshold signal — direction metrics carry zero predictive weight in this campaign, final. IN FLIGHT: fw135m_dg8 (bsign + dgain_top 8, s250000 -> 440k, 2-GPU, ~28h) — the treated 135M number; predicts endpoint ~3.25-3.29 if the 50% rate cut integrates (untreated 3.3839, BP 3.0902). GPU3: dg32top curve tail (log-law predicts ~72%). ### RESULT 70 (2026-07-29): FIRST REAL LEVERS — displacement amplification recovers the leak (dg4top +34% at cos-0.83 price, CE judges); the beta benefit factorizes to 100% DISPLACEMENT, 0% force; ffn carries ~half with an attn-ffn interaction term. Battery 6: leak_dg4top closes +34% (top-half state-formation d x4; despite gate cos 0.83 — FA-logic vindicated, CE is the judge); leak_dg2all +24% with tail IDENTICAL to b6d (3.5558 == 3.5558): doubling displacement-only reproduces doubling beta exactly -> the 1/beta dose response is entirely the DISPLACEMENT channel (force/SNR contributes nothing measurable); leak_mixffn +48% (attn 67 + ffn 48 > 100: overlapping/non-additive at C768, unlike R44's clean 72M additivity — an interaction term exists). MECHANISM PORTRAIT (refined): a displacement-threshold phenomenon distributed across top-half attn AND ffn nonlinearities — response components that finite probe displacement under-reaches, linearly redeemed by amplifying displacement through ANY route. Not softmax logits (R69). Pure estimator-side treatment exists (dgain: model function, inference, cost all untouched; hardware = per-layer nudge amplifier gain, standard analog practice). Battery 7 IN FLIGHT: dose curve dg8top/dg16top + stack dg4top-x-beta6e-3. If the curve reaches 60-80% recovery before contamination turns, the 135M treatment recipe = bsign + dgain-top(opt) at 1.0x, full-run validation next. ### RESULT 69 (2026-07-29): SATURATION-BLINDNESS REFUTED BY ITS OWN TREATMENT TEST; the leak survives exclusion #11; dgain (displacement-only amplification) is the live arm. Battery 4 (beta dose): leak INVERSE in beta — 1.5e-3: 0.0615, 3e-3: 0.0462, 6e-3: 0.0352 (~1/beta additive signature). This motivated the saturation-blindness hypothesis (sharp softmax switches invisible to finite displacement; the sibling Hopfield-EP "saturated units EP!=BPTT 80-130deg" lesson at transformer scale). Battery 5 (user-approved, matched-cap BP bar): logit softcap 30/50 (Gemma-2 style; firing certificate: cap=1 destroys the model, val 6.45) closes 0.0000 of the leak — relative leak 0.0462 IDENTICAL to 4 decimals. Saturation-at-the-logit level is NOT the mechanism. BONUS: the cap's own price under BP ~ 0 in-window (allbp+cap30 3.5205 vs allbp 3.5206) — the user's "cap hurts performance" concern priced at ~zero here (Gemma prior confirmed). SURVIVING CONSTRAINTS: top-half (100%), attn-dominant (2/3), ~1/beta, immune to: logit cap, odd/even bias, read noise, K, precision, head surgery, bottom blocks, magnitude/clip, Muon, direction (FA argument). Threshold nonlinearity still indicated by 1/beta but NOT at softmax logits — candidates: ffn gates (mixffn arm measuring its share), norms, qk-norm. Battery 6 IN FLIGHT: leak_dg4top (top-half state-formation d x4, cotangents true-d; smoke shows the price: gate cos 0.83 from anchor second-order — CE judges), leak_dg2all (displacement-vs-force discriminator vs the b6d arm which doubled BOTH: b6d 3.5558), leak_mixffn (completes the R44-style partition additivity). ### RESULT 68 (2026-07-28): LEAK LOCALIZED — 100% in the TOP-HALF block gradients, ~2/3 attention; the SAME map as R44 at 72M, 6x stronger at width. One structure, two scales. Battery 3 (guaranteed-signal partition, bars allbp 3.5206 / EP 3.5668): mixtop closes 100% (tail = allbp bar exactly), mixbot 0% (tail = EP bar exactly), mixattn 67%. The R44 endgame decomposition (77/0/41-37) reproduces on the leak instrument at 135M. The 0.999-cos paradox now has an address: top-half sharp-attention EP reads carry a ~1e-5/step systematic CE deficit, invisible to every per-step metric, killed only by BP substitution. Suspect state after batteries 1-3: NOT odd-order bias (bsign/centered), NOT read noise (centered/fp32), NOT relax depth (K8), NOT the head (headmix/hlr), NOT bottom transmission (mixbot 0%). Remaining channels for a top-local, symmetric-in-beta(?), BP-substitutable deficit: anchor-displacement curvature interaction at sharp softmax, or d-signal structure at the top read. BATTERY 4 (beta dose 1.5e-3/6e-3, same harness): leak ~ beta -> displacement channel confirmed -> treatment = top-local beta reduction (one flag); leak flat in beta -> displacement out, next = top-local read variants. ### RESULT 67 (2026-07-28): INSTRUMENT RECALIBRATED (1/3 harness confound, 2/3 physics) + BATTERY 2 ALL-NEGATIVE — the head-throttle is a SYMPTOM, not the cause; the 0.046 physics leak survives every single-component surgery. Localization battery (guaranteed-signal) in flight. Battery 2: leak_headmix (head on TRUE BP grads, config-verified) closes 0%; head_lr_mult 1.67/2.5 closes -1%/-3%. R66's causal chain RETRACTED (the 60% throttle measurement stands, but fixing it buys nothing — the head lags because of something upstream, or harmlessly). INSTRUMENT CONTROL leak_allbp (EP harness + ALL-BP grads, same data order): 3.5206 vs casc_bp_train 3.4982 vs EP arms 3.5668 -> the 0.0685 raw leak = 0.0224 harness/data-order artifact (33%) + 0.0462 gradient-physics (67%). Battery 1-2 conclusions survive on the recalibrated denominator: five EP configs all at 3.5668 = 0% of 0.0462 closed. Lesson banked: cross-harness comparisons REQUIRE the same-harness control arm first (amp-probe lesson, trainer edition). Battery-arm discipline: reference configs must pin every factor explicitly. BATTERY 3 (localization, must-split by construction): bpmix blocks:6-11 / blocks:0-5 / attn. ### RESULT 66 (2026-07-28): THE HEAD IS THROTTLED TO 60% — user's equivalent-LR hypothesis CONFIRMED and localized to ONE matrix; symmetrization/precision/depth all irrelevant to it. Displacement micro-probe (600 steps from s250000, ||dW||/||W|| per family): every family IDENTICAL EP-vs-BP (ratio 1.00-1.02) EXCEPT W_out: EP 0.0068 vs BP 0.0114 = 60%. Muon-managed blocks are magnitude-immune (msign); W_out lives on the Adam island where step ~ m/sqrt(v) — the EP head-read's per-coordinate temporal coherence is lower, v inflates ~2.8x, Adam throttles the loss-facing matrix by 40%. Explains R44 (gap lives head-side), the anneal freeze (head can't track features), and width-scaling (louder states -> stronger throttle). BATTERY 1 (all closed ~0% of the 0.0685 leak): centered 0%, fp32 +3%, K8 (arm died, rerun pending) — the throttle source is INVARIANT to estimator symmetry, precision, and relax depth: it is the shared structure of the EP W_out read, not removable noise. (cos flat 0.986-0.988 through the losing segment; EP gn 0.15, clip never fires — user's checklist all measured.) FA-argument (user): static direction quality exonerated a priori — FA learns at cos 0.3. BATTERY 2 IN FLIGHT: leak_headmix (--bpmix head = the head's share of the leak, diagnostic ceiling), leak_hlr17 / leak_hlr25 (--head_lr_mult, the zero-cost compensation candidate). If hlr claws back what headmix shows, the width fix is ONE hyperparameter. ### RESULT 65 (2026-07-28): THE LEAK INSTRUMENT WORKS — first reading +0.0685/6k at the SAME state (20x the trajectory-averaged rate: front-loaded/self-adapting), screening battery 1 in flight. Paired 6k segments from fw135m_bsign_s250000: BP tail val_mean 3.4982 vs EP-plain 3.5668. Caveats favoring conservatism: BP arm ran COLD momentum (weights-only resume) yet still outpaced by 0.0685 -> underestimate of BP's edge; the trajectory-integrated gap accrues 20x slower (EP re-equilibrates along its own path — the leak is partially self-healing, which is WHY the run survives at all). As a SCREENING instrument: 1.5h/arm, single GPU, huge SNR. BATTERY 1 (all from s250000 warm, 6k, vs the two bars): leak_cent (centered — if the exact- symmetric read still leaks, the entire estimator-bias family is out); leak_k8 (deeper relax — cos said saturated, the leak may live below cos resolution); leak_f32 (amp share on the same instrument, cross-check of the ~12% window estimate). ### RESULT 64 (2026-07-28): NOT A PLATEAU — A CONSTANT-RATE RELATIVE LEAK, 4-5x wider at C768; every named suspect cleared AT PROBE RESOLUTION; the leak-rate probe is the new instrument. LEDGERS CLOSED: 72M BP n=3: 3.2884/3.2943/3.2820 -> mean 3.2882+-0.006. Honest 72M gap: bsign(n=2) +0.076 (7.9% ppl); cent(n=1) +0.044. — fp32 A/B (both arms died at s275000 torch.save: DISK FULL, 100%/10T shared NFS; pruned 458GB of intermediate ckpts, 652 files, manifest runs/PRUNE_MANIFEST_260728.txt, 39 landmarks kept): fp32@3e-3 window val_mean 3.5538 vs amp 3.5891 vs fp32@1.5e-3 3.5843 -> amp DEMOTED to minor contributor (~12% of the deficit); R63's "amp is THE mechanism" corrected. SUSPECTS CLEARED (with resolution caveats): systematic bias at C768 <= sem (probe_bias nb=12: plain ~ sem everywhere; w2_b11 0.81% vs sem 0.51% = the one 1.6x flicker); Muon amplification NONE (msign cos 0.9987-0.9994 vs raw 0.9998-1.0000); sigma/scale runaway NONE (tok_rms IDENTICAL to 3 decimals; BP head LOUDER: sig1 744 vs EP 620 — the louder-states trajectory is normal training, not EP pathology). THE SHAPE (schedule-fraction matched): 72M gap hovers +0.04..0.09 and the anneal CLOSES it (0.087@80% -> 0.047@100%); 135M gap grows near-linearly +0.10/0.16/0.20/0.28/0.29 — a per-progress leak, 4-5x the 72M rate, that the anneal cannot outrun; "plateau" = leak rate catching BP's descent rate. Required per-step magnitude: ~1% CORRELATED discrepancy — exactly at/below every probe's floor. The 0.999-cos-yet-diverging paradox is the finding. IN FLIGHT: leak-rate paired probe — SAME weights (s250000), EP vs BP, 6k steps each, direct DCE readout; calibrates the leak instrument for recipe search (anything that cuts measured leak-rate is a candidate; full-run validation only after). ### RESULT 63 (2026-07-27): MECHANISM CLOSED — THE AMP CERTIFICATE BREAKS AT WIDTH AND THE C768 CORRIDOR INVERTS MID-RUN; every 135M EP failure mode falls into place. Chain of three probes (all GPU0, hours): (1) blockcos @s150000: fp32 EP gradient PERFECT (cos 0.9999-1.0000, every block, both channels) -> per-step fidelity exonerated at healthy states. (2) K-sweep @s250000/s435000: K=3 == K=8 at 0.999 -> K-sufficiency transfers; K exonerated. (3) amp_gate rerun on C768 ckpts — THE SMOKING GUN: amp cos @1e-3 = 0.9079 (C512: 0.9682 == fp32), @3e-3 ~0.95 (C512: 0.9878), @1e-2 0.9761 (C512: 0.9966). The bf16 state-noise is fixed-relative, the nudge displacement ~beta, the states are LOUDER at C768 (sigma 661 vs 475) -> the amp read-noise floor RISES with width. The BBP "no digital detectability floor" result was fp32-ONLY; under amp there is a real, width-scaling floor. THE INVERSION: amp floor wants beta >~3e-3 for clean reads; the fp32 ceiling sinks to ~1e-3 late -> from mid-run onward NO legal beta exists under amp at C768. All three EP arms explained: 3e-3 = over-ceiling; 1.5e-3/7.5e-4 = under the amp floor (b15's unknown limiter FOUND — its in-corridor-but-noise-rotten reads). The 0.294 gap = 200k+ steps of systematically degraded reads, not any single blowup. RECIPE CONSEQUENCES for width: (a) fp32 relax (tax->0, cost 1.5x) — the safe fix, A/B pending; (b) bf16 sweeps + fp32 polish sweeps (~1.2x) — needs a probe variant (naive amp_last is WORSE: one fp32 rebuild on bf16-settled states reads the bf16 offset as displacement); (c) large-beta escape is CLOSED by the sinking ceiling. The 1.5x amp lever that funded the cost model at 1-3B is width-broken — COST_MODEL must be re-quoted (fp32 or polish-mix rates) before any external number is repeated. NEXT: 72m_bp_s3 clean (GPU0, closes 72M n=3); amp-vs-fp32 C768 A/B replay from s250000 when GPUs free. ### RESULT 62 (2026-07-27): THE C768 ANCHOR — BP 3.0902; the EP gap EXPLODES 6x at the first ladder rung (0.047 -> 0.294, ~34% ppl). Width problem is 100% EP-specific; recipe/data/ shape exonerated (BP gains the full expected -0.198 from 72M). fw135m_bp (exact EP twin, 440k/2.7B tok, seed 1): 3.0902. Ledger: 72M BP 3.2884/3.2943 (n=2, band 0.006) | 135M EP best-of-3-arms 3.3839 (3e-3 lineage; 1.5e-3 rescue 3.4643; 7.5e-4 endgame 3.4805). HONEST CAVEAT: no clean C768 EP run exists — first run spent its second half over a sunken ceiling, b15 was a mid-run rescue with constant backstop halvings and an unknown limiter; "current recipes fail at width" is measured, "EP fails at width" is NOT yet. The scaling thesis now hangs on whether a width-correct recipe exists inside the C768 corridor. NEXT (GPU0, user holds 1/3): gap DECOMPOSITION at C768 — probe_blockcos on fw135m_bsign s150000 (healthy phase): if the gap still lives in top-block transmission but scaled up, the width amplification of the known transmission bias is the target; then 72m_bp_s3 clean restart. ### RESULT 61b (2026-07-25): CORRECTION — the sentinel's 21 DIVERGE alerts were FLOOR-NOISE FALSE POSITIVES; b15's corridor at 1.5e-3 was HEALTHY through 300k+; its true limiter is UNKNOWN. Row audit: s255000 @1.5e-3 resK=1.40e-6 (= fp floor, fully converged), rho_tail 1.0032 — the documented rho~noise/noise-at-floor artifact, which I had discounted manually at s200000 but failed to encode in the sentinel verdict rule (instrument-made ghost evidence). R61's claim "compliant lineage also sank its ceiling to 1.5e-3" is RETRACTED. What actually bounded b15 (best 3.4643; r18316 amp-noise backstop halvings at res>0.02 = frequent transient beta dips; late val 3.7-3.9, gn 0.32) is an OPEN question, deferred behind the BP anchor per the user's order. Probe verdict rule fixed: rho>1 only counts when resK is materially off-floor (>1e-4). Wider lesson for the dossier: at C768 the ONLY trustworthy stability instrument so far is the offline fp32 probe read by resK-to-floor; every in-band and rho-based verdict has now produced at least one false conviction. ### RESULT 61 (2026-07-25): EP-SIDE DIAGNOSTICS HALTED FOR THE BP ANCHOR (user call: "这种情况下应该先跑bp baseline"). b15 postmortem seals the C768 picture: the COMPLIANT beta=1.5e-3 lineage ALSO sank its ceiling to 1.5e-3 (sentinel: operating row DIVERGE for 5 consecutive ckpt probes; r18316 backstop halvings; best 3.4643 @387.8k, still short of the 3.3839 first run). The C768 sinking chases beta down — every EP arm so far (3e-3 fixed, 7.5e-4 endgame, 1.5e-3 compliant) fails differently, and NONE of it is interpretable without knowing what BP does at this width. ATTRIBUTION FIRST. IN FLIGHT (all 3 GPUs, BP battery): fw135m_bp s1 (GPU0) + s2 (GPU1) — exact EP twins, 440k/2.7B tok, ~41h each; fw72m_bp_s2 -> s3 chained (GPU3, ~11h each) — closes the 72M honest-gap denominator (currently n=1). Readout: (a) 135M BP endpoint vs 72M BP 3.2884 = does the RECIPE scale at all (if BP also disappoints, the width problem is shared, not EP's); (b) honest 135M gap = EP best vs BP twin mean; (c) 72M gap ledger goes n=3-vs-n=2. ### RESULT 60b (2026-07-24): THE RES METER IS WIDTH-BLIND UNDER AMP — in-band ratcheting is impossible at C768; the instrument is the offline fp32 probe. Ratchet arm killed at launch (6 knocks in the first step, beta -> 4.7e-5 death spiral: gate 1e-3 sat BELOW the bf16 floor). Measured amp-ensemble in-run residual (resdiag, 60-step diagnostics): healthy s250000 median 9.9e-3 vs dead-corridor s345000 median 1.2e-2 — the fp32 4-decade separation (1.3e-6 vs 2.6e-4+) collapses to 1.2x under bf16 forward noise. The old 0.02 constant "worked" at C512 by sitting above the amp floor; at C768 NO res threshold separates healthy from sick. The fp32-probe-vs-amp-trainer ensemble mismatch strikes again (BBP lesson, meter edition). CONTINUATION: fw135m_b15 — s250000, FIXED beta 1.5e-3 (3x margin at 250k ceiling ~5e-3; in-corridor at 300k ~3e-3; damaged-lineage endgame reading 1-1.5e-3 = lower bound), bsign, res_gate 0.02 as NaN backstop only, 3-GPU, ~27h. SENTINEL: every new ckpt gets an fp32 rho-probe at {1.5, 2.25, 3}e-3 — alert if the operating row goes DIVERGE (the calibrated instrument, out-of-band). Beta policy doctrine for the ladder: measured fixed points + offline-probe surveillance; in-band adaptive control is dead at width under amp. ### RESULT 60 (2026-07-24): s345000 CORRIDOR IS EMPTY (end075 3.4805 < original, worse); closure located at ~275-300k; CONTINUATION = bsign+ratchet+res_gate from s250000. Two code fixes en route (sign-wipe combo bug + configurable gate), both fire-certified. end075 (beta 7.5e-4 from s345000): best 3.4805, val climbing to 3.9, gn doubled — LOWER beta did WORSE than the original 3e-3 over the same segment. LR-scheduler suspect CLEARED (resume fast-forward exists, line 813). Diagnosis: at s345000 the corridor is EMPTY — ceiling ~1e-3 (contraction) while the state's elevated relax-noise floor (resK stuck 2.6e-4 at ALL beta, 100x above healthy) pushes the effective lower bound UP past it. Over-ceiling operation bakes in state damage (the 72M simple-lineage lesson, width edition). CLOSURE PROBE (resK-to-floor reading, rho-noise verdicts discounted): s250000 CLEAN at 3e-3 (resK 1.3e-6, ceiling ~5e-3) | s300000 boundary (3e-3 resK 2.2e-5) | s345000 dead. Crossing ~275-300k. Last healthy resume point = s250000. FIXES: (1) --res_gate flag (0.02 was a C512 constant; C768 ran 100k steps semi-converged under it; per-width rule ~100x healthy K=3 floor -> 1e-3 for C768); (2) bsign+beta_simple SIGN-WIPE bug — the beta_simple override sat after the sign flip and silently disabled bsign when both were on (never manifested: 72M ratchet ran plain). Flip moved after the override. COMBO_CERT_PASS: forced-knock smoke shows sign flips + knocks + persistent halving together. IN FLIGHT: fw135m_ratchet — resume s250000, bsign + halve-only ratchet + res_gate 1e-3, beta0 3e-3, 3-GPU, 190k-step replay (~27h). Expected: knocks track the sinking ceiling (~290k, ~1.5-2e-3, endgame ~1e-3), anneal harvested inside the corridor. Verdict = best vs 3.3839 (bsign fixed-3e-3) and vs 72M band 3.364. ### RESULT 59b (2026-07-24): C768 CEILING MEASURED — the run spent its second half 2-3x ABOVE the ceiling, and the ceiling sinks UNDER BSIGN at this width. R59's "graze" and "stability transfers" both CORRECTED. rho-grid @C768 (fine grid 0.5-20e-3, K=30): ceiling beta* = 8-12e-3 @s150000 (3e-3 healthy, 4x margin) -> ~1e-3 @s345000 (3e-3 = 3x OVER; 1.5e-3 diverges, resK 1.1e-2) -> 1-1.5e-3 @s435000. So: (1) not a graze — sustained over-ceiling operation, fail-soft only because bsign cut the bias feedback and guards were silent (plain_nf regime: no explosion, no learning); (2) the ceiling sinks even with bsign at C768 — sharpening itself (sigma 569->661) drives it; 72M's "debiasing removes the sink" narrows to "removes the BIAS-driven component" — width adds a second, bias-independent sinking term. NEW LADDER PHYSICS: ceiling scale drops with width (~1/sigma^2); at 270M+ even 1e-3 may cross mid-run -> the recipe needs per-stage beta descent gated by measured ceilings (the ratchet/descent axis returns, this time with the instrument to calibrate it). Also non-transferring: the res>0.02 legality constant (C768 ran half its schedule at res~1e-2, under the alarm line, while the harvest died). end15 (beta 1.5e-3) KILLED pre-data (measured above-ceiling at the resume state); fw135m_end075 relaunched: s345000, beta=7.5e-4 (below both late readings), final 95k replay. ### RESULT 59 (2026-07-24): FW135M COMPLETED 3.3839 — stability TRANSFERS (zero events, 440k, bsign), but the OPERATING POINT does not: beta=3e-3 grazed the C768 ceiling through the endgame and the anneal harvest was lost. Scaling verdict DEFERRED pending the beta-fix replay. Numbers: best 3.3839 @347.2k (79% of schedule; final 93k deep-anneal steps bought NOTHING; val_mean flat 3.593->3.581 over the last 300k). Grazing signature, monotone by window: cos 0.990->0.981->0.968, drift 0.0097->0.0173, gn 0.146->0.197, while sigma climbed to 661 (72M endgame: 475, cos 0.995). Mechanism: wider model -> larger sigma -> ceiling ~1/sigma^2 LOWER, and the 72M-calibrated beta=3e-3 met it. bsign held (zero d-skips, no storm, no collapse) = the stability claim extends to 135M; but estimator quality degraded in the graze and the last-fifth harvest vanished. Endpoint sits INSIDE the 72M seed band (bsign mean 3.364+-0.03) = a FLAT scaling step as measured, n=1, attribution impossible without the 135M BP twin (collaborator, now urgent). LESSON (ladder doctrine): window constants do NOT transfer across width — every rung needs a beta re-gate (rho-grid at that width), exactly as BASELINE_SPEC prescribed for depth events. Candidate width rule: beta ~ 1/sigma_end^2 (C768 sigma_end/C512 ratio^2 = 1.94 -> beta ~1.5e-3). IN FLIGHT: (a) probe_cx3_rhogrid on fw135m ckpts s150000/s345000/s435000, fine beta grid — the direct ceiling measurement at C768 (physics certificate for the graze); (b) fw135m_end15: resume s345000 (just before best), beta 1.5e-3, replay the final 95k (2-GPU, ~16h). Harvest returns (best << 3.3839) -> diagnosis confirmed + better crown + the width-scaling beta rule enters the recipe; no return -> C768 endgame is genuinely harder, escalate to estimator (centered tail) or schedule work. ### RESULT 58 (2026-07-22): BSIGN SEED-2 — STABILITY REPLICATED (n=2, zero events both runs); CE mean 3.3637 +- 0.029, s1 was the lucky draw (dip-statistics bite #4). fw72m_bsign_s2 DONE 3.3924: full 234k, d0/g0/r0 end to end under the res-alone gate — the corridor-holds claim is now two-for-two; the load-bearing stability result is sealed. CE honest ledger at 72M now: bsign mean 3.3637 (n=2) | cent 3.3318 (n=1) | plain2 3.3437 (n=1, stalled) | BP 3.2884 (n=1). Single-seed spread at 72M is +-0.03-0.06 -> every n=1 number above carries that band; the bsign-vs-BP gap is UNQUOTABLE until BP seeds land (collaborator first task, BASELINE_SPEC). Flagship pair (cent s1 vs BP s1, DCE 0.043) stands as a sealed-run fact. 135M migrated to 3-GPU DDP (B8x3) from s60000 after s2 freed the pair. ### RESULT 57 (2026-07-21): RATCHET VERDICT — pure retreat cannot outrun the bias-driven sink; the staircase IS the plain-lineage ceiling trajectory, and it chases beta to the bf16 floor. The trio closes: debiasing is the only 1.0x recipe that holds. fw72m_ratchet (user design: halve-on-crossing only, no climb, plain estimator, res-alone gate, guards silent): mechanically flawless — zero explosions, knocks fired exactly as designed. Staircase: 3e-3 held to 88.4k (first knock -> 1.5e-3), held to 131.2k, then CASCADE (5 knocks in 3.2k steps -> 4.7e-5 by 134.4k), then 2.3e-5@140k, 5.9e-6@152k, 1.8e-7@154k = sub-bf16- cliff, training dead-walks (val 3.42 best -> 4.28 noise-walk). KILLED at 181.4k (staircase complete; remaining steps carried no information). GPU0 -> fw72m_bp_s2 (honest-gap seed). READING: under the SAME res-alone gate, bsign saw ZERO events in 234k while ratchet's ceiling collapsed through 4+ decades from 131k. Same estimator-family difference as plain2-vs-bsign, now in the adaptive-beta version: the sink is a property of the BIASED estimator's trajectory, not of the beta policy. No beta policy — fixed (plain2), descending schedule (would-be betacos), or measured retreat (ratchet) — survives it; removing the bias (bsign) removes the sink. Bonus instrument value: the staircase is the first continuous in-vivo ceiling-trajectory measurement (a plain lineage's ceiling vs step curve, free of controller-climb artifacts). LADDER STANDS: cent 3.3318 (1.67x) | bsign 3.3349 (1.0x, zero events) | plain2 3.3437 (stall) | ratchet 3.4200 (stall@~131k). 135M gate now = bsign_s2 confirmation only (102k/234k, clean). ### RESULT 56 (2026-07-21): BSIGN CROWN SEALED 3.3349 — THE 1.0x-COST CROWN EXISTS; the sink is BIAS-DRIVEN (user's "no effect" bet resolved AGAINST, happily). Sign randomization alone (Scellier-Bengio random nudge sign, one flag, zero extra phases) held the corridor for the FULL 234k schedule with ZERO guard events end to end (d-telemetry 0 at every step, under the STRICTER res-alone gate in silent mode) — where plain2 (identical recipe minus the sign flip) skip-stormed from 217k (84-89%/step, 12840 events, training frozen after 199.8k). LADDER: cent 3.3318 (1.67x) | bsign 3.3349 (1.0x) | plain2 3.3437 (1.0x, stalled@200k) | c190 hybrid 3.3732. bsign TIES the centered crown within 0.003 (noise) at 60% of its cost. best@203.7k — the SAME dip step as cent's best; endgame keeps taking updates (val ~3.46 at 234k, healthy oscillation, final gate cos 0.9958). sigma trajectory identical to plain/plain2 (475 endgame) => sigma exonerated again; the causal chain closes: single-sided systematic bias (measured 0.3-1%@3e-3, RESULT 55) -> trajectory steering -> ceiling collapse; remove the bias in EXPECTATION (per-step sign flip) and the collapse never comes. The entire controller saga's answer was a coin flip on beta's sign. Note vs BP: bsign gap 0.047 vs cent's 0.043 — cent remains the flagship gap number; bsign is the cost claim. NEXT: fw72m_bsign_s2 LAUNCHED (seed discipline — three single-seed crowns have been retired this campaign; no external claim before n>=2). PSIGN (per-sample sign, RESULT 55-verified unbiased, batch-averaged) = the scale-native production form, candidate recipe for the 300M stage. Ratchet arm continues as the no-debiasing control (its staircase should track the bias-driven sink that bsign no longer has). ### RESULT 55 (2026-07-20): SINGLE-SIDED BIAS DIRECTLY QUANTIFIED (the BBP-thread gap) — real, linear in beta, 0.3-1% of grad norm @3e-3; PER-SAMPLE SIGN + BATCH AVERAGE = UNBIASED (user's claim VERIFIED at zero measurable cost). probe_bias @plain2_s35000, 16 batches, beta {3e-3,1e-2}, selfcheck 3e-9: plain rel_bias |mean(gEP-gBP)|/|mean(gBP)|: w2_b11 0.0035->0.0114 (x3.26 ~ beta ratio 3.33, 2.7sigma over sem 0.0013->0.0042); qkv_b0 0.0106->0.0325; all four layers >sem with ~x3.1-3.4 scaling => the Laborieux O(beta) term measured directly on a 72M transformer. PSIGN (per-sample random nudge sign, read re-flipped): bias -> noise floor in EVERY row (w2_b11: 0.0014 vs sem 0.0014), both betas; sem unchanged vs plain => debiasing is FREE (no extra phase, no variance penalty at B=8; improves with B — the scale-native estimator, supersedes per-batch bsign in theory). Batch size alone does NOT debias (bias survives expectation); it only debiases AFTER the per-sample sign flip makes the O(beta) term batch-antisymmetric — user's "sign+大batch=无偏" is the correct composition. PRE-REGISTERED corollary: bias is only 0.3-1%, so if the RUNNING bsign crown fails to stop the ceiling sink (user predicts it will fail), cent's armor is NOT debiasing — it must be a variance/geometry property of the symmetric read; new puzzle. In flight: fw72m_bsign (GPU1/3, per-batch sign), fw72m_ratchet (GPU0 queued, halve-only). ### RESULT 54 (2026-07-20): THE CEILING SINKS ∝1/σ² AS TRAINING SHARPENS — no fixed β survives; "fixed β never explodes" was FALSE (it STALLS: endgame skip-storm = frozen training). β must DESCEND. User caught it: "固定β永不炸个屁!你一直skip还训不训了?" — dead right. EVIDENCE: plain2 fixed-3e-3 full run — best 3.3437 hit at step 199800, then skip rate SURGES: 84%/step @224k, 89%/step @231k. Endgame (200-234k) = 34k steps of ~zero effective updates; the "234k crown 3.3437" is really a 200k result. skip ≠ safe, skip = STALL. DENSE ρ-PROBE (probe_cx3_rhogrid, K=30 asymptotic — corrects Codex's K=8 β*=0.7 which mistook slow divergence for convergence): true ceiling β* (max β whose resK reaches the fp floor) SINKS: s35000 β*≈0.12 (σ242) | s95000 β*≈0.12 | s150000 β*≈0.08 (σ443) | s230000 β*<0.03 (σ473). MECHANISM: training sharpens the model (σ 242→473, ~2×); a sharper block Jacobian makes the nudge relaxation harder to converge → ceiling drops. It's ∝1/σ^~2. The sunk endgame ceiling (<0.03, approaching the 3e-3 training β) is WHY fixed 3e-3 skip-storms late. CODE: σ-scaling EXISTS (line 473: β=β0·SIG0²/σ²) — the ORIGINAL wall-1 design already tracks σ — but line 477 floor=3e-3 PINS it (plain2 wanted β↓1.94e-3 @σ473, floor forced 3e-3). My beta_simple "full ownership" that DELETED σ-scaling was exactly backwards. But naive un-flooring fails too: early σ 3.9→242 (62×, network growing structure NOT sharpening) would crush β∝1/σ² to 8e-7 and starve training — which is WHY the floor existed. The real tension: early σ-growth (structure, don't cut β) vs late σ-growth (sharpening, DO cut β); σ magnitude can't separate them. FIX (least-assumption): SCHEDULE β descent by progress like LR — --beta_cos_min added (cosine β from --beta to min, bypasses σ-scaling+floor). Calibrated to the measured ceiling: early 2e-2 (<0.12), endgame 8e-4 ( the entire adaptive-controller line (servo/ride/wsync/simple) was the wrong AND dangerous axis; drop it. - CEILING SINKS with training: β*=0.7 @s35000 (mid), but endgame ceiling ~3e-3 (plain2's 11527 skips at fixed 3e-3 = the sunk ceiling biting). So no single fixed β is "safe" in the sense of never skipping — but skipping ≠ exploding. - β-CURVE FLAT (RESULT bcurve, s35000+10k, controller-free): 1e-3→3.6305, 3e-3→3.6207, 1e-2→3.6255, 3e-2→3.6133. Δ0.017 over 30× = seed noise. NO SNR benefit to large β in the digital sim (consistent with the BBP no-floor finding). Large β only COSTS (single-sided bias ∝β, Laborieux; less headroom below the sinking ceiling). POLICY (corrected — my earlier "fixed large β" was backwards, user caught it): three constraints (bias∝β↓, sinking-ceiling↓, SNR-floor absent in digital per BBP) all point SMALL. Use fixed small β; controller removed. Digital ceiling never binds via explosion, only via defensive skips that a small β minimizes. HARDWARE is the only regime with a real SNR floor (additive readout noise → β*_HW by ENOB) — a separate calc, not settable a priori. VALIDATION: fw72m_b1e3 (fixed 1e-3, full 234k, controller-free) queued vs plain2 fixed-3e-3 (best 3.3437). hibeta 0.2/0.4 arms still measuring slow-accumulation (hypothesis B) at ρ≈0.95. plain2 crown = current best 1.0x-cost result at 3.3437 (vs cent 3.3318 @1.67x). ### RESULT 52 (2026-07-19): 42M 3x3 MATRIX COMPLETE — cent's "zero-gap" was LUCKY-SEED; estimator choice does NOT move final CE. cent seeds: 1.2334(s1)/1.2716/1.2579 mean 1.2543. Full table: BP 1.2141 | ride 1.2449 | cent 1.2543 | plain 1.2558 — the three EP recipes are statistically indistinguishable (means within 0.011, seed spread +-0.02-0.04); RESULT 46's cent0.02 AND rho>0.9) — "res huge but rho<0.9" (large-displacement CONVERGING relax, the big-beta mode) passed as legal: 35.4k-36k = 600 CONSECUTIVE first-try-legal verdicts while beta climbed x392 (=1.01^600 exactly) into damaging gradients (cos 0.999->0.9556->0.90). Plus: res/rho used rank-0 bcast while drift used global max (rank-1 crossings committed); NaN compares silently legal; kretry rescue didn't persist halvings (upward bias, 0% causal here); up-step fired before gn/second-drift guards; bsimp not checkpointed. AIMD arithmetic EXONERATED with exact numbers: 4k-34k halving rate 0.01423/step ~= ln(1.01)/ ln(2)=0.01436 equilibrium. 4-ARM REPLAY (s35000, 400 steps): plain2-lineage control @3e-3 = r0 zero events cos 0.9989; simple-lineage @3e-3 = r868 storm (its ceiling had sunk to ~1e-3, sigma 268 vs plain2's 242 — chronic boundary-riding damage is REAL as the amplifier); @9e-2 = d207, cos 0.569, K->8. Crown-3 shared the same contract failure (cos -0.21, drift 0.067, zero skips). FIXES APPLIED (commit): OR-gate + NaN-illegal + ddp_max; kretry persist; up-step at commit point; bsimp in ckpt. Consequence: future endgame runs will reject on res>0.02 alone (stricter, fail-visible). fw72m_plain2 crown resumed s95000 on the OLD gate (safe: its fixed 3e-3 never enters the res-huge/rho-low hole; its risk class res>0.02 AND rho>0.9 was provably caught, wsync-arm r10/r21@190.3k). simple-v2 (fixed gate) ready for a clean-state validation arm pending user go. ### RESULT 50 (2026-07-18): c190 SEALED 3.3732 — centered rescues the plain-lineage state through the full death window (beta pinned 3e-3, min graze 2.54e-3, zero skips). 2x2 complete: the disease is NOT baked into the s190000 weights; single-sided CONTINUED updates collapse the ceiling. HYBRID recipe (plain bulk -> centered tail) = 3.3732 at ~1.13x cost; carries the ~0.04 bulk transmission tax vs full-cent 3.3318 (1.67x). GOVERNOR AUTOPSY (user's verdict, correct): crown-3's kill was an adaptive-beta BUG stack — (a) cap hard bottom saturated 18k steps above the collapsed ceiling [fixed: cap_floor 0], (b) res/rho illegality gate wired only to beta_sync, crown-3 ran bare -> 18k garbage steps ACCEPTED on drift-gate-only [fixed: gate always-on when beta_cap_rho>0], (c) probe-validated protections (ride-v2/beta_sync) sat on the shelf [scheduling error, R42]. KNEE ARM CANCELLED (user: multi-constant fix rejected); replaced by --wsync = synchronous weight-step acceptance in the user's beta_sync idiom (snapshot params+momentum pre-step; next relax's SAME _legal gate judges the new state; illegal -> (p+snap)/2 halvings -> full revert + skip; zero new constants, model function untouched). CROWN-3b IN FLIGHT: fw72m_plain2, from scratch, fixed stack (cap_floor 0 + always-on gate + wsync 3), plain est, 234k, DDP GPU1+3. Detonation caveat: wsync SNAPSHOT path exercised every step; ROLLBACK path code-reviewed but first live fire will be ~196k. OPEN (user thinking + my proposal pending): ceiling-HUGGING beta controller — current AIMD-style (slow climb / halve) never sits at the ceiling; proposal = ODE-step-size-style servo on the linear plant rho ~= G*beta (G = rho^/beta observable every step, free): beta_next = safety*cap_rho / EMA(G) — deadbeat inversion + Gustafsson-PI smoothing, rejection gate stays as backstop. To validate in a cheap arm BEFORE any crown use. ### RESULT 49 (2026-07-18): CYCLE-GAIN DECOMPOSITION — the explosion anatomized (user directive). The loop is a ~23-stage product sitting near-critical by construction; plain heats ONE stage (b8's sharp head h6, logit 94 vs cent 75) until a MARGINAL crossing; the catastrophe was the governor's hard-bottom RESPONSE, which damaged b1 (vjp-entry gain 1.3 -> 14.9). probe_cycle (mode-aligned per-stage gains, sweeps 15/16 diffs, beta=3e-3): structure (BOTH lineages): down-chain vjp product ~50-72x, up-chain rebuild ~13-15k x, top read gH ~1e-6 (beta*NBT*H_top) -> round-trip O(1). Biggest single amplifier = b8: 2.9 (vjp) x 3.8 (fwd) ~ 11x round-trip; b7 next. Same map as R44/R47. plain vs cent @195k: same structure ~1.2-1.4x hotter, concentrated at the b8 stage (2.92 vs 2.50) and its carrier head b8-h6 (share 0.21, max|logit| 94 vs cent 75, ent 0.4); healthy runs CARRY 50-57-logit heads fine (cent b10-h7=57, b11-h0=51) -> threshold ~80+. plain @210k (sick): gV entry-stage 0<-1 EXPLODED 1.31 -> 14.91 (b1 sigma(J) 16->34->44 in statej) = the noise-walk DAMAGE, not the original cause. Pre-collapse growth was modest (res0 +55% over 40k steps); G's 300x conflates marginal crossing with post-crossing damage. BP-batch-noise control: per-batch BP u1 vs mean-BP u1 overlap 0.388/0.446 == EP arms' 0.37-0.41 -> the "shared noise" is minibatch sampling, identical for BP; Q1 sealed. FIX ARM ARMED: --logit_knee 80 (piecewise: identity below 80, slope 0.2 above; healthy heads untouched incl cent's own carriers; cuts the runaway head's switching Jacobian 5x). fw72m_plain_knee = exact plain_nf replay + knee, queued behind c190 (smoke then DDP launch). Also on the menu (user decision): est_auto = centered-on-demand only when rho_ema near cap (pays 1.39x on ~10-20% of steps, avg ~1.05x) — mechanism-blind but c190-proven lever. ### RESULT 48c (2026-07-18): WHY the digital floor can't be BBP — the THETA-READ HAS NO 1/beta CHANNEL, structurally. The BBP transition lives exactly where readout is DIFFERENTIAL (analog). probe_bbp3 (6-bit STOCHASTIC rounding, fresh draw per estimate; frozen = one draw for free+ nudged [trainer-faithful], split = independent draws [decorrelated bound]): R beta-FLAT 0.3-1.4 in BOTH modes, identical tables — even fully decorrelated operator noise produces NO 1/beta bulk. Mechanism: our estimator reads <-d/(beta*NBT), dO/dtheta> — d is proportional to beta, the beta CANCELS in the cotangent; quant/rounding errors enter as RELATIVE anchor/transmission shifts O(q), never divided by beta. The (E+-E-)/2beta catastrophic-cancellation channel simply does not exist in this code path. => Across ALL FIVE digital ensembles probed (fp32, bf16-amp, det-quant, stoch-quant frozen, stoch-quant split): no detectability floor; wall-1-in-sim = window geometry + bf16 displacement cliff. WHERE BBP IS REAL: hardware whose readback IS a differential measurement (delta-I/delta-V between phases / beta) — T64's transpose/column read IS differential -> additive ADC noise / beta = true 1/beta bulk -> beta*_HW = nu_ADC(sqrt m + sqrt n)/sigma1(g) stands as the design equation, calibrate nu from ENOB 8.51b. Caveats: R is an iid-frame index; 6-bit noise is STRUCTURED (correlated with signal directions: R<1 yet overlap 0.3-0.9, not the BBP zero) — structured noise degrades gracefully, R conservative. qb6 beta-buyback mechanism REASSIGNED to open (not 1/beta-noise; candidate: anchor-displacement-relative channel) — measurable later, not blocking. ### RESULT 48b (2026-07-18): GENERALIZED BBP (amp semantics, measured-spectrum, no iid assumption; user-directed): R = spike/edge > 1 EVERYWHERE down to beta=1e-4 — EP's extra noise sits 5-50x BELOW the signal spike in bf16-amp too. The digital floor is NOT a detectability transition; the binding single-step detection limit is BATCH noise (shared with BP). probe_bbp2 @s150000, 8 layers x 8 batches x beta {1e-4..1e-2}: R@1e-4 = 8.7-46.8 per layer, FLAT toward low beta (no 1/beta bulk: our theta-read never differences energies numerically — beta enters via state displacement, so bf16 bites as an all-or-nothing ROUNDING CLIFF at beta*||d|| ~ ulp(z) [RESULT 11's death], not as a noise bulk). u1-overlap ~0.37 (b0-b8) / 0.66-0.80 (b11) beta-INDEPENDENT => dominated by batch sampling noise (missing control column: per-batch BP u1 vs mean-BP u1 — same-batch cos(EP,BP)~0.99 implies EP u1 ~ BP u1, so 0.37 is the batch-noise number, not an EP defect). WHERE the sharp BBP transition genuinely lives: (a) ANALOG readout (ENOB additive bulk -> T64 design equation stands), (b) aggressive weight/ compute QUANTIZATION — qb6 beta-buyback is an in-hand crossing of the edge; controlled demo = rerun bbp2 with 6-bit qcomp, expect R to cross 1 and overlap to jump with beta. (c) bf16 cliff as the degenerate case. beta_floor's empirical benefit now fully attributed to window geometry (distance from cliff + T2-sensitivity), NOT detectability. ### RESULT 48 (2026-07-18): BBP FLOOR AUDIT — the fp32 SIMULATOR has NO additive detectability floor (a=0 all layers, top-direction overlap 1.000 down to beta=3e-4); BBP becomes a HARDWARE design equation. probe_bbp @plain s150000, 24 layers (qkv+w2 x12), 6 batches x beta {3e-4,1e-3,3e-3,1e-2}: additive coefficient a = 0 within fit resolution EVERYWHERE; residual error is beta-independent (b ~ 1-7e-7 entry-std, negligible vs sigma1(g) 0.002-0.01); u1-overlap 1.000 at all beta (slight dip at 1e-2, e.g. 0.898 on b0.w2 = finite-beta bias, wall-2 leakage, NOT noise). => The empirical floor's benefit never came through the detectability channel (consistent with the r-sweep multiplicative finding); the corridor closed from the CEILING side only. Wall-1-as- BBP is real where noise IS additive iid: ANALOG readout. Design equation for T64: beta*_l = a_HW(sqrt m + sqrt n)/sigma1(g_l), a_HW from ENOB 8.51b / 32uV window numbers -> quantitative minimum-nudge spec ("big nudge free" upgraded from empirical to closed form). Muon footnote: sub-threshold beta + msign orthogonalization = full-magnitude noise injection (why the hard-bottom burn was toxic: gn pinned ~0.3 by normalization, cos -0.76). ### RESULT 47 (2026-07-18): THE GROWING QUANTITY NAMED — one relaxation mode living in blocks 8-11 (90% of residual mass, BOTH lineages); its gain crosses 1 at ~205-210k in plain, flat in cent. Weight-space and free-state audits both exonerated. Three-probe chain: (1) specaudit (weights): ALL spectral scales flat-to-falling; survivor cent carries LARGER vpath/ffn/logit maxima; qk-gain creep +3% both lineages pre-collapse -> weights exonerated. (2) statej (free-state per-block sigma(J), FD-JVP): chain-product hypothesis REFUTED — cent@230k product (3.1e16) exceeds plain@death (3.9e16 @195k) and sails; free-state norms don't discriminate. (3) rhorelax (governor meter replicated offline, K=30, fixed batch): res0 plain 0.0020->0.0031 (2.1x cent's flat 0.0015) pre-collapse, 0.0093@210k (rho_tail 1.02, non-contracting), 0.44/divergent@230k; cent converges ~1e-6 at every step through 230k. MODE PROFILE: 0.15/0.21/0.26/0.28 on blocks 8-11 — identical in both lineages at all steps. Same structure as RESULT 44's gap localization: finite mode gain = transmission error (CE gap); gain->1 = corridor death. One object, two symptoms. State-side discriminator: b8 max|logit| runs away in plain (70->94 by 195k) vs bounded <=80 in cent (statej data). TARGETED 1.0x CANDIDATE: logit softcap on blocks 8-11 only (~50*tanh(l/50); healthy p99.9 ~30-50 -> cap inactive except on the runaway tail). Test = one s190000 resume arm, 4h. ### RESULT 46 (2026-07-18): ENDGAME-42 TABLE — BP is lr-robust (reference stands); plain@3e3 honest 3-seed gap +0.042; ride's edge is MID-phase only. Late probes (cent s150000 -> 160k): fix_late 3.4448 vs rv2as_late 3.4464 — tie/fade (mid-stage ride won by 0.004; late it buys nothing). Riding harvests only where floor << ceiling; late-phase floor 3e-3 is already at the ceiling. Recipe note: ride = early/mid tool. 42M plain@3e3 seeds: 1.2371(s1)/1.2731/1.2572, mean 1.2558 +-0.018 — s1 was the LUCKY draw (mirror of BP where s1 was the weak one; dip-statistics bites both ways). The old s1-vs-s1 "plain ~ cent" read was two outliers shaking hands. BP lr sweep (s1): 7e-4 1.2119 / 1e-3 1.2311 / 1.4e-3 1.2136 -> flat plateau across 2x lr; no better setting exists in range; BP reference = 3-seed mean 1.2141 (fair, user's best-setting question answered at 42M). HONEST 42M LADDER (best val, n-seed means): BP 1.2141 (n=3) | cent 1.2334 (n=1, s2/s3 QUEUED on GPU0) | ride 1.2449 (n=3) | plain3e3 1.2558 (n=3). Gap ordering cent < ride < plain matches the 72M stability ladder — same estimator hierarchy at both scales. ### RESULT 45 (2026-07-18): NO-FLOOR VERDICT — the corridor TRULY CLOSES for the plain lineage (ceiling falls 260x with no shelf); fail-soft property of pure tracking CONFIRMED. Plain judged DEAD as a full-schedule 72M recipe; centered-rescue 2x2 cell in flight. fw72m_plain_nf (resume plain s190000, cap_floor 0, exact mirror): COMPLETED 234k, best 3.3838 — set at ~195k (matches crown-3's pre-death 3.3836) and NEVER improved after: the cap crushed beta 3e-3 -> 9.3e-4@200k -> 2e-4@208k -> 1e-4@220k -> 1.1e-5@232k, monotone with NO SHELF. The ceiling is really gone (not a hard-bottom artifact); with res>0.02 persisting at beta=1e-5 the meter is reading FREE-PHASE contraction loss (state property, beta-independent) — consistent with the 07-14 wall-2 regression at 72M/FW ~195k. FAIL-SOFT CONFIRMED: gn <=0.4 / drift 0.012 / zero skips / no spike / weights un-poisoned (val drifts 3.5 -> 4.1 by SNR-dead noise-walk) — vs crown-3's hard-bottom fail-BURN (gn 29.5, cos -0.76, drift 0.39, spike). cap_floor 0 is the strictly safer governor policy; the hard bottom bought crown-3 its extra 0.011 of best (floor-fight learning 196-204k) at the price of late-run poisoning. DECISION (user pre-committed): plain cannot hold the 72M/FW full schedule -> centered. Cheapest form = the decisive 2x2 cell fw72m_plain_c190 IN FLIGHT (centered est on the SAME plain s190000, cap_floor 0, otherwise identical): centered holds the ceiling -> HYBRID recipe (plain bulk + centered tail, ~1.13x cost) is live and 2x2 completes as state-rescuable-by-estimator; centered also collapses -> plain bulk bakes the disease in before 190k -> full-run centered only (crown stays fw72m_cent 3.3318 at 1.67x). ### RESULT 44 (2026-07-18): GAP-SOURCE LOCALIZATION — the residual EP gap lives in the TOP-half blocks, split ~evenly attn/ffn; bottom-half EP grads are loss-equivalent to BP. bpmix 5-arm, all from cent s200000 through the 215k endgame frame (shared frame; internal A/B only, absolutes not comparable to crown lineage). BP reference (allbp) 3.4020; all-EP 3.4245 -> window gap 0.0225. Swapping module groups to TRUE BP grads: mixbot (blocks 0-5) 3.4244 -> recovers ~0% of the gap mixtop (blocks 6-11) 3.4071 -> recovers 77% mixattn (attn, all) 3.4152 -> recovers 41% mixffn (ffn, all) 3.4161 -> recovers 37% (41+37 ~= 77+0: additive decomposition holds) Reading: bottom-half EP fidelity is a SOLVED problem at this scale (exact-BP substitution buys nothing); the entire residual gap concentrates in blocks 6-11, attn and ffn about equally. Tension with RESULT 25 (transmission bias compounds toward the bottom): bottom grads may be worse in cosine yet loss-irrelevant in this window — update magnitude/role concentrates near the head. HW consequence: fidelity budget (quant/noise/beta) should be weighted toward the top half. ### RESULT 43 (2026-07-18): CROWN-3 CAUSE OF DEATH = CEILING CROSSED, THEN THE CAP'S HARD BOTTOM HELD BETA ABOVE THE COLLAPSED CEILING FOR 18k STEPS (disguised wall-2, user's read). Estimator-toxicity story DEMOTED; endgame A/B is a TIE. Endgame A/B (from cent s200000, 15k steps): end_plain 3.4245 vs end_cent 3.4244 — dead tie, both healthy (gn 0.147 / drift 0.009 / 0 skips; plain: one -12% cap graze @213.9k, instant recovery). Estimator CE-equivalence now 3-way replicated (42M tail, 72M mid-tail, 72M endgame); single-sided is NOT locally toxic on a healthy state. "不是plain不行" (user) confirmed. Death timeline (simple-bounds-first per user): first cap bite @196k with ALL health metrics normal — leads every symptom by 8k+ steps (gn 204k, cos-negative 217k, drift 216k, skips 225k). Cap wrestles floor 196-204k (best 3.3728@203.7k set DURING the fight), then pins at the HARD BOTTOM 0.05*beta = 1.5e-4 from ~216k with rho^>0.9 & res>0.02 continuously (v2 absolute-gate = bites genuine). rho^ is a per-sweep residual contraction ratio — a WEIGHT-STATE property beta cannot restore — so the governor had no lever and sat pinned ABOVE the true ceiling. Zero r-skips until 225k: the relax never overtly diverged; death by sustained sub-ceiling-less operation + SNR starvation, not explosion. Controls: cent late-phase (150k-234k): ONE graze (221.8k). end_plain (plain est, cent state): one graze in 15k. => ceiling collapse is a plain-LINEAGE-STATE property. sigma nearly equal at 200k (467.7 vs 463.0) => the lineage difference lives in block contraction ||J||, not W_out. RESULT 42's "bias->sharpening positive feedback" narrative DEMOTED (user's evidentiary bar: no direct literature for invisible-accumulation stories). Canonical citation that DOES cover one-sided-fails/symmetric-rescues: Laborieux et al. 2021 (O(beta) estimator bias, deep-net failure, +-beta symmetric fix = our centered). ACTION: --cap_floor flag added (default 0.05 = legacy; 0 = pure ceiling-tracking, three hard bottoms released). fw72m_plain_nf IN FLIGHT: resume plain s190000 (last healthy ckpt), cap_floor 0, exact crown-3 mirror, replay the 190k->234k death window. Survive => hard bottom was the killer, plain viable at 1.0x (crown-3b = plain+no-floor). Beta free-fall + skip-storm stall (weights preserved, fail-stop) => corridor truly closed => centered is the fix. R42's est_late-flip candidate is MOOT either way (endgame estimator doesn't matter — tie). ### RESULT 42 (2026-07-18): CROWN-3 VOIDED (user call) — slow-burn endgame instability the guards accepted; the SINGLE-SIDED estimator emerges as the common factor in both 72M deaths. User caught what skip-counting missed: gn rising / cos falling / CE rising from ~200k, spike at 231.7k, partial recovery — RESULT 41's "zero guard events = healthy" reading RETRACTED. Forensics: bcap DID fire (beta 3e-3 -> 1.4e-3 @200k -> cap floor 1.5e-4 by 215k) but starvation cannot cure a state disease (ride-deadlock replay, this time with NO governor overshoot). The poison route: drift 0.25-0.32 updates ACCEPTED (hard threshold 0.5; only 2 skips) with garbage quality (cos as low as -0.76) at starved beta. sig peaked 483. PATTERN: original fw72m died 195k@1e-3; plain died ~200k@3e-3; cent sailed 234k@3e-3. Common factor across deaths = single-sided estimator late-phase, NOT the beta value. HYPOTHESIS (promoted): the single-sided O(beta) bias at large sigma*||J|| forms a positive feedback (bias -> sharpening -> larger bias); centered's symmetric read cancels that term = STABILITY ARMOR, not just CE polish. The 0.041 "discrepancy" and this death are one phenomenon. MISSED PROTECTION (my scheduling error): beta_sync and drift_adapt existed on the shelf (probe-validated) but were not given to the crown — "conservative" in the wrong dimension. drift-EMA was 0.006 -> a 12x wall (0.072) would have rejected every poison accept. ADJUDICATION IN FLIGHT: end_plain vs end_cent from clean s200000 (does single-sided sicken in 15k endgame steps?) + bpmix 5-arm localization. CROWN-3b deferred until they land; candidate recipe: plain bulk + sync/drift-wall armor + CENTERED ENDGAME (est_late, meaning flipped). Ledger: the 72M crown remains fw72m_cent 3.3318 (gap 0.043); a 1.0x-cost crown DOES NOT EXIST yet. ### RESULT 41 (2026-07-18): CROWN-3 SEALED — fw72m_plain 3.3728 (gap 0.084 at 1.0x cost); a 0.041 plain-vs-cent discrepancy opens the ENDGAME question. fw72m_plain (from scratch, plain estimator, fixed bf_late 3e-3@20k, bcap defensive): 234k steps, best 3.3728, zero guard events. Ledger: BP twin 3.2884 | cent 3.3318 (gap 0.043) | plain 3.3728 (gap 0.084) | original blown 3.7117. - Still 5x better than the blown-schedule number at 1.0x estimator cost; but 0.041 WORSE than cent — which contradicts the mid-tail A/B (plain==cent at s175-185k, RESULT 33). - Two candidate explanations: (a) 72M single-seed trajectory noise (no seed distributions at this scale); (b) centered has a REAL edge specifically in the cosine ENDGAME (cent's best came at 203.7k; the A/B never tested that segment; bias/beta ratio grows as LR->0?). - ADJUDICATOR LAUNCHED: endgame A/B from cent-lineage s200000 -> 215000 (15k steps through the LR endgame), cent+mirror vs plain, GPU1/3 parallel. If cent wins there: est_late flips its meaning — plain early, CENTERED endgame (the exact opposite of the original est_late design); recipe = plain bulk + centered finish at ~1.05x cost. If equal: (a) seed noise, cent's full-run edge unexplained, more seeds needed at 72M. ### RESULT 40 (2026-07-17): RIDE-V2 PROBE VERDICT — the floor-jump bug WAS the whole disease; sync-accept adopted as default insurance. 4 arms, 72M mid-stage (cent s50000 -> 60000, 10k steps): fix 3.5715 | rv2a 3.5688 (beta->1.24e-2, 0 skips) | rv2ad 3.5688 | rv2as 3.5674 (beta 1.06e-2, 0 skips). - Bug-fixed naked ride surfs SAFELY at 72M mid (4x beta, zero guards) and edges the fixed control. - drift-wall (d): zero interventions in a healthy climb — costless insurance. - sync-accept (user's synchronous acceptance): best of four, more conservative beta — adopted into the default governor stack (a + sync; d optional). - Honest bounds: inter-ride deltas (0.001-0.004) are single-seed probe noise; the firm claims are safety + >=control. Late-stage pair (s150000) queued to complete the picture. ENDGAME-42 QUEUE launched on GPU0: fix_late + rv2as_late -> plain3e3 seeds 2/3 (EP-recipe distribution, RESULT 39) -> BP lr mini-sweep 0.7e-3/1.4e-3 (fairness, RESULT 36). ### RESULT 39 (2026-07-17): SEED CAMPAIGN VERDICT — ride_s1 was a lucky draw, AND RESULT 26's "statistical zero-gap" was anchored on BP's weakest seed. CLAIM CORRECTED. | run | s1 | s2 | s3 | mean +- sd | |---|---|---|---|---| | stage1b_ride (EP governor) | 1.2016 | 1.2738 | 1.2592 | 1.2449 +- 0.038 | | stage1b_bp (twin) | 1.2311 | 1.2050 | 1.2063 | 1.2141 +- 0.015 | (protocol note: bp_s2/s3 ran --amp like all EP arms; bp_s1 was the original fp32 run.) 1. ride's 1.2016 = the deepest dip of a HIGH-VARIANCE family (range 0.07) — the single best number of all six runs, but by distribution BP leads (means 0.031 apart, ~2 sd). 2. ERRATUM to RESULT 26/35 zero-gap phrasing: bp_s1 (1.2311) is the WEAKEST of BP's three seeds. Against the BP 3-seed mean (1.2141): cent-full gap +0.019, plain3e3 +0.023 — SMALL gap, not statistical zero. The honest C512 statement until EP-recipe seeds land: "EP within ~0.02 of the BP seed-mean; family distributions overlap at the tails." 3. NEEDED for the final number: EP-recipe seeds (cent/plain3e3 currently n=1) — queued after the ridev2 probes on GPU0; deck S9's zero-gap line to be softened in the next revision. 4. The BP-fairness lr mini-sweep remains queued behind that (RESULT 36 protocol). ### RESULT 38 (2026-07-17): 2D UPDATE-FIELD ANALYSIS (professor's suggestion) — the finite-beta field IS non-conservative, but only where training doesn't live; EP and BP reach DIFFERENT, equally deep basins. Method: fieldviz.py — full-model PCA plane over 27 snapshots of 5 runs (49% trajectory variance); 13x13 grid; loss contours + three update fields (BP, EP@3e-4, EP@1e-2; 2-batch averaged; block-subspace projection); finite-difference curl; ride<->BP linear interpolation. | field | curl RMS | |---|---| | BP (conservative reference / noise floor) | 4.38e-6 | | EP @ 3e-4 (operating beta) | 4.41e-6 (= floor) | | EP @ 1e-2 | 1.37e-5 (3.1x floor) | Findings: 1. NON-CONSERVATIVITY IS REAL AND LOCALIZED: at beta=1e-2 a coherent positive-curl blob appears in the steep high-loss region (large states/gradients -> large bias field c(theta)); BP and EP@3e-4 show pure noise. NEAR THE BASINS — where training actually operates — the EP field stays conservative-within-noise even at 10x operating beta. Same-batch BP control shows no blob => EP-specific, not data noise. 2. Two-level echo for the dynamics paper: parameter-space curl concentrates exactly where the loop-gain hazard lives (big sigma*||J||) — the state-space non-conservativity story reappears one level up, scaled by beta. Training under finite-beta EP = a mildly non-gradient flow that is gradient-like precisely in the region the governor keeps it in. 3. DIFFERENT BASINS, EQUAL DEPTH: ride(s50k) 1.4325 vs BP(s55k) 1.4156 with an 8.25 barrier (~random-level) between them — same seed/init, methods branch at step 0 and land in linearly disconnected but equally good minima. The EP family (plain/ride/cent) clusters in one plane region, BP in another. CONTROL PENDING: BP-s1 <-> BP-s2 interpolation (does BP-vs-BP also barrier? decides whether different-basin is EP-specific or generic) — runs when the seed chain delivers bp_s2. Cheap follow-ups queued: EP<->EP interp (expected flat); curl-vs-beta scaling curve at a near-basin point and a steep point (the quantitative bridge figure for aep-dynamics). Caveats: 2D slice (49% variance); curl from 2-batch-averaged fields on a 13x13 grid; barrier statement is linear-path only (no permutation alignment attempted). ### RESULT 37 (2026-07-17): RIDE FAILS AT 72M — weight poisoning via permissive guards during beta surfing; crown-3 relaunched as PLAIN from the clean ckpt. Timeline (fw72m_ride): healthy surf to ~48k; at 72M the governor never found the 42M-style smooth hover — beta banged between 1.5e-4 and 1.85e-2 (rho-hat is noisy/laggy at this scale). ~54k: first skip jolt (4->17). By 60k beta was PINNED at the cap floor (1.5e-4) yet nearly every step drift-rejected at K=8 — relaxation diverging at 20x smaller beta than had been stable for 50k steps => the WEIGHT STATE itself was sharpened/poisoned, not a beta-level wall. Mechanism: at big beta, displacement (and any semi-diverged garbage) scales with beta, but the drift-accept threshold (0.5) is beta-INDEPENDENT — near-threshold accepts during 1e-2-scale surfing carry beta-scaled damage into the weights. 42M never showed this because its window is so wide the excursions stayed benign. CE never recovered (val 3.77 -> 4.0-4.2, best frozen 3.5797); training deadlocked (governor cannot cure a state disease by lowering beta). INTERVENTION: killed at ~75k; relaunched as fw72m_plain from the last clean ckpt (s50000, skips=3 era): plain estimator, FIXED bf_late 3e-3@0, bcap 0.9 defensive-only (ride=1.0). Resume health: val 3.7354, zero skips, cos 0.9965. ETA ~15h. LESSONS (ride-v2 design, to be tuned on cheap probes, NEVER on crown runs): (a) beta-scaled accept threshold (drift_max ~ f(beta)) or update-norm clip during high-beta; (b) climb hysteresis: after any wall contact, cooldown + re-climb at a fraction of the last stable beta (no immediate re-surf); (c) rho-hat smoothing (EMA) before governor decisions — the raw per-step meter is too noisy at 72M; (d) hard beta_ride headroom set from the last-known stable beta, not a fixed 30x. Status: ride stays VALIDATED at 42M (RESULT 35); at 72M it is a failed-first-attempt with a diagnosed mechanism — an honest boundary datum for the governor line, not a retraction of it. ADDENDUM (same day, deeper forensics — CORRECTS the mechanism story and the intervention): - THE FIRST WOUND WAS A CONTROLLER DESIGN BUG, not gradual poisoning: during the calm 3e-4 era (steps 0-20k) the ride cap climbed to saturation (30x). When bf_late jumped the floor to 3e-3 at step 20000, beta = floor x cap = 0.09 ON THE FIRST HIGH-BETA STEP (log: step 20000 beta=0.0900). Everything after 20k is contaminated -> the "s50000 clean" call is RETRACTED (it was judged by skips; damage precedes symptoms). - The bad resume confirmed it: from s50000 with fixed 3e-3, bcap immediately crushed beta to the cap floor (1.5e-4) and relative drift stayed 0.02-0.04 even at that tiny beta = sharpened-state disease inherited from the lineage. - INTERVENTION v2: fw72m_plain relaunched FROM SCRATCH (s15000-salvage saves only ~1.5h; a clean lineage is worth more). Recipe = cent-proven schedule with plain estimator: bf_late 3e-3@20k, bcap 0.9 defensive-only. The stale wandb run was deleted; bad-resume log kept as fw72m_plain_badresume.log. - ride-v2 lesson (e), the binding one: the cap multiplier composes with FLOOR JUMPS — cap must be defined relative to effective beta (or reset/clamped at any floor change), never allowed to pre-charge against a low floor. ### RESULT 36 (2026-07-17): PRE-REGISTRATION — ride seed campaign + fw72m_ride crown-3 + the BP-fairness note. User call: "多跑几条 ride, 在 72M 跑 ride — 说不定真比 BP 好, 因为我们没扫 BP 的最佳 setting." Launched: - fw72m_ride (GPU1+3, DDP, from scratch, seed 1): PLAIN estimator + ride governor (bf_late 3e-3@20k, bcap 0.9, ride 30, up 1.01), 234k steps. Predictions: (a) governor surfs, few guard events; (b) best <= 3.33 (plain==centered + governor edge); (c) wall-clock ~20h (plain speed 3.3 it/s vs cent 1.96). This is the 1.0x-cost crown. - stage1b_ride_s2/_s3 + stage1b_bp_s2/_s3 (GPU0 chain): seed distributions for BOTH sides of the comparison. Metrics: best AND tail-median AND (new) fixed large-eval on FINAL weights — both trainers now save the final-step ckpt (protocol upgrade, this commit). - FAIRNESS NOTE (user's point cuts both ways): the BP twin has never been tuned (mirrored EP lr/ schedule). Any EP-vs-BP ordering claim requires a BP lr/schedule mini-sweep — queued as the next GPU0 item after the seed chain. Until then, "EP matches BP" is claimable; "EP beats BP" is not, regardless of seed outcomes. - FAIRNESS PROTOCOL AT SCALE (settled 2026-07-17, user Q "1B+ can't sweep — what's accepted?"): (1) BP anchor at every scale = the PUBLISHED community recipe for the architecture (OLMo2's own tables) + citation — stronger than any self-sweep; (2) proxy-scale sweeps (42M, 300M) for BOTH methods validate the anchor AND produce lr-sensitivity curves that bound the "BP could be better" risk quantitatively; (3) at 1B+ both sides run transferred settings, zero on-site tuning — fairness = symmetric protocol, not asymmetric best-effort; (4) optional strongest form: muP/muTransfer from the 300M rung (decide at 300M); (5) selling point: EP's extra knob (beta) is governor-self-tuned — "one self-tuning knob at scale" vs BP's lr tables. ### RESULT 35 (2026-07-16): FULL-EPOCH TRIO SEALED — 1.0x-cost zero-gap confirmed; the ride governor wins on BOTH metrics; dip-statistics caveat formalized. | run (C512 full epoch) | best val | tail median (last 6k) | |---|---|---| | BP twin | 1.2311 | 1.2823 | | **stage1b_ride (governor)** | **1.2016** | **1.3034** | | stage1b_cent (centered) | 1.2334 | 1.3241 | | stage1b_plain3e3 (plain, 1.0x) | 1.2371 | 1.3259 | | stage1b_est15 (switch rule) | 1.2374 | — | - plain3e3 1.2371 = the zero-gap-class result at 1.0x estimator cost: RECIPE SIMPLIFICATION CONFIRMED (plain single-sided + big-beta schedule; centered's full-epoch edge 0.0037 < band). - ride beats every EP arm on BOTH best AND median (median edge over cent: 0.021) — the governor's dynamic schedule (big-beta bulk, auto-taper endgame, beta down to 1.7e-4 at the end; 20 gn-guard skips in 58.8k) genuinely improves the steady state. - HONESTY RULE (formalized): ride's best 1.2016 < BP's 1.2311 is VARIANCE-ASSISTED (beta surfing raises val variance -> deeper dip harvest under the best-of-noisy-val convention). By tail median BP still leads ride by 0.021. Do NOT claim EP tax +0.0223, down from +0.0348 at beta 1e-3. Direction confirmed (a wall-1 epsilon component exists), magnitude ~1/3 reduction for 10x beta — a second, non-beta-scaling component remains (weight-grid roughness is a function-space perturbation, not just read noise). Spec consequence: the 8-bit operating point stands; do NOT plan on beta rescuing 6-bit devices; ENOB acceptance bar stays ~7. ### RESULT 33 (2026-07-16): TRIPLE VERDICT — centered CE-neutral at 72M too; CE-vs-beta monotone through 1e-2; the RIDE GOVERNOR beats every static beta. 1. 72M A/B (resume fw72m_cent s175000 -> 185000, single-GPU B24, identical schedule): plain 3.4288 vs centered 3.4289 — IDENTICAL. The 42M decoupling replicates at crown scale: the recipe does not need centered. plain + governor = 1.0x cost at 72M. 2. Static beta dose-response (plain, 42M tail): 2e-3 1.2611 / 3e-3 1.2590 / 5e-3 1.2578 / 1e-2 1.2571; centered@1e-2 1.2569 (neutral again). CE STILL IMPROVING at 1e-2 — no bias bite anywhere in the explored range; the 42M window extends beyond 1e-2. 3. arm_ride30 (two-sided governor, 30x headroom): best 1.2562 — BEATS every static arm. The governor surfed dynamically: 3e-3 -> ~6e-2 excursions mid-tail -> taper to ~7e-4 at the end, ZERO guard events, cos 0.992 held. It discovered a dynamic beta schedule no hand-tuning found (high-beta bulk + late taper aligned with the cosine-LR endgame). - RECIPE CONSEQUENCE: estimator upgrades (centered/centfast/centmirror) demote to cos-telemetry tools and insurance; the working recipe trends to PLAIN single-sided + ride governor = 1.0x estimator cost. 8B ledger returns to the 3.2x base. - Full-epoch validations: stage1b_est15 SEALED 1.2374 (+0.0063 vs BP — the beta-coupled switch rule lands at the seed-band edge; consistent with the big-beta cluster 1.233-1.238); stage1b_plain3e3 (46k+) and stage1b_ride (GPU3) in flight. - beta-buyback arm LAUNCHED (qb6_b1e2: 6-bit compute quant at beta 1e-2; RESULT 30's mechanism prediction — if the +0.035 tax shrinks, quantization noise is confirmed as a wall-1 epsilon term purchasable with beta; control = bab_p1e2 1.2571). ### RESULT 32 (2026-07-16): CROWN RESEALED — fw72m_cent 3.3318 vs BP 3.2884: honest gap 0.043, 10x below the blown-schedule number. fw72m_cent complete: 234k steps / 1.44B FineWeb tokens, from scratch, fully BP-free, DDP 2xA6000. | | best val CE | note | |---|---|---| | BP twin | 3.2884 | @195.6k | | **fw72m_cent** | **3.3318** | @203.7k — still improving in the final quarter | | original fw72m | 3.7117 | @104.8k, then 90k-step starve + 195k blow | - Registered predictions (RESULT 24): (a) no blow — ZERO skipped steps in 234k ✓; (b) beats 3.7117 ✓ (by 0.38); (c) gap 0.25-0.35 — BEAT 8x (0.043). - beta history: 3e-4 ramp (first 20k) -> 3e-3 for ~214k steps; the ride never ended — bcap grazed twice (2.76e-3 / 2.65e-3 single lines) and recovered instantly. cent's trajectory NEVER met the wall the original hit at 1e-3@195k: the ceiling is trajectory-specific, and the healthy window-riding path kept its margin. (Honest note: this also means the 72M ceiling-descent curve from the original run does NOT transfer across recipes.) - Best kept improving to 203.7k (no 105k-style freeze): the SNR account balanced. - Samples (fw72m_cent_s205000): FineWeb register, "can-read" PASS (runs/fw72m_cent_samples.txt). - HEADLINE: largest BP-free transformer LM (72.11M x 1.44B tokens), gap to matched BP twin 0.043 (~1.3% relative), one governor knob, standard architecture, standard inference. ### RESULT 31 (2026-07-16): DECOUPLING VERDICT — the tail win is ALL beta; centered is CE-neutral at 42M tail. RESULT 22's mechanism reading corrected. The missing arms7 control (user-demanded): plain single-sided @ beta 3e-3 flat tail. | arm (45k->55k) | best | note | |---|---|---| | ctl (plain @1e-3) | 1.2678 | | | **arm_plain_f3e3 (plain @3e-3)** | **1.2590** | = centered to the 4th decimal | | arm_cent_f3e3 (centered @3e-3) | 1.2591 | | | arm_centmirror (mirror @3e-3) | 1.2590 | downstream parity of the mirror trick CONFIRMED | beta share of the tail effect: 101%. Readings: - RESULT 22's causal story ("O(beta^2) bias lets beta ride high") is REFUTED at this scale: plain rides 3e-3 equally well. The O(beta) secant bias at 3e-3 is measurable in cos but COSTLESS in CE — third independent confirmation that direction-space error does not price CE; only SNR does. - The zero-gap driver in stage1b_cent (1.2334) is therefore suspect of being pure-beta too: **stage1b_plain3e3 full epoch queued** (post-crown, GPU1/3). If it lands ~1.233, the C512 zero-gap recipe simplifies to plain + big-beta = 1.0x cost, and centered/centfast/centmirror demote to cos-telemetry tools pending a scale where bias binds. - 72M check queued (post-crown A/B): resume fw72m_cent s175000, 10k steps, centered-continue vs plain-switch at identical schedule — decides whether the crown recipe needs centered at all. - Auto-triggered beta ablation RUNNING (share>=0.5 rule): plain @ {2e-3, 5e-3, 1e-2} + centered @1e-2 — maps CE-vs-beta and where plain's O(beta) bias finally bites; 1e-2 arms double as 42M wall-2 probes (guard telemetry free). ### RESULT 30 (2026-07-16): QUANTIZATION Delta-vs-Delta MATRIX — equal at the operating point, EP-SPECIFIC tax below it. BP mirror suite complete (bp_qctl baseline 1.2171 — resumed-tail+amp beats the original fp32 twin, which is exactly why in-family baselines were required; user's methodology point vindicated). | injection | Delta_EP (vs 1.2678) | Delta_BP (vs 1.2171) | EP/BP ratio | |---|---|---|---| | qup 10-bit | +0.012 | +0.003 | ~4x | | qup 8-bit | +0.042 | +0.015 | ~3x | | qup 6-bit | +0.092 | +0.068 | ~1.4x | | qcomp 8-bit | **+0.002** | **+0.000** | **both ZERO** | | qcomp 6-bit | +0.035 | +0.004 | ~8x | | qcomp 4-bit | +0.140 | +0.061 | ~2.3x | Readings: 1. AT THE T64 OPERATING POINT (8-bit compute + digital master): quantization is free for BOTH. The BP QAT toolbox transfers AT this point; T64 green light unconditional on this axis. 2. BELOW it, the tax is EP-SPECIFIC (1.4-8x faster degradation): quantization roughness enters the finite-beta measurement chain as an extra epsilon in the wall-1 SNR term (the estimator protects an O(beta)-scale signal; BP has no such small signal). A blanket "EP inherits BP quantization behavior" claim is REFUTED below 8 bits — publishable Stage-0 finding. 3. MECHANISM PREDICTION (designed, pending GPU): bigger beta should buy back quantization tolerance (noise-to-displacement ratio ~ 1/beta) — one arm: qcomp6 + beta 1e-2 tail vs qcomp6@3e-3 (+0.035). If confirmed, the window story absorbs quantization as another epsilon term. 4. Procurement consequence: acceptance ENOB bar STAYS ~7 (do NOT relax to 6 on BP intuition — EP@6-bit compute is +0.035 real). ### RESULT 29 (2026-07-16): qcomp VERDICT — 8-bit COMPUTE quantization is FREE (T64 green light); centmirror ships (1.39x). qcomp arms (compute on DAC-grid weights, fp32 master = word-streaming / shadow accumulation), same protocol, vs ctl 1.2678: | bits | qcomp (T64 scenario) | qup (naked resident) | |---|---|---| | 8 | **1.2696 (+0.0018 = ZERO)** | 1.3096 (+0.042) | | 6 | 1.3026 (+0.035) | 1.3598 (+0.092) | | 4 | 1.4077 (+0.140) | — | - 8-bit DACs + digital master = tax-free at 42M: the Y3/pc AD7528 line and the T64 word-streaming architecture pass gate #1b. 6-bit compute has a real but moderate tax; 4-bit heavy. - BP mirror suite RUNNING (bp_qctl 49.5k, then qup10/8/6 + qcomp8/6/4): the Delta-vs-Delta verdict (EP-specific or generic) lands tonight; RESULT 28's naked-cell reading stays provisional till then. - ESTIMATOR COST ENGINEERING sealed: --centmirror (the -beta pass initialized as the MIRROR of the +beta solution at the shared free anchor + 1 polish sweep; second free pass and K-1 sweeps deleted). Gradtest cos 1.000000000 vs sequential centered (relerr 2.2e-5). Quiet bench B12/amp: plain 4.389 it/s | sequential centered 2.547 (1.72x) | centfast 2.684 (1.64x) | **centmirror 3.169 (1.39x)**. With est_late@80%: amortized ~1.08x — centered is now essentially free at scale (8B ledger: 3.2x -> ~3.45x vs BP). ### RESULT 28 (2026-07-16): STAGE-0 HW GATE #1 — naked analog-resident updates need >10 bits; T64's word-streaming scenario measured next. Protocol: arms7 (resume s45000 -> 55000, ctl 1.2678). --qup_bits = weights snapped to an ABSOLUTE per-tensor grid after every update, stochastic rounding (= analog-resident cells, NO shadow). | levels | best CE | tax vs ctl | |---|---|---| | fp32 (ctl) | 1.2678 | — | | 10-bit | 1.2795 | +0.012 | | 8-bit | 1.3096 | +0.042 | | 6-bit | 1.3598 | +0.092 | Monotone dose-response; clean (zero guards). Readings: - NAKED resident-cell training (updates quantized at write, no digital shadow) needs >=10-12 bits — this is the measured version of why in-memory analog UPDATE machines die on write resolution. - T64 is NOT this scenario: word-streaming keeps the fp32 master in Zynq DDR; the 8-bit DAC is a COMPUTE element. The correct T64 gate = --qcomp_bits (compute on grid-snapped weights, fp32 master gets updates; mathematically identical to resident-cells + 24b shadow accumulator). qcomp8/6/4 arms running (gate #1b). - Procurement impact: the Y3/pc AD7528 line is unaffected either way (its role is compute); the acceptance ENOB bar keys on the qcomp verdict. - CAVEAT (user, 2026-07-16): the fp32-EP control conflates "quantization tax" with "EP-specific quantization tax". BP MIRROR SUITE chained (bp_qctl + qup10/8/6 + qcomp8/6/4 on the BP twin, same resume protocol). Decision metric = Delta_EP(bits) vs Delta_BP(bits): equal deltas => generic quantized-training problem => the BP QAT/shadow toolbox transfers to EP unchanged. Reading-one above ("naked cells need >=10-12b") is PROVISIONAL until the BP control lands — naked quantized writes likely hurt BP comparably. ### RESULT 27 (2026-07-15): K-SATURATION ACROSS TRAINING (Alexi's challenge answered with data). Challenge (Alexi Gladstone): K=3 nudge sweeps seems very few; as training roughens the landscape, more sweeps may be needed absent a convexity argument. Probe: probe_blockcos.py on fw72m_cent ckpts s5000/s50000/s100000, K in {1,2,3,8}, beta 3e-3, 4 batches. | ckpt | K=1 | K=2 | K=3 | K=8 | |---|---|---|---|---| | s5000 | 1.0000 | 1.0000 | 1.0000 | 1.0000 | | s50000 | 1.0000 | 0.9998 | 0.9997 | 0.9997 | | s100000 | 1.0000 | 0.9994 | 0.9995 | 0.9995 | Verdicts: (a) fixed point reached by K~2-3 at EVERY stage; K3=K8 to 4 decimals — no K-starvation trend over 100k steps. (b) What grows with training is the FIXED POINT's own O(beta) transmission bias (1.0000 -> 0.9995) — the window story, not truncation; K cannot treat it, beta/centered can. (c) ANCHOR stays exactly 1.0000 at all stages: RESULT 25 decomposition holds at 72M throughout. (d) K=1 constant 1.0000 = the frozen-state BP degeneracy (digital shortcut, not physics). (e) Cost note: K=2 is already ~at the fixed point -> potential ~25% nudged-phase saving (validation item, recipe unchanged for now). The correct response to a roughening landscape is governed beta (the certificate is the live rho-hat spectral meter), not more sweeps — at the edge, extra sweeps amplify (w2_adapt). ### ERRATUM to RESULT 21 (found 2026-07-15): the original fw72m's best val 3.7117 was first hit at step 104,800 (log line 1055), NOT at s185000 (that was the checkpoint used for sampling). Best-val progression: 3.8769@30.9k -> 3.8328@61.3k -> 3.8004@75.3k -> 3.7789@83.3k -> 3.7117@104.8k -> flat for the remaining ~90k healthy steps until the 195k blow. The original schedule was therefore SNR-starved from ~105k on (wall-1), before it was killed by wall-2. Strengthens the window story; blowup figure annotation corrected. (fw72m_cent at the same 104.8k mark: best 3.4523 and still descending.) ### RESULT 26 (2026-07-15): C512 ZERO-GAP SEALED — full-epoch centered erases the 0.050 gap. stage1b_cent (C512 42.75M, --est centered whole epoch, bf_late 3e-3@15k, amp, seed 1, GPU0) completed all 58800 steps: best val CE 1.2334. | run | best | gap vs BP | |---|---|---| | BP twin (stage1b_bp_muon) | 1.2311 | — | | stage1b_cent (THIS) | 1.2334 | +0.0023 << seed band ±0.006 -> STATISTICAL ZERO | | cent tail-only arm (45-55k) | 1.2591 | +0.028 | | original EP (plain, decay floor) | 1.2808 | +0.050 | Decision tree (RESULT 23 pre-reg): strongest branch hit — the whole 0.050 was a recipe/noise account. ZERO-GAP now holds at C128, C192, and C512 (C256/C384 tiers pending seeds). The ladder's "fit-depth tax" is REMOVABLE by the window recipe at this scale. Note: in-run gate_cos still declines late (wandb strip) while CE stays matched — third confirmation that direction-cosine is not the CE-relevant metric (C128 lesson). Cost: centered whole-epoch 1.72x. est_late would buy most of it back; the switch-point rule is the remaining tuning question. fw72m_cent (35%, best 3.5560, leading by 0.24) tests the same claim at 72M/FineWeb. ### EXTERNAL NOTES (2026-07-15, user-relayed from a sibling Hopfield-EP library project): 1. Hard-sigmoid Hopfield needs state CLAMPING: units with drive>1 equilibrate exactly ON the rho breakpoint (sliding-mode equilibrium); unclamped discrete Euler chatters around it with O(1) amplitude forever. -> Taxonomy exhibit #3 for the dynamics paper (solver-artifact family: Hopf / loop-gain / sliding-mode chatter). -> OPERATIONAL RULE for Stage-0 fault injection: when injecting rail/saturation, CLAMP states — otherwise you measure integrator chatter, not device physics. 2. Saturated units make finite-beta EP legitimately diverge from BPTT (measured 80-130 deg): dead units are invisible to BP, the nudge can revive them. NOT a bug — a hard-rho property. -> METHODS RULE: gradient-equivalence gates require smooth activations (our gates comply). -> Discussion candidate: finite-beta as exploration/repair (dead-head revival probe, low prio). -> HW note: at device rails, EP's systematic deviation points OFF the rail — likely a robustness bonus for analog training. 3. Classic EP + momentum collapses (MNIST 85% -> 48%; plain SGD required): momentum integrates the estimator's non-zero-mean bias. UNIFIES with RESULT 22's dose-response: bias/signal large -> fatal (their setting); small (governed beta + centered, mom 0.95) -> harmless; pushed to 0.99+ at low-SNR tail -> harmful again (+0.019/+0.059/+0.076). Momentum tolerance is NOT an EP property — it is purchased by bias control. ### RESULT 25 (2026-07-15): ERROR-SOURCE DECOMPOSITION — the bias lives ENTIRELY in between-block transmission; within-block reads are exactly lossless. Probe: probe_blockcos.py (GPU0, s45000 C512 ckpt, 4 batches, fp32). Exact factorization {anchor: free/nudged} x {cotangent: exact-c/EP-transmitted-d}; all four corners share one code path (selfcheck (free,c) vs BP = 1.000000). | corner | meaning | cos vs BP (beta 1e-3) | |---|---|---| | (nudged, d) = EP | the training estimator | 0.9987 | | (free, d) = TRANS-only | transmission error alone | 0.9987 (== EP, per-block profile identical) | | (nudged, c) = ANCHOR-only | local-read displacement alone | 1.0000 (1.000 x12 blocks) | Depth profile (0=bottom..11=top): 1.000 at top -> 0.997 at bottom, ~3e-4 loss per hop, monotone compounding — the between-block fingerprint (also explains why L3->L12 gates barely differ: 12 hops x 3e-4). K-sweep twist: K=1 -> EVERYTHING 1.0000 (d derived at free states = exact vjp chain = BP reproduced through block-local ops). K=3 = K=8 = 0.9987/0.9988 -> the error is the O(beta) DISPLACEMENT OF THE SELF-CONSISTENT nudged solution (saturates immediately; NOT settling truncation — K8 doesn't help; NOT local curvature — ANCHOR=1). beta 1e-3 vs 3e-3: 0.9987 vs 0.9989 (insensitive — this bias term is far below the noise term's beta-sensitivity). Readings: - ANSWER to "block内EP过程 vs block间PC连接": the bias is 100% transmission (PC-connection side); the within-block theta-read at displaced anchors is exactly free. (Digital-twin statement; on analog hardware within-block reads acquire device noise instead.) - WHY CENTERED WINS, mechanistically: +/-beta averaging symmetrizes the self-consistent displacement — kills the odd O(beta) term of exactly the one error source that exists. - The K=1 degeneracy is a digital-only shortcut (free-state vjp chain = BP-with-local-ops; a referee would rightly kill the BP-free claim for K=1). K>=2 = the physically-faithful self-consistent settle; its bias price is 0.9987 = not the binding constraint (noise is). - C128's gate cos 0.805 (ladder) is the NOISE term at small gn, not this 0.999-level bias; RESULT 23: that noise only costs CE at deep fit. ### RESULT 24 (2026-07-14): fw72m_cent LAUNCH — window-aware centered crown rerun (pre-registered). User call: "72m从头跑centered试试看". From-scratch 234k-step FineWeb rerun applying RESULT 22: original fw72m flags EXCEPT --est centered, --bf_late 3e-3 --bf_late_at 20000 (ride the arms-winner beta instead of 1e-3@60k), --beta_cap_rho 0.9 (v3 threshold: attack only near true divergence; v2's 0.7 conflated slow contraction and starved beta to 2e-5 in the c2 segment). Early phase keeps the stock ramp + 3e-4 floor (the early beta dip is a sigma-transient stabilizer, not a bias fix). Correctness gate BEFORE launch: ddp_gradtest with centered+amp — cos(DDP 2-rank, single-GPU big-batch) = 0.999987714, relerr 5.0e-3 (bf16 band; fp32 harness was 0.999999999). Chained behind the gap-scaling ladder on GPU1+3 (launcher polls GS markers). Cost: centered ~1.8x step time. Predictions: (a) no 195k-style blow — beta_t = min(3e-3, ceiling(t)) tracked by bcap instead of a fixed floor crossing the falling ceiling; (b) best val beats 3.7117 (bias down one order at matched loop gain + tail SNR up); (c) honest gap vs BP twin 3.2884 lands ~0.25-0.35 (registered guess). Failure mode to watch: bcap-0.9 rides too close to the edge -> guard storms (kretry/drift) without progress; lever = drop threshold toward 0.8, resume from last 5k ckpt. ### RESULT 23 (2026-07-14): GAP-SCALING SUITE — PRE-REGISTRATION (launched, results pending). Question: how does the EP-BP epoch gap scale with model width under the FROZEN stage1b recipe? Design: C ∈ {128, 192, 256, 384} x L12 H8 T256 B24, tinystories_bpe, full data-matched epoch (58800 steps — identical token stream for every size), EP = frozen stage1b recipe + --amp (lr 1e-3, beta 3e-3, K3, floor 3e-4, bf_late 1e-3@15k, kretry 8, olmo2, wd 0.1, muon, cosine, warmup 500, seed 1); BP twin = casc_bp_train.py mirrored flags + --amp. Anchor points already measured: C512 = 1.2808 vs 1.2311 (gap 0.050); fw72m (different data) 3.71 vs 3.29. Runs: gs_ep_c{128,192,256,384} + gs_bp_c{...}, wandb project ept-tinystories-gapscaling. PRE-REGISTERED PREDICTIONS (before any result): - H-A (user hypothesis): bigger = more robust to update noise -> gap DECREASES with C. - H-B (loop-gain): wall-2 gain grows with sigma*||J|| chains -> gap INCREASES with C at frozen beta; expected WEAK below 42M (window still wide — zero skips at C512). - H-C (detune): recipe tuned at C512 -> smallest C off-tuned -> gap inflated at C128 for uninteresting reasons. - Registered call: mild H-A trend, gap(C128) ~= 0.06-0.10 falling to 0.050 at C512, possible C128 outlier from H-C. Falsifier that matters: monotone INCREASING gap -> H-B active even sub-42M -> per-size beta recalibration becomes mandatory before any scaling claim. - Noise floor: seed band ~±0.006 (amp 3-seed); single seed per size -> differences <0.01 are NOT interpretable; if the trend lands inside the band, extremes get 3 seeds before any conclusion. RESULTS (2026-07-15, all 8 runs sealed, wandb ept-tinystories-gapscaling 8/8 live-streamed): | C | params | EP best | BP best | gap | final gate cos | skips | |---|---|---|---|---|---|---| | 128 | ~3.4M | 1.5562 | 1.5519 | 0.0043 (=0 in noise) | 0.805 | 25 (all gn) | | 192 | ~6.9M | 1.4257 | 1.4222 | 0.0035 (=0 in noise) | 0.996 | 0 | | 256 | ~11.4M | 1.3502 | 1.3353 | 0.0149 | 0.998 | 0 | | 384 | ~24.4M | 1.2898 | 1.2766 | 0.0132 | 0.996 | 0 | | 512 | 42.75M | 1.2808 | 1.2311 | 0.0497 | ~0.994 | ~0 | (128-384 pairs amp-consistent; C512 anchor pair fp32-consistent; amp lossless ±0.006 -> tiers comparable.) VERDICT — my registered call was WRONG, and so is every simple hypothesis on the list: - H-A (bigger = more noise-robust -> gap shrinks): FALSIFIED. Gap RISES in tiers toward C512. - My registered call (mild H-A, 0.06-0.10 at C128): FALSIFIED. C128 gap is ZERO. - H-C (recipe detuned at small C -> C128 inflated): FALSIFIED. Smallest sizes are the cleanest. - H-B as loop-gain/wall-2: NOT the mechanism here — zero skips at 192-384, no guard activity, cos flat ~0.996 across 192-384. Nothing wall-2-shaped below 42M. - The pattern that survives: the EP tax tracks FIT DEPTH, not width per se. Capacity-bound runs (C128/192, high floor) pay ~nothing; as runs become optimization-bound the tail-SNR tax appears (0.013-0.015 at 256/384) and compounds at C512 (0.050) where gn falls lowest. The C128 anomaly nails the point from the other side: gate cos 0.805 + 25 gn-guards, yet ZERO gap — direction noise alone does not cost validation CE when the loss floor is capacity-set. Consistent with RESULT 22: the binding constraint is late-phase SNR at low |g|, and centered@big-beta (which won exactly there) is the validated counter. fw72m_cent (RESULT 24, running) tests it at 72M. - Caveats before this becomes a paper figure: single seed per size (256/384 gaps ~2x band — likely real, certify with 3 seeds at C256 and C512); gap measured at matched STEPS on the same token stream (matched-data, not matched-compute); C512 anchor recipe is the tuning point. ### RESULT 22 (2026-07-14): TAIL-SNR SEVEN ARMS — centered estimator at BIG beta WINS; momentum-on-ghat REFUTED. Setup: resume stage1b_ep_muon_s45000.pt, run 45k->55k (10k tail steps, the low-|g| regime where wall-1 bites), all --amp, common recipe; one knob per arm. Reference: original run at s55000 = 1.2808 (fp32; data order differs after resume, so judge arms vs arm_ctl, not vs 1.2808). | arm | tail beta | knob | best val CE | last gate cos | verdict | |---|---|---|---|---|---| | arm_ctl | 1e-3 flat | none (muon .95) | 1.2678 | 0.994 | baseline | | arm_cent_f3e3 | 3e-3 flat | --est centered | **1.2591** | **0.997** | **WINNER (-0.009)** | | arm_adamw_f1e3 | 1e-3 flat | --opt adamw | 1.2664 | 0.994 | tie (-0.001) | | arm_rich_f1e3 | 1e-3 flat | --est richardson | 1.2830 | 0.986 | LOSES (+0.015) | | arm_mom99_f1e3 | 1e-3 flat | muon_mom 0.99 | 1.2867 | 0.994 | LOSES (+0.019) | | arm_mom99_f3e4 | 3e-4 flat | mom 0.99, beta/3 | 1.3266 | 0.993 | LOSES (+0.059) | | arm_mom995_f1e4 | 1e-4 flat | mom 0.995, beta/10 | 1.3440 | 0.973 | LOSES (+0.076) | Readings: - CENTERED at 3x beta wins BOTH CE and cos: O(beta^2) bias lets beta ride high -> readout noise /3 -> SNR up. Cost: 2 nudged phases, measured ~1.8x step time (15.1 -> 8.2 it/s). THE key that opens the wall-1 tail lock. - MOMENTUM-on-ghat REFUTED in all three doses: at matched beta it loses 0.019; using momentum to BUY lower beta (the sqrt-N averaging idea) loses monotonically more (0.059, 0.076). The noise is not zero-mean-averageable at the update level the way the hypothesis needed (Muon orthogonalization + staleness at decaying LR). - Richardson loses at matched beta: the 2g(b)-g(2b) combination amplifies variance ~sqrt(5)x — strictly dominated by centered-at-big-beta. - AdamW == Muon at the tail (1.2664 vs 1.2678): the NS-orthogonalization noise-amplification suspicion is NOT confirmed; no reason to switch (Muon carried the 0.050 epoch gap). - ctl at FLAT 1e-3 (1.2678) beats the original decaying schedule at s55000 (1.2808): more evidence the tail wants BIGGER beta, not smaller — consistent with the rising SNR floor picture. - CROWN-RERUN NOTE (72M and up): wall-2 caps beta from ABOVE there, so "centered + big beta" must become "centered + beta pinned at the bcap ceiling" — bias falls to O(beta^2) at unchanged loop gain. Momentum is off the table; bcap-v3 (attack threshold ~0.92) remains the other half. ### RESULT 21 (2026-07-14): CROWN SEALED — 72.11M x 1.44B tokens, fully BP-free, largest to date. fw72m_c2 reached 234,000 steps = the full Chinchilla budget (segments: 0-195k original schedule + 195-215k blow-recipe + 215-234k bcap-v2; final segment skips=1, drift 0.016 — the wall managed). **Best model: val CE 3.7117 (s185000, ~1.13B tokens) vs BP twin 3.2884 -> headline gap ~0.43** (disclosed as un-recalibrated transfer; mechanism = beta window, RESULT 19/20). "NENG-KAN" GATE PASSED on FineWeb register: fluent, on-topic, syntactic English (factual coherence not expected at 72M on open web; BP twin equally confused). CLAIM NOW LIVE: **the largest neural network fully trained without backpropagation to date (72.11M > KHS 62.7M), and the first transformer LM at that scale** — crown + first stack together. Samples: runs/fw72m_samples.txt; gen tool now data-aware. wandb: team eqprop-llm-training, split projects (ept-tinystories-42m / ept-fineweb-72m) + reports; workspace default-visibility gotcha (newest-10) fixed by replaying headline runs last. ### RESULT 20 (2026-07-14): WALL-ZONE PROBE VERDICTS — damping REFUTED (2 doses + adapt), beta-down SAILS; crown completion launched. Six arms, s195000 -> 201000 (the full crossing zone), identical data/recipe otherwise: | arm | val@201k (best) | skips | drift@end | verdict | | ctl floor-1e-3 | 4.55 (3.87) | 611 | 0.460 | crossing REPLICATED (probe validity) | | geta07 damp-0.7 | 4.57 (3.90) | 2373 | 0.481 | WORSE than ctl — damping refuted, dose 1 | | **blow floor-3e-4** | **3.94 (3.82)** | **6** | **0.015** | **sails the wall; best quality of all arms** | | geta05_k5 | 4.25 (3.89) | 3 | 0.451 | still at ceiling — damping refuted, dose 2 | | adapt (relax_tol, kmax 12) | 5.79 diverging (killed) | — | 0.494 | iterate-longer AMPLIFIES a divergent map (12 sweeps of gain>1 vs 3) — third monotone-spectrum witness | | bcap v1 (rho-cap) | 4.12 (3.91) | 0 | 0.013 | SURVIVES but starves: cap slammed beta to 2e-5 (meter reads noise/noise~1 at tiny residuals -> never recovers). v2 needs an absolute-scale gate on rho + slower attack | CONCLUSIONS: (1) spectrum is MONOTONE-POSITIVE (damping mathematically can't fix; 3 independent witnesses); (2) beta-reduction is THE working lever (loop gain ~ beta, linear); (3) the wall-1 fix (raising the floor to 1e-3) directly CAUSED the wall-2 crossing — the two walls are one beta-window story; (4) beta 3e-4 is NOT SNR-starved at this scale (blow's best 3.82 beats ctl's pre-wall plateau) — the "floor must be 1e-3" calibration was another absolute-constant transplant error. LAUNCHED: fw72m_c = crown completion from s195000 with the blow recipe (late floor 3e-4), 2xDDP, ETA ~3.3h; arms7 (tail-SNR momentum/centered/richardson/adamw suite) sequential on GPU1 overnight. ### RESULT 19 (2026-07-14): WALL-2 THEORY SESSION — loop-gain decomposition, the beta WINDOW, and three new controllers. **Mechanism formalized.** The nudged solve is a fixed-point iteration whose per-sweep error gain factorizes as **rho ~ beta x ||output curvature|| x ||down J^T chain|| x ||up J chain||** (follow the error once around the clamp-closed loop: top-force Hessian ~ sigma(W_out)^2, force chain down, rebuild chain up). Feedforwardness does NOT protect the iteration — the beta-clamp + force chain CLOSE a loop through the stack. The continuous flow stays unconditionally stable (solver wall, not physics wall; analog has no ceiling). Confirmations: f3e3 (3x beta -> earlier crossing), w2_blow (beta down -> drift 0.062->0.027), sig telemetry (sigma 430 vs stage1b ~90 => curvature factor ~20x). **The beta operating WINDOW**: beta_min (wall-1 SNR floor, rises as true grad shrinks) < beta < beta_max (wall-2 loop-gain ceiling ~ margin/(sigma^2 ||J||^2), falls as training sharpens). fw72m died because the fixed floor 1e-3 ended up ABOVE the falling ceiling. Window at 195k was still OPEN below 1e-3 (blow healthy AND beating ctl on val) — the crash was constant-transplanting, not a closed window. **Probe mid-flight verdicts (s195000 -> 201k):** ctl replicating the crossing trajectory; blow (3e-4) healthiest (drift 0.027, val 3.909 < ctl 4.073); geta07 FAILING (drift pinned 0.498, val 4.236) => monotone/positive-spectrum divergence suspected — damping structurally ineffective there; geta05_k5 = second damping dose (clean refutation if it also fails); w2_bcap + w2_adapt chained. **New machinery (committed):** (1) adaptive relax (--relax_tol: sweep-to-tolerance + rho-triggered geta backtrack + FINAL FULL-STEP graphed round — the E-read identity (z-o)=d REQUIRES undamped last substitution; mixing there leaks iteration residual into E (gn 1e5 bug, caught by smoke)); (2) rho^ meter (per-sweep residual ratio = free live loop-gain gauge, GOV['rho']); (3) **beta-cap-by-loop-gain (--beta_cap_rho): rho^>thresh -> cap *= 0.8, cap OVERRIDES the floor** (the ceiling can sit below the floor near the wall; survival first). Paired smokes at the wall ckpt: adaptive 0.771 vs legacy 0.754 cos (no regression). **Control-map placement (aep-dynamics toolbox, second in-vivo transfer):** old drift>0.5 guard = lagging column (silent through all precursors); rho^ servo = leading/online column; candidate construction-column addition = **sigma-cap on W_out** (one clamp hits BOTH walls' drivers: the wall-2 curvature factor AND the wall-1 governor collapse). K exonerated again at 72M (K3 vs K8 cos 0.910 vs 0.899, probe at s140000). ### RESULT 18 (2026-07-14): WALL-2 RETURNS AT SCALE — fw72m diverged at ~195k steps; f3e3 confirms beta-stress. Timeline (fw72m, 72M/FineWeb/32k): healthy to 168k (drift 0.013-0.017, skips 0, best val 3.7117 @~184k ~= 1.13B tokens); precursor drift SPIKE 0.044 @172k (transient); boundary contact 192-196k (drift 0.055-0.069, first skips); MASS CROSSING @200k (drift pinned 0.46-0.475 vs guard 0.5, skips +1000/4k = 25% reject rate, val 3.98 -> 5.5 divergence on the biased surviving subset). **f3e3 branch (floor 3e-3 from 140k) crossed the SAME wall EARLIER and harder (drift rejects 14.7k, train 27) — nudged displacement ~ beta => beta is the relaxation stress amplifier. The user's "beta 大了不收敛" is this, live.** Both runs killed (ckpts intact to s215000; best-val assets preserved). LESSONS: (1) the 42M/TinyStories architecture cure (norm placement) DELAYED the crossing past 58.8k there but did NOT eliminate the mechanism — theta drift crossed at 195k in the bigger/harder regime; "wall-2 ELIMINATED" is rescoped to "wall-2 delayed beyond horizon at 42M/TS". (2) drift-creep + spike is the leading indicator (the aep-dynamics control-map discipline is now LOAD-BEARING for the LLM line — first in-vivo transfer of the paper's machinery at scale). (3) My blowup watchers were blind (awk field bug, $6='val' string) — fixed pattern: grep -oE. FIX PROBES (running, from healthy s195000 through the full wall zone to 201k): w2_ctl (replicate), w2_geta07 (damped mixing 0.7 = the principled contraction restorer), w2_geta05_k5, w2_blow (floor back to 3e-4 = beta-stress direct test). Crown status: best ckpt 3.7117@1.13B tokens is a trained artifact but the clean crown re-run waits for the fix verdict. ### RESULT 17 (2026-07-13): fw72m LAUNCHED (the 62.7M-crown run) + size ladder complete + resume upgraded. - **fw72m**: 72.11M (L12 C512 @32k vocab) x 1.44B FineWeb-Edu tokens (Chinchilla), --amp, first NCCL 2-GPU DDP production run (GPU0+3, B12/rank = global B24 preserving recipe semantics). Startup: world=2, cos 0.9999, zero skips, **2.92 it/s -> ETA ~22h**. Crown context: KHS VGG10 = 62.7M (verified from their Table 6: convs 9.2M + dense 25088->2048 = 51.4M + head 2.05M = 62.65M). BP twin queued for GPU1 (waiting on user's phasescan). Corpus: 9.99B tokens, doc-level shuffled (9.67M docs, seed 1234, val re-drawn 20M disjoint) per user directive. - **nsize ladder DONE** (TinyStories 4k-vocab, 3k steps, fp32, seed 1): c256 1.9265 / c512 1.7920 / c768 1.7913 / c1024 1.7876 — at FIXED 3k steps quality saturates with width (data/steps-bound, expected); purpose = noise-robustness probe ckpts (runs/nsize_*_s3000.pt x4). Probe script queued. - **Resume upgraded to exact**: MultiOpt gains state_dict/load_state_dict; trainer saves 'opt' in every ckpt and restores it on --resume (chunked HPC jobs no longer lose Adam/Muon state). - **Delta storage recon**: /scratch 1.5T/1.5T FULL, /work/hdd/bfqt over quota -> 300M data transfer BLOCKED until space found (own old-ept footprint = first cleanup candidate). A100x4 queue ~8.5d, A40 same-day (48h cap -> needs exact resume, now DONE). ### RESULT 16 (2026-07-13): STAGE-2 DATA PIPELINE LIVE + FINEWEB SMOKE PASSED. `prepare_fineweb.py`: FineWeb-Edu sample-10BT -> 32k ByteLevel BPE (<|eot|> id 0) -> uint16 bins, tinystories_bpe format, `--data` flag added to both trainers. SMOKE_READY in 11 min (download 5min @80MB/s, tokenizer train 39s on 1.5GB, shard0 5min = 755M tokens; val = first 20M, disjoint). Full 14 shards -> ~10.5B tokens (running, ~50 min ETA at 754M/5.4min per shard). **fw_smoke (L12 C512 32k-vocab = 72.11M, --amp, 400 steps, GPU1): CE 10.51 -> 5.84, cos(EP,BP) 0.9999@0 / 0.9951@100 / 0.9995@400, ZERO skips, drift 0.002.** The estimator + amp + beta-governance (sig grew 3.9->58, beta floored by step 100 — wall-1 machinery engaged correctly on the harder corpus) all transfer to real web text at 4x vocab unchanged. Speed 0.89 it/s at this shape (bigger head). Stage-2 recipe question OPEN for user: T=1024 (web-native context) vs T=256 (strict TinyStories comparability) for the 300M run. NOTE: the "BP twin 2.9746" line in DONE prints is the stale TinyStories reference (cosmetic); no fineweb BP twin exists yet. ### RESULT 15 (2026-07-12): bf16 MIXED PRECISION (--amp) VALIDATED — lossless at 4k, 1.56x wall-clock. The /2-class cost lever, same-day pipeline: amp_gate.py static gate -> trainer flag -> 3-seed A/B. SEMANTICS (why this lives while naive-cast --bf16 is dead): params/states/displacements/E-accum stay fp32; ONLY block forwards run under autocast(bf16). RESULT 11's naive-cast death = pure STATE quantization (wall-1: beta-displacement below bf16 resolution) — exactly as diagnosed. - Gate (stage1b s55000, fp64 cosine): amp cos(EP,BP_fp32) 0.9682 vs fp32-EP 0.9687 (zero loss); beta=3e-3 -> 0.9878, 1e-2 -> 0.9966 (bigger beta ACTIVELY better — wall-1 SNR physics); bf16 fwd valCE -0.0002; BP_amp baseline 0.9993. amp_last (fp32 final rebuild) buys nothing -> amp_all everywhere; the E-subtraction term is not binding at production beta (fbnoise-tolerance prediction from RESULT 14 held: relative noise on forces is invisible). - 3-seed 4k A/B (bsign flagset + --amp): 1.7086/1.7569/1.7562 mean 1.7406 vs fp32 3v3 mean 1.7314 (Delta +0.009 inside the seed-noise band; amp_s1 BEAT the BP+Muon mean 1.7098). In-trainer bp_gate cos 0.9999 at step 0. Zero guard events. - SPEED (solo GPU3/A6000, C512): amp 2.785 it/s vs fp32 1.789 it/s = 1.56x wall-clock; grows with width (tensor-core-bound share) -> treat 1.5x as the floor for 1-3B on H100. - ~~CAVEAT + confirm step~~ **EPOCH CONFIRM SEALED (2026-07-12 late): stage1b_amp DONE best val CE 1.2868 vs fp32 1.2808 (Δ+0.006, inside the 0.02-0.03 best-of-noisy-val band); zero guard events over 58.8k; 2.80 vs ~1.79 it/s = the 1.56x held for the full epoch.** amp = Stage-2 default, full confidence. EMAIL_BEN_DRAFT2 gate #6 CLEARED (the sent "validated this week" claim is now closed at epoch scale). - Cost consequence: COST_MODEL.md v2.1 (sourced July-2026 prices: market H100 $1.87-2.99/GPU.h, AWS p5e blocks $4.97/GPU.h post-hike) — with amp measured, 3B-Chinchilla ~$40k / 7Bx20B ~$30k on AWS blocks: BOTH inside the $50k envelope individually. amp is the Stage-2 default. ## RESULT 75 (2026-07-30): dg128 全程验证完赛 — 口径分裂,class-2 归档 Matched 三口径(同窗 434-440k, n=61/臂): | 臂 | best | 尾窗中位 | |---|---|---| | dg128 (250k 起治疗) | 3.1387 | 3.3380 | | bsign 未治疗 | 3.3839 | 3.5649 | | BP twin | 3.0902 | 3.2051 | - best 口径: gap +0.049, 回 72M 带(+0.076)以内, 关闭 84%; **尾窗口径: gap +0.133, 未回带, 关闭 63%**。 规范裁决用尾窗; best 是噪声评估的下探偏置(三臂都低 0.11-0.18, 含 BP → 口径病不是臂病)。 - 仪器 97% vs 全程 63% 的差值解释候选: (a) **续跑只治后 190k 步** — 前 250k 未治疗欠账已烙进权重 (dg128 resumed from untreated s250000, 探索性设计的固有混杂); (b) 剂量需求随走廊下沉增长, 6k 仪 器段低估长程需求。谱仪轨迹轴 (c768f23 vs c768f91) 直接裁决 (b); 裁 (a) 需 from-scratch treated 臂 (~22h×3GPU, 等 geo 电池 + 谱仪结果再决定是否值得)。 - 稳定性结论无保留: ×128 位移 190k 步零发散(drift 稳 0.008), 大位移在完整训练尺度安全。 - 按 07-29 条令: class-2 exploratory, 不入 scaling 曲线; 价值 = "泄漏在全程可部分赎回" + 稳定性 + 上述口径教训。