summaryrefslogtreecommitdiff
AgeCommit message (Collapse)Author
2026-07-10RESULT 9: wall-2 ELIMINATED by OLMo2 arch — epoch cleared 10-14k with zero ↵Yuren Hao
guard events (old arch: blowup@12100); cos-erosion watch armed
2026-07-10RESULT 8 prelim: Muon+beta_floor on EP = 1.7316 (-0.098 vs AdamW, n=1, zero ↵Yuren Hao
skips); retraction closed; BP+Muon control + s2 launched
2026-07-10RESULT 7: OLMo2 parity PASS (BP 1.8335 vs EP 1.8643, gap +0.031 within ↵Yuren Hao
metric noise); arch worth ~0.08 CE both columns; epoch auto-launched; Pascal unbanned (canaries green), farm running 2xBP epochs + Muon retest
2026-07-10RESULT 6 addendum: K8 also blew at 12600 -> wall-2 is true contractivity ↵Yuren Hao
crossing (K delays only); kretry demoted to mitigation; defense = OLMo2 arch > lr channel > damped-fb/jacreg
2026-07-10RESULT 6: wall-2 = EP-specific marginal under-convergence at the ↵Yuren Hao
contractivity edge (A skips16/B skips2/C skips1/D-BP clean); ship --kretry (retry marginal batches at K=kmax); launch OLMo2 parity matrix with parity-gated epoch autolaunch Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-10HW dossier: OLMo2-standard block analog audit (net win; no new ↵Yuren Hao
non-reciprocity; 3 new E-tier tolerance items)
2026-07-10OLMo2-standard architecture (user directive): norm-after-sublayer RMSNorm ↵Yuren Hao
blocks, full-width QK-norm, RoPE(500k), SwiGLU, no-bias, untied head, final RMSNorm, 0.02 init, grouped wd, optional z-loss (plumbed through nudge+readout+gate for EP exactness); EP smoke 42.75M cos=1.0000 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-10AUDIT: retract confounded Muon verdict; tone down parity claims (n=3, ↵Yuren Hao
best-of-noisy-val); scope depth-tax claim to 4k horizon; guard-split skip telemetry (skd/skg); add BP control arm diag_D_bp (+--resume in BP trainer); record resume confounds Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-10Stage-1 epoch blowup @12100: sig story REFUTED (only +8%, cos fine till ↵Yuren Hao
after); leading indicator=drift-guard skips -> contractivity bifurcation in nudged relaxation (cascade Hopf wall); add resume/sig0/final_ln; A/B/C diagnostic launched
2026-07-10RESULT 4: QK-norm validated parity-preserving (EP 1.8868 <= BP 1.9066); ↵Yuren Hao
qk_norm complements but does NOT replace beta-floor (sig_tok still grows); Stage 1 TinyStories epoch launched
2026-07-10QK-norm: replace nn.MHA with explicit SDPA attn + --qk_norm (OLMo2-style, ↵Yuren Hao
analog-friendly); cancel epoch, insert 8-run qk validation, stage roadmap TinyStories-epoch->FineWeb-Edu->OLMo2
2026-07-10RESULT 3: K4 depth-parity SEALED EP-favorable (EP 1.8907 <= BP 1.9192 @ ↵Yuren Hao
L12xC512); full-epoch generation run auto-launched (6.9h)
2026-07-10Add --cosine (warmup->cosine decay to 0.1x lr) to both cascade trainers for ↵Yuren Hao
full-epoch runs
2026-07-10CORRECTION: D1a did not die (misread mangled proc scan); originals completed ↵Yuren Hao
s1 1.9745/s2 2.0013/s3 2.1188 + muon 2.7515; beta-floor fix independently found on free GPU1 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-10RESULT 2: beta-floor confirmed as L12 depth fix (3e-4 holds cos=1.0 vs ↵Yuren Hao
control collapse 0.896); K4 3-seed verdict launched vs BP 1.9192 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-10K refuted as L12 lever (cos K-invariant K3==K8); BP 3-seed sealed 1.919; ↵Yuren Hao
mechanism = finite-beta SNR collapse; add --beta_floor/--beta_fixed + launch floor sweep Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09D1a autopsy + K-ladder diagnostic: parent-death (not nohup'd) killed 4 arms; ↵Yuren Hao
L12 cos-erosion signal; relaunched K3/K8/BPs3 nohup Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09notes refresh: plan status stamp (K1-K3 sealed, D-tier in flight), ↵Yuren Hao
ONBOARDING cascade-pivot section Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09FINDINGS: 2026-07-11 consolidated hardening entryYuren Hao
2026-07-09HW: cascade re-pricing under the reuse doctrine — depth becomes a memory ↵Yuren Hao
line-item (time-mux one trainable block), nudged duty drops ~50x, leash instrumentation deleted, SRAM-CIM alone suffices for Demo-1 (device fab optional) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09HW economics retraction: doctrine currency restored (COTS-stitch + FPGA + ↵Yuren Hao
UIUC collab, $5-20k Demo-1 BOM); optimizer state digital-side by design => Adam ~free, Muon unblocked at demo scale, choice is algorithmic Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09HW BoM correction: Lion/SGD ruled out for EP training (empirical, abl_lion2 ↵Yuren Hao
blew @600 + prior tests) — factored-Adam is the MANDATORY analog baseline, not premium tier Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09BP-free formal audit (5.6-sol): EP update verified local (108 bitwise ↵Yuren Hao
zero-influence checks); fixes — gate_every<=0 = true off switch, governor reaction now opt-in (--gate_govern, default observe-only), dFdtheta restored as self-sealing invariant-test surface; test_bp_free.py 4/4 green Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09HW: optimizer BoM audit per cascade block — Adam analog-compatible via ↵Yuren Hao
factored-AGC (+2x memory, zero ADC), Muon priced out of analog (matrix-matrix NS + full-gradient ADC tax), Lion native baseline Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09A0.4 precision gate (TF32 harmless cos 0.9946==fp32; pure-bf16 cos 0.9427) + ↵Yuren Hao
Muon hybrid optimizer (--opt muon) wired into both trainers; D1a flagship matrix launched (L12xC512, 3xBP + 3xEP + Muon arms) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09FINDINGS: K3 sealed — matched-tuning parity n=3 (EP 2.0500+-0.015 vs BP ↵Yuren Hao
2.0530+-0.004), the three-act retraction arc documented Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09casc_eq_train: exact mode (dtop_every=1) now DEFAULT — bisection showed ↵Yuren Hao
the v7 dedup cost ~4% CE at tuned lr; fast mode kept as opt-in. K3 verdict: exact equilibrium-EP 2.0481 vs BP 2.053+-0.004 = matched-tuning PARITY Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09HW dossier: analog softmax/attention (12 verified claims) — Elfadel-Wyatt ↵Yuren Hao
1993 reciprocal 'entropic resistor' softmax (EP-composable, PAR escape realized), charge-domain KV gain-cells (Leroux 2025), the analog-training gap confirmed unclaimed; +-1% mismatch->+-1% softmax linear tolerance coeff Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09casc_eq_train v7: two algorithmic dedups — d_top refreshed every other ↵Yuren Hao
round (O(beta^2) error), last-rebuild graphs reused by the theta-readout (kills one full graphed chain per step); fp32 untouched for fair BP twinning Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09casc_eq_train v6: graph-reuse fb (rebuild graphs feed next round's vjps), ↵Yuren Hao
--compile flag, tok-sigma amortized every 25 steps — 2.23 -> 2.68 it/s on 1080 (3.28x BP) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09FINDINGS: final C1 verdict — equilibrium fb-EP 2.2949 beats BP 2.3694 ↵Yuren Hao
(same init/seed/steps), cos 1.0000 throughout, 3.9x BP cost Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09FINDINGS: cascade pivot one-day arc — bridge, solver saga, fb ↵Yuren Hao
message-passing, init root cause, diagC 2.9021 result Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09cascade root cause found: tied readout + default N(0,1) embedding init = ↵Yuren Hao
pathological top-CE stiffness once predictions sharpen (sigma_tok~76 vs GPT-standard 0.02 giving ~1.6); add --untie and --tok_init; diag arms A/B/C Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09casc_eq_train v4: quality-governed estimator — periodic measured ↵Yuren Hao
cos(EP,BP) drives K/beta (spend when quality drops, relax when abundant); guards reduced to sanity-only Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09casc_eq_train v3: adaptive beta (tok-sigma stiffness compensation), ↵Yuren Hao
adaptive-K fb with contraction test, non-contraction bail, in-training cos(EP,BP) telemetry Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09cascade equilibrium solver solved: fb (forward-backward message passing) K=3 ↵Yuren Hao
beta=0.003 — gates 1.0000/0.9990/0.9946 across BP trajectory, L12 0.9999; trainer v2 single-sided EP readout + divergence guard (~5x BP) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09casc_eq_train: equilibrium-mode cascade-EP trainer (GS-reverse solver K ↵Yuren Hao
sweeps, two-phase ±β, EP readout at relaxed states) — the true-EP route C1 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09cascade: retire zil per user directive (EP route = equilibrium two-phase); ↵Yuren Hao
probe gains --init_sweep (state warm-start, EP-clean readout) and --sopt adam; B2 solver sweep Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09cascade: zil interleaved scheme = exact gradients at all depths (cos 1.0000 ↵Yuren Hao
L6-24, trajectory 0.9998+); casc_ep_train zil trainer; naive-relaxation depth failure documented Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09cascade ablation program: plan doc (5 claims, tiers 0-4), probe v2 ↵Yuren Hao
(jacobi/gsf/gsr schemes, full-theta gate, multi-batch), casc_bp_train ckpt producer Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09cascade_probe: layered-energy EP on L distinct standard blocks — free ↵Yuren Hao
equilibrium == plain forward (standard LLM inference), two-phase grad gate cos 0.9968 vs BP (L3 C128) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09lt_ep_train: --gov_keepjr now keeps the jr floor through reg_delay too ↵Yuren Hao
(fastfull arm: jr always-on, resreg post-warmup) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09backfill_wandb: replay historical run logs into W&B (ept-c512 / ept-tol / ↵Yuren Hao
ept-33m) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09lt_ep_train: --rr_floor (resreg floor during free/delay phase) for govfloor armsYuren Hao
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-09lt_ep_train: optional W&B mirroring (--wandb/--wandb_run), failure-proof ↵Yuren Hao
logging of val/ema/best/res/jr/lr/rho/gov state Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-07FLAGSHIP: warm_fast 1.7065 final — beats tuned BP anchor by 0.086, ↵Yuren Hao
param-matched twin by 0.257 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-07Wall-3 theorem addendum: arXiv:2603.26969 (PAR) hardens the J^T obstruction; ↵Yuren Hao
price list + Demo-M reframe Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-07SPECTRAL GOVERNOR (--governor): the from-scratch recipe candidate, ↵Yuren Hao
dip-map-calibrated Dip screening (4/4 seeds have dips; first-dip cluster 1200-1700, width 100-300 steps, depth |lam|~0.998 = s2000-class; s2 shows excursion->dip in sequence): reg-free until a scout (warm lead_rho, every gov_k=100, deep-400 subbatch) flags rho<0.985 after gov_min=1000, then ARPACK-certify all top-3 |lam|<1 -> leash ON (pair regs). Post-engagement: sustained rho_scout>1.02 x2 -> rescue boost (resreg x3 for 500 steps; the proven hr2 maneuver). abl_delay's fixed-2000 death explained: it engaged AFTER the dip cluster, mid-excursion. Two governor seeds launched on 107 (trained-state engagement is Pascal-safe per hr2 precedent). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-06dp_ep smoke PASSED: 2-GPU lockstep clean, ~107% linear scaling (allreduce ↵Yuren Hao
negligible vs seconds-long EP steps) val 6.0->5.83/100 steps at effective B48, res nominal, ranks never diverged. Speed package final ledger: pack 1.47x X DP ~linear (X optional AA-v2 1.3x). dip_screen.py: ARPACK screener for the dip-farm trajectories (running on freed 1080s, 2 chains). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
2026-07-06Tier-3 gates: Anderson math-yes/impl-no (parked for v2); bf16polish UNSAFE ↵Yuren Hao
near-edge (eval-only); dp_ep.py ready Anderson: res 25-35x deeper per budget but 5.8x slower (naive history stacks + per-iter safeguard eval) — v2 = ring buffers + periodic safeguard, est +1.3x on the speed tier. bf16+20polish: 1.41x free phase, res parity, BUT z-diff 1.2e-3 — near-marginal operators contract too slowly for a 20-step polish (0.998^20≈0.96), same magnitude as the TF32 kill verdict and the estimator's 50%-sensitivity input. Predicted by our own depth/noise theory. Flags kept with warnings; neither ships for training. dp_ep.py: manual-allreduce EP DP (controller in lockstep, aligned collectives), smoke pending freed 1080s. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn