| Age | Commit message (Collapse) | Author |
|
best-of-noisy-val); scope depth-tax claim to 4k horizon; guard-split skip telemetry (skd/skg); add BP control arm diag_D_bp (+--resume in BP trainer); record resume confounds
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
after); leading indicator=drift-guard skips -> contractivity bifurcation in nudged relaxation (cascade Hopf wall); add resume/sig0/final_ln; A/B/C diagnostic launched
|
|
qk_norm complements but does NOT replace beta-floor (sig_tok still grows); Stage 1 TinyStories epoch launched
|
|
analog-friendly); cancel epoch, insert 8-run qk validation, stage roadmap TinyStories-epoch->FineWeb-Edu->OLMo2
|
|
L12xC512); full-epoch generation run auto-launched (6.9h)
|
|
full-epoch runs
|
|
s1 1.9745/s2 2.0013/s3 2.1188 + muon 2.7515; beta-floor fix independently found on free GPU1
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
control collapse 0.896); K4 3-seed verdict launched vs BP 1.9192
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
mechanism = finite-beta SNR collapse; add --beta_floor/--beta_fixed + launch floor sweep
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
L12 cos-erosion signal; relaunched K3/K8/BPs3 nohup
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
ONBOARDING cascade-pivot section
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
|
|
line-item (time-mux one trainable block), nudged duty drops ~50x, leash instrumentation deleted, SRAM-CIM alone suffices for Demo-1 (device fab optional)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
UIUC collab, $5-20k Demo-1 BOM); optimizer state digital-side by design => Adam ~free, Muon unblocked at demo scale, choice is algorithmic
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
blew @600 + prior tests) — factored-Adam is the MANDATORY analog baseline, not premium tier
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
zero-influence checks); fixes — gate_every<=0 = true off switch, governor reaction now opt-in (--gate_govern, default observe-only), dFdtheta restored as self-sealing invariant-test surface; test_bp_free.py 4/4 green
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
factored-AGC (+2x memory, zero ADC), Muon priced out of analog (matrix-matrix NS + full-gradient ADC tax), Lion native baseline
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
Muon hybrid optimizer (--opt muon) wired into both trainers; D1a flagship matrix launched (L12xC512, 3xBP + 3xEP + Muon arms)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
2.0530+-0.004), the three-act retraction arc documented
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
the v7 dedup cost ~4% CE at tuned lr; fast mode kept as opt-in. K3 verdict: exact equilibrium-EP 2.0481 vs BP 2.053+-0.004 = matched-tuning PARITY
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
1993 reciprocal 'entropic resistor' softmax (EP-composable, PAR escape realized), charge-domain KV gain-cells (Leroux 2025), the analog-training gap confirmed unclaimed; +-1% mismatch->+-1% softmax linear tolerance coeff
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
round (O(beta^2) error), last-rebuild graphs reused by the theta-readout (kills one full graphed chain per step); fp32 untouched for fair BP twinning
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
--compile flag, tok-sigma amortized every 25 steps — 2.23 -> 2.68 it/s on 1080 (3.28x BP)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
(same init/seed/steps), cos 1.0000 throughout, 3.9x BP cost
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
message-passing, init root cause, diagC 2.9021 result
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
pathological top-CE stiffness once predictions sharpen (sigma_tok~76 vs GPT-standard 0.02 giving ~1.6); add --untie and --tok_init; diag arms A/B/C
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
cos(EP,BP) drives K/beta (spend when quality drops, relax when abundant); guards reduced to sanity-only
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
adaptive-K fb with contraction test, non-contraction bail, in-training cos(EP,BP) telemetry
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
beta=0.003 — gates 1.0000/0.9990/0.9946 across BP trajectory, L12 0.9999; trainer v2 single-sided EP readout + divergence guard (~5x BP)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
sweeps, two-phase ±β, EP readout at relaxed states) — the true-EP route C1
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
probe gains --init_sweep (state warm-start, EP-clean readout) and --sopt adam; B2 solver sweep
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
L6-24, trajectory 0.9998+); casc_ep_train zil trainer; naive-relaxation depth failure documented
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
(jacobi/gsf/gsr schemes, full-theta gate, multi-batch), casc_bp_train ckpt producer
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
equilibrium == plain forward (standard LLM inference), two-phase grad gate cos 0.9968 vs BP (L3 C128)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
(fastfull arm: jr always-on, resreg post-warmup)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
ept-33m)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
logging of val/ema/best/res/jr/lr/rho/gov state
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
param-matched twin by 0.257
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
price list + Demo-M reframe
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
dip-map-calibrated
Dip screening (4/4 seeds have dips; first-dip cluster 1200-1700, width
100-300 steps, depth |lam|~0.998 = s2000-class; s2 shows excursion->dip in
sequence): reg-free until a scout (warm lead_rho, every gov_k=100, deep-400
subbatch) flags rho<0.985 after gov_min=1000, then ARPACK-certify all top-3
|lam|<1 -> leash ON (pair regs). Post-engagement: sustained rho_scout>1.02
x2 -> rescue boost (resreg x3 for 500 steps; the proven hr2 maneuver).
abl_delay's fixed-2000 death explained: it engaged AFTER the dip cluster,
mid-excursion. Two governor seeds launched on 107 (trained-state engagement
is Pascal-safe per hr2 precedent).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
negligible vs seconds-long EP steps)
val 6.0->5.83/100 steps at effective B48, res nominal, ranks never diverged.
Speed package final ledger: pack 1.47x X DP ~linear (X optional AA-v2 1.3x).
dip_screen.py: ARPACK screener for the dip-farm trajectories (running on
freed 1080s, 2 chains).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
near-edge (eval-only); dp_ep.py ready
Anderson: res 25-35x deeper per budget but 5.8x slower (naive history stacks
+ per-iter safeguard eval) — v2 = ring buffers + periodic safeguard, est +1.3x
on the speed tier. bf16+20polish: 1.41x free phase, res parity, BUT z-diff
1.2e-3 — near-marginal operators contract too slowly for a 20-step polish
(0.998^20≈0.96), same magnitude as the TF32 kill verdict and the estimator's
50%-sensitivity input. Predicted by our own depth/noise theory. Flags kept
with warnings; neither ships for training. dp_ep.py: manual-allreduce EP DP
(controller in lockstep, aligned collectives), smoke pending freed 1080s.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
(0.91x, cos 0.94, avg free); compile demoted
Full-ep_step wall times on quiet A6000 (warm s2000, B24), res parity across
all 8 configs. compile only 1.12x at this shape (historical 1.46x was a
different workload split); FULL cmp_sdpa saves 4% over eager at t80 — not
worth the guard complexity. tforce_sdpa added (flash baked into compiled
graph, flag-free so grad paths never see SDPA). bp_lm --tie probe in flight.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
tooling
trend-aware stop + plateau averaging recovers the semi-convergence victims
exactly as predicted. Estimator pack now: holofast+sdpa+t2sel80(+holoavg).
Also: bp_lm stdinit/beta2/sched flags (anchor archaeology), dipfarm_freezer,
staged tol_sweep.sh (gated on hr2 verdict), bp_sweep.sh.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
(semi-convergence fix, ungated)
BP+EMA 2x2 (local, healthy env): qknorm {1.9951, 1.9977}, no-qknorm {1.9728,
1.9732} -> parameterization-matched BP twin tops at ~1.97 vs EP 1.7888: the
equilibrium computation's iteration/depth dividend = 0.18 CE from identical
parameters. External anchor (tuned depth-1 BP, 1.7921): EP at parity.
bp_lm on 107/2.3.1 gave 1.9823 ~= local -> plain backward exonerated on the
pascal env; the 107 divergence (pair/floss/resreg dead by step 600) narrows
to the EP reg/estimator loop. Single-step fingerprints all match (a/b/c/d) ->
suspected intermittent kernel issue; discriminators in flight (torch-2.7 env
probe + delay_hr2's step-2000 reg-on transition).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
SEMI-CONVERGENCE (not r, not anchor, not transients)
Refutation chain: r-sweep flat (0.02-0.4); deep anchor (res 100x tighter) no
gain; kappa brake monotonically harmful. t2sel window sweep at s2000:
40->80 lifts ALL batches (mean 0.889->0.936, truncation confirmed); past 80
SEMI-CONVERGENT (batch-dependent optimum; the inc-argmin t_best rule fails on
rotating slow modes -> batch2 degrades 0.956->0.875 at 320). Early stopping
IS the regularizer; iteration count = reg parameter.
Shipping insight: holofast + sdpa + t2sel80 ~= old default wall-clock with
cos 0.89->0.94. warm_fast (record) already ran t2sel=80 vs proven-scratch 40
— a real +0.05-cos hidden difference in the lineage table.
Next lever: trend-aware stopping / t_best-neighborhood averaging.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
nudge-SNR hypothesis refuted
The 50% single-shot sensitivity is MULTIPLICATIVE (near-marginal T2 dynamics'
state-sensitivity; signal and error scale with r together) — not additive
noise divided by 2r. cos ceiling ~0.88 at near-edge states is set by T2
truncation + state sensitivity (0.98 at deeply-contracted states). hr-0.2's
empirical wins are NOT estimator SNR. Hardware upside: algorithm indifferent
to r across 20x -> nudge amplitude free to fight ADDITIVE readout noise.
Surviving accuracy levers: kappa nbrake (Tikhonov, targets the real culprit),
tail-window averaging over argmin selection.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|
|
phase, z* parity 4e-7
Scoped via blk._sdpa set only inside relax()'s loop (grad paths jvp/vjp/resreg
keep the manual attention: no forward-mode-through-flash risk). Combined with
--holofast: ~1.51x full-step exact-math tier.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
|