From 411870d64f00181ea41a8059f4b53f5b232447c0 Mon Sep 17 00:00:00 2001 From: Yuren Hao Date: Sun, 12 Jul 2026 05:44:18 -0500 Subject: E-tier wave-1 ledger (RESULT 14 + HW map rows) + cost model v2 recalibrated on measured H200 datapoint Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn --- docs/campaign/CASCADE_ABLATION_PLAN.md | 24 +++++++++++++++++ docs/campaign/COST_MODEL.md | 47 ++++++++++++++++++++++++++++++++++ docs/hardware/COMPONENT_HW_MAP.md | 10 ++++---- 3 files changed, 76 insertions(+), 5 deletions(-) create mode 100644 docs/campaign/COST_MODEL.md (limited to 'docs') diff --git a/docs/campaign/CASCADE_ABLATION_PLAN.md b/docs/campaign/CASCADE_ABLATION_PLAN.md index 61860f6..8b52369 100644 --- a/docs/campaign/CASCADE_ABLATION_PLAN.md +++ b/docs/campaign/CASCADE_ABLATION_PLAN.md @@ -538,3 +538,27 @@ s1 alone beat the BP mean). Ledger of the residual-0.05 epoch gap after three pr Carry `--bsign_rand` and a future centered mode as Stage-2 A/B flags; the gap question re-opens at 300M/real-corpus where it means something. Effort pivots to: (1) Stage-2 data pipeline (FineWeb-Edu + 32k tokenizer), (2) E-tier tolerance suite on the idle farm (hardware track / UIUC outreach feed). + +### RESULT 14 (2026-07-12): E-TIER WAVE-1 — full analog-fault tolerance ledger at stage1b s55000. +`etier_probe.py`, farm GPUs 2/3/7 (shards A/B/C), stage1b_ep_muon_s55000.pt (clean valCE 1.2678), +B=8 eval batches; metrics = faulted valCE (Δ vs clean), cos(EP_faulted, BP_faulted) [self-consistency +of the learning signal under fault], cos(EP_faulted, BP_clean) [direction vs the ideal update]. + +| fault (component) | mild | medium | severe | verdict | +|---|---|---|---|---| +| wq — weight quant (crossbar #3) | 8b: +0.004 / 0.956 | 6b: +0.051 / 0.812 | 4b: +2.23 / 0.05 | **8b FREE, 6b marginal, 4b dead → ≥7b effective is the binding spec** | +| fnoise — fwd additive state noise (softmax/relax #5) | 1e-3: 0.000 / 0.975 | 3e-3: 0.000 / 0.974 | 1e-2: +0.001 / 0.970 | **FREE at 1% — looped-era 1e-3 cliff does NOT transfer to cascade** | +| divmis — divisive-norm mismatch (#4/#7) | 1%: 0.000 / 0.971 | 3%: +0.003 / 0.948 | 10%: +0.042 / 0.785 | 3% (routine matching) FREE; 10% marginal | +| rope — phase error rad (#2) | 0.01: 0.000 / 0.973 | 0.03: +0.001 / 0.967 | 0.1: +0.012 / 0.923 | 0.03 rad FREE; ~2° I/Q accuracy suffices | +| gilbert — gate gain error (#6) | 1%: 0.000 / 0.974 | 3%: 0.000 / 0.971 | 10%: +0.006 / 0.945 | **FREE at 10%** — translinear practice is comfortably inside | +| fbnoise — nudge/error-channel noise | 1e-2: 0.969 / 0.975 | **1e-1: 0.951 / 0.957** | 3e-1: 0.764 / 0.768 | **10% relative noise on the ERROR CHANNEL is FREE** (cos 0.95) — the r-indifference/large-nudge gift, now measured on cascade | + +Reading: (a) the only hard constraint is crossbar weight precision (≥7b effective — inside standard +SRAM-CIM capability; 6b rescue = wave-2 quant-aware co-training); (b) everything dynamic — forward +noise 1%, error-channel noise 10%, gate/divider/phase mismatch at routine device tolerances — is +FREE at this scale. cos(EP,BP_faulted) stays ~0.97 under every non-fatal fault: the EP estimator +tracks whatever network the faults define, i.e. learning co-adapts to the fault (the analog-training +thesis in one number). CAVEAT: static probes at a trained checkpoint (eval CE + one-step gradient +direction), not training-under-fault; wave-2 = co-training with faults injected from step 0 +(expectation from the literature and from (c): tolerances IMPROVE). Feeds COMPONENT_HW_MAP.md +(per-row status updated) + UIUC outreach dossier. diff --git a/docs/campaign/COST_MODEL.md b/docs/campaign/COST_MODEL.md new file mode 100644 index 0000000..2b9d8b6 --- /dev/null +++ b/docs/campaign/COST_MODEL.md @@ -0,0 +1,47 @@ +# Cost model v2 (2026-07-12) — recalibrated on a measured rental datapoint + +Supersedes the v1 table in EMAIL_BEN_DRAFT2.md (drafts ≤5) and the numbers quoted in the +2026-07-11 scale discussion. **v1's error: anchored $/FLOP on AWS-A100-savings TF32 +(~$19.7k/EF) — the wrong reference class for 2026.** Market H100/H200 rentals are 2.5–3× +cheaper per FLOP. + +## Anchors + +| anchor | value | source | +|---|---|---| +| **Empirical all-in ceiling** | **$16.7k/EF** | user's real run: 1B × 10B tok BP on rented H200s ≈ $1k all-in; 6ND = 0.06 EF. Implied throughput only 42–58 TF/s eff (4–6% of bf16 peak) — an UN-optimized run, so this is a ceiling, not a target | +| Market rental rate | $2–3.5/GPU·h | H100/H200 marketplace (Vast/RunPod class), 2026 | +| Our stack's utilization | ~19% of TF32 peak | measured on A6000 (14 TF eff / 75 TF TF32 peak), eager TF32 | +| → projected H200 throughput | 75–125 TF/s eff | 15–25% × 494 TF TF32 peak | +| → **working rate (TF32 path)** | **$6–9k/EF** (mid $7.5k) | $2.5–3/h ÷ 75–125 TF/s | +| EP/BP FLOP ratio | **3.2×** (measured) | EP ≈ 19.2N FLOPs/tok vs BP 6N; single-sided, K=3 | +| Failure allowance | 1.3× | failure modes characterized; ckpt every 5k steps | +| AWS multiplier | 1.5–2.5× on-demand; 1.2–1.5× spot/capacity-block | vs marketplace $/GPU·h | +| Mixed-precision upside | ÷~2 | proper autocast (bf16 matmul, fp32 states + E-accum) — UNVALIDATED for EP; naive-cast is dead (cos ≤0.67) | + +Cross-check: v1's per-EF rate ($19.7k) ≈ the empirical ceiling ($16.7k) — v1 rows were not +order-of-magnitude wrong, they were "un-optimized-run" priced. The row deltas people remember +($1k vs $12k) decompose as: ×2 tokens (20B vs 10B) × 3.2 EP × 1.4 buffer ≈ ×9. + +## Table (market rates, TF32 path, ×1.3 buffer; AWS on-demand ≈ +50–150%) + +| run | EP FLOPs (19.2ND) | cost | note | +|---|---|---|---| +| 300M × 6B (recipe validation) | 0.035 EF | **~$400** | | +| **1B × 20B Chinchilla EP** | 0.384 EF | **~$4k** | ≤$8k at the empirical-ceiling rate | +| 1B × 20B BP control | 0.12 EF (6ND) | **~$1.5k** | user's datapoint directly: 10B = $1k | +| **3B × 60B full-Chinchilla EP** | 3.46 EF | **~$30–40k** | inside $50k alone (market); AWS on-demand $45–60k → needs spot or bf16 | +| **7B × 20B (GPT-3-style) EP** | 2.69 EF | **~$25–35k** | inside envelope | +| 7B × 140B full-Chinchilla EP | 18.8 EF | $110–170k | out of scope; dedicated funding | + +## Consequences + +1. **$50k envelope reaches ONE of {3B full-Chinchilla, 7–8B reduced-token} after the 1B milestone** — + the scale conversation upgrades from speculative to budgetable. +2. The **calibration day (~$500 on Ben's AWS instances)** now resolves a 2× spread + (working rate vs empirical ceiling, AWS vs market) — worth more than before. +3. **bf16 mixed-precision validation is the highest-leverage engineering item**: ÷2 across the + table brings BOTH large options comfortably inside AWS on-demand pricing. Distinct from the + dead naive-cast: fp32 states + fp32 energy accumulation, bf16 matmuls only, gate = cos(EP,BPTT). +4. Throughput assumptions are conservative (eager, no compile/flash in the nudged phase); + the speed-profile levers (--holofast, --sdpa ≈ 1.5×) are not priced in. diff --git a/docs/hardware/COMPONENT_HW_MAP.md b/docs/hardware/COMPONENT_HW_MAP.md index 95c5689..ff47123 100644 --- a/docs/hardware/COMPONENT_HW_MAP.md +++ b/docs/hardware/COMPONENT_HW_MAP.md @@ -11,11 +11,11 @@ Companion docs: `HW_RESEARCH_FINDINGS.md` (softmax dossier, OLMo2 analog audit, | # | Component | Computes | Analog primitive | Reuse source (no tapeout) | Feedback path (nudged phase) | Tolerance status | |---|---|---|---|---|---|---| | 1 | Token embedding | table lookup | **digital** (SRAM lookup + DAC line drivers) | any MCU/FPGA + DAC array | none (input boundary) | n/a | -| 2 | RoPE | fixed per-position 2×2 rotations on q,k | **I/Q quadrature mixer**: cos/sin DDS + multiplying DACs | century-old RF practice; COTS mixer/DDS chips | transpose = rotation by −θ — **same mixer, negated sin** | phase error ≈ position blur — expect very tolerant; **E-tier item (queued)** | -| 3 | qkv / attn-proj / SwiGLU w1,w3,w2 / head | matrix–vector products | **crossbar arrays** (the bulk of all compute) | Demo-0/1: **SRAM-CIM (Shanbhag DIMA class)**; alternatives: Mythic flash CIM, IBM HERMES PCM eval | J^T reads = **bidirectional crossbar access** (the known PAR price; standard in every analog-EP proposal) | wq8 safe / wq6 marginal (looped-era static-tolerance); **re-run on cascade in E-tier** | -| 4 | QK-norm (full-width RMS) | q,k ← q/rms(q)·g | **divisive normalization**: square-law devices + KCL current sum + divider/AGC | same primitive class as softmax normalization periphery | Jacobian symmetric — self-transpose, no extra hardware | divider mismatch/offset — **E-tier item (queued)** | -| 5 | Softmax (causal) | exp + row-normalize | **Elfadel–Wyatt "entropic resistor"**: subthreshold-exp devices + KCL normalization (log-sum-exp co-content) | classic analog-VLSI circuit family; small arrays demonstrated | softmax Jacobian diag(p)−pp^T is symmetric; the QKV coupling around it is the non-reciprocal part | ±1% device mismatch → ±1% softmax error (linear); QK-norm bounds its input range. Dynamic-noise cliff (fnoise ≥1e-3, looped-era) **re-test on cascade = E-tier** | -| 6 | SwiGLU gate | silu(w1x) ⊙ (w3x) | sigmoid = differential pair (native); ⊙ = **Gilbert/translinear multiplier** per unit | textbook standard cells | transpose of ⊙ = multiply by the partner signal — cheap | per-multiply mismatch/noise, 1408 wide — **E-tier item (queued)** | +| 2 | RoPE | fixed per-position 2×2 rotations on q,k | **I/Q quadrature mixer**: cos/sin DDS + multiplying DACs | century-old RF practice; COTS mixer/DDS chips | transpose = rotation by −θ — **same mixer, negated sin** | **E-tier w1: 0.03 rad FREE; 0.1 rad ΔCE +0.012, cos_clean 0.92** — ~2° I/Q phase accuracy suffices (easy) | +| 3 | qkv / attn-proj / SwiGLU w1,w3,w2 / head | matrix–vector products | **crossbar arrays** (the bulk of all compute) | Demo-0/1: **SRAM-CIM (Shanbhag DIMA class)**; alternatives: Mythic flash CIM, IBM HERMES PCM eval | J^T reads = **bidirectional crossbar access** (the known PAR price; standard in every analog-EP proposal) | **E-tier w1 (cascade): wq8 FREE (ΔCE +0.004, cos_clean 0.96); wq6 MARGINAL (+0.051, 0.81); wq4 DEAD (+2.23)** — ≥7b effective weight precision is the binding spec; wave-2 lever = quant-aware co-training | +| 4 | QK-norm (full-width RMS) | q,k ← q/rms(q)·g | **divisive normalization**: square-law devices + KCL current sum + divider/AGC | same primitive class as softmax normalization periphery | Jacobian symmetric — self-transpose, no extra hardware | **E-tier w1: 3% mismatch FREE (ΔCE +0.003, cos_clean 0.95); 10% MARGINAL (+0.042, 0.78)** — 3% device matching is routine | +| 5 | Softmax (causal) | exp + row-normalize | **Elfadel–Wyatt "entropic resistor"**: subthreshold-exp devices + KCL normalization (log-sum-exp co-content) | classic analog-VLSI circuit family; small arrays demonstrated | softmax Jacobian diag(p)−pp^T is symmetric; the QKV coupling around it is the non-reciprocal part | ±1% device mismatch → ±1% softmax error (linear); QK-norm bounds its input range. **E-tier w1: forward additive state noise σ=1e-2 FREE (ΔCE +0.001, cos 0.97) — the looped-era 1e-3 cliff does NOT transfer to cascade** | +| 6 | SwiGLU gate | silu(w1x) ⊙ (w3x) | sigmoid = differential pair (native); ⊙ = **Gilbert/translinear multiplier** per unit | textbook standard cells | transpose of ⊙ = multiply by the partner signal — cheap | **E-tier w1: 10% gain error ΔCE +0.006, cos_clean 0.945 — FREE** at translinear-practice tolerances | | 7 | RMSNorm ×2 per block (norm-after-sublayer) | bound each sublayer output | divisive normalization (as #4), **no mean-subtraction path** (cheaper than LayerNorm) | same as #4 | symmetric Jacobian | as #4; ~4L+1 normalizers total = the largest NEW periphery count | | 8 | Residual stream | z + branch outputs | **current-summing bus (KCL node)** | wires | pass-through | norm-after bounds every injection — the arch change is itself the tolerance fix | | 9 | Final RMSNorm | bound pre-readout state | as #4 | as #4 | symmetric | bounds the **ADC** dynamic range at the digital boundary | -- cgit v1.2.3