summaryrefslogtreecommitdiff
path: root/docs
diff options
context:
space:
mode:
authorYuren Hao <yurenh2@illinois.edu>2026-07-12 05:44:18 -0500
committerYuren Hao <yurenh2@illinois.edu>2026-07-12 05:44:18 -0500
commit411870d64f00181ea41a8059f4b53f5b232447c0 (patch)
tree391b22448b2c2398e7857e3c2af39abc9cd037c2 /docs
parent155478b7ffdd9015fde13884b1ebff9b92148a20 (diff)
E-tier wave-1 ledger (RESULT 14 + HW map rows) + cost model v2 recalibrated on measured H200 datapoint
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
Diffstat (limited to 'docs')
-rw-r--r--docs/campaign/CASCADE_ABLATION_PLAN.md24
-rw-r--r--docs/campaign/COST_MODEL.md47
-rw-r--r--docs/hardware/COMPONENT_HW_MAP.md10
3 files changed, 76 insertions, 5 deletions
diff --git a/docs/campaign/CASCADE_ABLATION_PLAN.md b/docs/campaign/CASCADE_ABLATION_PLAN.md
index 61860f6..8b52369 100644
--- a/docs/campaign/CASCADE_ABLATION_PLAN.md
+++ b/docs/campaign/CASCADE_ABLATION_PLAN.md
@@ -538,3 +538,27 @@ s1 alone beat the BP mean). Ledger of the residual-0.05 epoch gap after three pr
Carry `--bsign_rand` and a future centered mode as Stage-2 A/B flags; the gap question re-opens at
300M/real-corpus where it means something. Effort pivots to: (1) Stage-2 data pipeline (FineWeb-Edu +
32k tokenizer), (2) E-tier tolerance suite on the idle farm (hardware track / UIUC outreach feed).
+
+### RESULT 14 (2026-07-12): E-TIER WAVE-1 — full analog-fault tolerance ledger at stage1b s55000.
+`etier_probe.py`, farm GPUs 2/3/7 (shards A/B/C), stage1b_ep_muon_s55000.pt (clean valCE 1.2678),
+B=8 eval batches; metrics = faulted valCE (Δ vs clean), cos(EP_faulted, BP_faulted) [self-consistency
+of the learning signal under fault], cos(EP_faulted, BP_clean) [direction vs the ideal update].
+
+| fault (component) | mild | medium | severe | verdict |
+|---|---|---|---|---|
+| wq — weight quant (crossbar #3) | 8b: +0.004 / 0.956 | 6b: +0.051 / 0.812 | 4b: +2.23 / 0.05 | **8b FREE, 6b marginal, 4b dead → ≥7b effective is the binding spec** |
+| fnoise — fwd additive state noise (softmax/relax #5) | 1e-3: 0.000 / 0.975 | 3e-3: 0.000 / 0.974 | 1e-2: +0.001 / 0.970 | **FREE at 1% — looped-era 1e-3 cliff does NOT transfer to cascade** |
+| divmis — divisive-norm mismatch (#4/#7) | 1%: 0.000 / 0.971 | 3%: +0.003 / 0.948 | 10%: +0.042 / 0.785 | 3% (routine matching) FREE; 10% marginal |
+| rope — phase error rad (#2) | 0.01: 0.000 / 0.973 | 0.03: +0.001 / 0.967 | 0.1: +0.012 / 0.923 | 0.03 rad FREE; ~2° I/Q accuracy suffices |
+| gilbert — gate gain error (#6) | 1%: 0.000 / 0.974 | 3%: 0.000 / 0.971 | 10%: +0.006 / 0.945 | **FREE at 10%** — translinear practice is comfortably inside |
+| fbnoise — nudge/error-channel noise | 1e-2: 0.969 / 0.975 | **1e-1: 0.951 / 0.957** | 3e-1: 0.764 / 0.768 | **10% relative noise on the ERROR CHANNEL is FREE** (cos 0.95) — the r-indifference/large-nudge gift, now measured on cascade |
+
+Reading: (a) the only hard constraint is crossbar weight precision (≥7b effective — inside standard
+SRAM-CIM capability; 6b rescue = wave-2 quant-aware co-training); (b) everything dynamic — forward
+noise 1%, error-channel noise 10%, gate/divider/phase mismatch at routine device tolerances — is
+FREE at this scale. cos(EP,BP_faulted) stays ~0.97 under every non-fatal fault: the EP estimator
+tracks whatever network the faults define, i.e. learning co-adapts to the fault (the analog-training
+thesis in one number). CAVEAT: static probes at a trained checkpoint (eval CE + one-step gradient
+direction), not training-under-fault; wave-2 = co-training with faults injected from step 0
+(expectation from the literature and from (c): tolerances IMPROVE). Feeds COMPONENT_HW_MAP.md
+(per-row status updated) + UIUC outreach dossier.
diff --git a/docs/campaign/COST_MODEL.md b/docs/campaign/COST_MODEL.md
new file mode 100644
index 0000000..2b9d8b6
--- /dev/null
+++ b/docs/campaign/COST_MODEL.md
@@ -0,0 +1,47 @@
+# Cost model v2 (2026-07-12) — recalibrated on a measured rental datapoint
+
+Supersedes the v1 table in EMAIL_BEN_DRAFT2.md (drafts ≤5) and the numbers quoted in the
+2026-07-11 scale discussion. **v1's error: anchored $/FLOP on AWS-A100-savings TF32
+(~$19.7k/EF) — the wrong reference class for 2026.** Market H100/H200 rentals are 2.5–3×
+cheaper per FLOP.
+
+## Anchors
+
+| anchor | value | source |
+|---|---|---|
+| **Empirical all-in ceiling** | **$16.7k/EF** | user's real run: 1B × 10B tok BP on rented H200s ≈ $1k all-in; 6ND = 0.06 EF. Implied throughput only 42–58 TF/s eff (4–6% of bf16 peak) — an UN-optimized run, so this is a ceiling, not a target |
+| Market rental rate | $2–3.5/GPU·h | H100/H200 marketplace (Vast/RunPod class), 2026 |
+| Our stack's utilization | ~19% of TF32 peak | measured on A6000 (14 TF eff / 75 TF TF32 peak), eager TF32 |
+| → projected H200 throughput | 75–125 TF/s eff | 15–25% × 494 TF TF32 peak |
+| → **working rate (TF32 path)** | **$6–9k/EF** (mid $7.5k) | $2.5–3/h ÷ 75–125 TF/s |
+| EP/BP FLOP ratio | **3.2×** (measured) | EP ≈ 19.2N FLOPs/tok vs BP 6N; single-sided, K=3 |
+| Failure allowance | 1.3× | failure modes characterized; ckpt every 5k steps |
+| AWS multiplier | 1.5–2.5× on-demand; 1.2–1.5× spot/capacity-block | vs marketplace $/GPU·h |
+| Mixed-precision upside | ÷~2 | proper autocast (bf16 matmul, fp32 states + E-accum) — UNVALIDATED for EP; naive-cast is dead (cos ≤0.67) |
+
+Cross-check: v1's per-EF rate ($19.7k) ≈ the empirical ceiling ($16.7k) — v1 rows were not
+order-of-magnitude wrong, they were "un-optimized-run" priced. The row deltas people remember
+($1k vs $12k) decompose as: ×2 tokens (20B vs 10B) × 3.2 EP × 1.4 buffer ≈ ×9.
+
+## Table (market rates, TF32 path, ×1.3 buffer; AWS on-demand ≈ +50–150%)
+
+| run | EP FLOPs (19.2ND) | cost | note |
+|---|---|---|---|
+| 300M × 6B (recipe validation) | 0.035 EF | **~$400** | |
+| **1B × 20B Chinchilla EP** | 0.384 EF | **~$4k** | ≤$8k at the empirical-ceiling rate |
+| 1B × 20B BP control | 0.12 EF (6ND) | **~$1.5k** | user's datapoint directly: 10B = $1k |
+| **3B × 60B full-Chinchilla EP** | 3.46 EF | **~$30–40k** | inside $50k alone (market); AWS on-demand $45–60k → needs spot or bf16 |
+| **7B × 20B (GPT-3-style) EP** | 2.69 EF | **~$25–35k** | inside envelope |
+| 7B × 140B full-Chinchilla EP | 18.8 EF | $110–170k | out of scope; dedicated funding |
+
+## Consequences
+
+1. **$50k envelope reaches ONE of {3B full-Chinchilla, 7–8B reduced-token} after the 1B milestone** —
+ the scale conversation upgrades from speculative to budgetable.
+2. The **calibration day (~$500 on Ben's AWS instances)** now resolves a 2× spread
+ (working rate vs empirical ceiling, AWS vs market) — worth more than before.
+3. **bf16 mixed-precision validation is the highest-leverage engineering item**: ÷2 across the
+ table brings BOTH large options comfortably inside AWS on-demand pricing. Distinct from the
+ dead naive-cast: fp32 states + fp32 energy accumulation, bf16 matmuls only, gate = cos(EP,BPTT).
+4. Throughput assumptions are conservative (eager, no compile/flash in the nudged phase);
+ the speed-profile levers (--holofast, --sdpa ≈ 1.5×) are not priced in.
diff --git a/docs/hardware/COMPONENT_HW_MAP.md b/docs/hardware/COMPONENT_HW_MAP.md
index 95c5689..ff47123 100644
--- a/docs/hardware/COMPONENT_HW_MAP.md
+++ b/docs/hardware/COMPONENT_HW_MAP.md
@@ -11,11 +11,11 @@ Companion docs: `HW_RESEARCH_FINDINGS.md` (softmax dossier, OLMo2 analog audit,
| # | Component | Computes | Analog primitive | Reuse source (no tapeout) | Feedback path (nudged phase) | Tolerance status |
|---|---|---|---|---|---|---|
| 1 | Token embedding | table lookup | **digital** (SRAM lookup + DAC line drivers) | any MCU/FPGA + DAC array | none (input boundary) | n/a |
-| 2 | RoPE | fixed per-position 2×2 rotations on q,k | **I/Q quadrature mixer**: cos/sin DDS + multiplying DACs | century-old RF practice; COTS mixer/DDS chips | transpose = rotation by −θ — **same mixer, negated sin** | phase error ≈ position blur — expect very tolerant; **E-tier item (queued)** |
-| 3 | qkv / attn-proj / SwiGLU w1,w3,w2 / head | matrix–vector products | **crossbar arrays** (the bulk of all compute) | Demo-0/1: **SRAM-CIM (Shanbhag DIMA class)**; alternatives: Mythic flash CIM, IBM HERMES PCM eval | J^T reads = **bidirectional crossbar access** (the known PAR price; standard in every analog-EP proposal) | wq8 safe / wq6 marginal (looped-era static-tolerance); **re-run on cascade in E-tier** |
-| 4 | QK-norm (full-width RMS) | q,k ← q/rms(q)·g | **divisive normalization**: square-law devices + KCL current sum + divider/AGC | same primitive class as softmax normalization periphery | Jacobian symmetric — self-transpose, no extra hardware | divider mismatch/offset — **E-tier item (queued)** |
-| 5 | Softmax (causal) | exp + row-normalize | **Elfadel–Wyatt "entropic resistor"**: subthreshold-exp devices + KCL normalization (log-sum-exp co-content) | classic analog-VLSI circuit family; small arrays demonstrated | softmax Jacobian diag(p)−pp^T is symmetric; the QKV coupling around it is the non-reciprocal part | ±1% device mismatch → ±1% softmax error (linear); QK-norm bounds its input range. Dynamic-noise cliff (fnoise ≥1e-3, looped-era) **re-test on cascade = E-tier** |
-| 6 | SwiGLU gate | silu(w1x) ⊙ (w3x) | sigmoid = differential pair (native); ⊙ = **Gilbert/translinear multiplier** per unit | textbook standard cells | transpose of ⊙ = multiply by the partner signal — cheap | per-multiply mismatch/noise, 1408 wide — **E-tier item (queued)** |
+| 2 | RoPE | fixed per-position 2×2 rotations on q,k | **I/Q quadrature mixer**: cos/sin DDS + multiplying DACs | century-old RF practice; COTS mixer/DDS chips | transpose = rotation by −θ — **same mixer, negated sin** | **E-tier w1: 0.03 rad FREE; 0.1 rad ΔCE +0.012, cos_clean 0.92** — ~2° I/Q phase accuracy suffices (easy) |
+| 3 | qkv / attn-proj / SwiGLU w1,w3,w2 / head | matrix–vector products | **crossbar arrays** (the bulk of all compute) | Demo-0/1: **SRAM-CIM (Shanbhag DIMA class)**; alternatives: Mythic flash CIM, IBM HERMES PCM eval | J^T reads = **bidirectional crossbar access** (the known PAR price; standard in every analog-EP proposal) | **E-tier w1 (cascade): wq8 FREE (ΔCE +0.004, cos_clean 0.96); wq6 MARGINAL (+0.051, 0.81); wq4 DEAD (+2.23)** — ≥7b effective weight precision is the binding spec; wave-2 lever = quant-aware co-training |
+| 4 | QK-norm (full-width RMS) | q,k ← q/rms(q)·g | **divisive normalization**: square-law devices + KCL current sum + divider/AGC | same primitive class as softmax normalization periphery | Jacobian symmetric — self-transpose, no extra hardware | **E-tier w1: 3% mismatch FREE (ΔCE +0.003, cos_clean 0.95); 10% MARGINAL (+0.042, 0.78)** — 3% device matching is routine |
+| 5 | Softmax (causal) | exp + row-normalize | **Elfadel–Wyatt "entropic resistor"**: subthreshold-exp devices + KCL normalization (log-sum-exp co-content) | classic analog-VLSI circuit family; small arrays demonstrated | softmax Jacobian diag(p)−pp^T is symmetric; the QKV coupling around it is the non-reciprocal part | ±1% device mismatch → ±1% softmax error (linear); QK-norm bounds its input range. **E-tier w1: forward additive state noise σ=1e-2 FREE (ΔCE +0.001, cos 0.97) — the looped-era 1e-3 cliff does NOT transfer to cascade** |
+| 6 | SwiGLU gate | silu(w1x) ⊙ (w3x) | sigmoid = differential pair (native); ⊙ = **Gilbert/translinear multiplier** per unit | textbook standard cells | transpose of ⊙ = multiply by the partner signal — cheap | **E-tier w1: 10% gain error ΔCE +0.006, cos_clean 0.945 — FREE** at translinear-practice tolerances |
| 7 | RMSNorm ×2 per block (norm-after-sublayer) | bound each sublayer output | divisive normalization (as #4), **no mean-subtraction path** (cheaper than LayerNorm) | same as #4 | symmetric Jacobian | as #4; ~4L+1 normalizers total = the largest NEW periphery count |
| 8 | Residual stream | z + branch outputs | **current-summing bus (KCL node)** | wires | pass-through | norm-after bounds every injection — the arch change is itself the tolerance fix |
| 9 | Final RMSNorm | bound pre-readout state | as #4 | as #4 | symmetric | bounds the **ADC** dynamic range at the digital boundary |