# Cost model v2 (2026-07-12) — recalibrated on a measured rental datapoint Supersedes the v1 table in EMAIL_BEN_DRAFT2.md (drafts ≤5) and the numbers quoted in the 2026-07-11 scale discussion. **v1's error: anchored $/FLOP on AWS-A100-savings TF32 (~$19.7k/EF) — the wrong reference class for 2026.** Market H100/H200 rentals are 2.5–3× cheaper per FLOP. ## Anchors | anchor | value | source | |---|---|---| | **Empirical all-in ceiling** | **$16.7k/EF** | user's real run: 1B × 10B tok BP on rented H200s ≈ $1k all-in; 6ND = 0.06 EF. Implied throughput only 42–58 TF/s eff (4–6% of bf16 peak) — an UN-optimized run, so this is a ceiling, not a target | | Market rental rate | $2–3.5/GPU·h | H100/H200 marketplace (Vast/RunPod class), 2026 | | Our stack's utilization | ~19% of TF32 peak | measured on A6000 (14 TF eff / 75 TF TF32 peak), eager TF32 | | → projected H200 throughput | 75–125 TF/s eff | 15–25% × 494 TF TF32 peak | | → **working rate (TF32 path)** | **$6–9k/EF** (mid $7.5k) | $2.5–3/h ÷ 75–125 TF/s | | EP/BP FLOP ratio | **3.2×** (measured) | EP ≈ 19.2N FLOPs/tok vs BP 6N; single-sided, K=3 | | Failure allowance | 1.3× | failure modes characterized; ckpt every 5k steps | | AWS multiplier | 1.5–2.5× on-demand; 1.2–1.5× spot/capacity-block | vs marketplace $/GPU·h | | Mixed-precision upside | ÷~2 | proper autocast (bf16 matmul, fp32 states + E-accum) — UNVALIDATED for EP; naive-cast is dead (cos ≤0.67) | Cross-check: v1's per-EF rate ($19.7k) ≈ the empirical ceiling ($16.7k) — v1 rows were not order-of-magnitude wrong, they were "un-optimized-run" priced. The row deltas people remember ($1k vs $12k) decompose as: ×2 tokens (20B vs 10B) × 3.2 EP × 1.4 buffer ≈ ×9. ## Table (market rates, TF32 path, ×1.3 buffer; AWS on-demand ≈ +50–150%) | run | EP FLOPs (19.2ND) | cost | note | |---|---|---|---| | 300M × 6B (recipe validation) | 0.035 EF | **~$400** | | | **1B × 20B Chinchilla EP** | 0.384 EF | **~$4k** | ≤$8k at the empirical-ceiling rate | | 1B × 20B BP control | 0.12 EF (6ND) | **~$1.5k** | user's datapoint directly: 10B = $1k | | **3B × 60B full-Chinchilla EP** | 3.46 EF | **~$30–40k** | inside $50k alone (market); AWS on-demand $45–60k → needs spot or bf16 | | **7B × 20B (GPT-3-style) EP** | 2.69 EF | **~$25–35k** | inside envelope | | 7B × 140B full-Chinchilla EP | 18.8 EF | $110–170k | out of scope; dedicated funding | ## Consequences 1. **$50k envelope reaches ONE of {3B full-Chinchilla, 7–8B reduced-token} after the 1B milestone** — the scale conversation upgrades from speculative to budgetable. 2. The **calibration day (~$500 on Ben's AWS instances)** now resolves a 2× spread (working rate vs empirical ceiling, AWS vs market) — worth more than before. 3. **bf16 mixed-precision validation is the highest-leverage engineering item**: ÷2 across the table brings BOTH large options comfortably inside AWS on-demand pricing. Distinct from the dead naive-cast: fp32 states + fp32 energy accumulation, bf16 matmuls only, gate = cos(EP,BPTT). 4. Throughput assumptions are conservative (eager, no compile/flash in the nudged phase); the speed-profile levers (--holofast, --sdpa ≈ 1.5×) are not priced in.