summaryrefslogtreecommitdiff
path: root/docs/campaign/COST_MODEL.md
blob: 2b9d8b660851bc1d0195e1af1536b83f610c3974 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
# Cost model v2 (2026-07-12) — recalibrated on a measured rental datapoint

Supersedes the v1 table in EMAIL_BEN_DRAFT2.md (drafts ≤5) and the numbers quoted in the
2026-07-11 scale discussion. **v1's error: anchored $/FLOP on AWS-A100-savings TF32
(~$19.7k/EF) — the wrong reference class for 2026.** Market H100/H200 rentals are 2.5–3×
cheaper per FLOP.

## Anchors

| anchor | value | source |
|---|---|---|
| **Empirical all-in ceiling** | **$16.7k/EF** | user's real run: 1B × 10B tok BP on rented H200s ≈ $1k all-in; 6ND = 0.06 EF. Implied throughput only 42–58 TF/s eff (4–6% of bf16 peak) — an UN-optimized run, so this is a ceiling, not a target |
| Market rental rate | $2–3.5/GPU·h | H100/H200 marketplace (Vast/RunPod class), 2026 |
| Our stack's utilization | ~19% of TF32 peak | measured on A6000 (14 TF eff / 75 TF TF32 peak), eager TF32 |
| → projected H200 throughput | 75–125 TF/s eff | 15–25% × 494 TF TF32 peak |
| → **working rate (TF32 path)** | **$6–9k/EF** (mid $7.5k) | $2.5–3/h ÷ 75–125 TF/s |
| EP/BP FLOP ratio | **3.2×** (measured) | EP ≈ 19.2N FLOPs/tok vs BP 6N; single-sided, K=3 |
| Failure allowance | 1.3× | failure modes characterized; ckpt every 5k steps |
| AWS multiplier | 1.5–2.5× on-demand; 1.2–1.5× spot/capacity-block | vs marketplace $/GPU·h |
| Mixed-precision upside | ÷~2 | proper autocast (bf16 matmul, fp32 states + E-accum) — UNVALIDATED for EP; naive-cast is dead (cos ≤0.67) |

Cross-check: v1's per-EF rate ($19.7k) ≈ the empirical ceiling ($16.7k) — v1 rows were not
order-of-magnitude wrong, they were "un-optimized-run" priced. The row deltas people remember
($1k vs $12k) decompose as: ×2 tokens (20B vs 10B) × 3.2 EP × 1.4 buffer ≈ ×9.

## Table (market rates, TF32 path, ×1.3 buffer; AWS on-demand ≈ +50–150%)

| run | EP FLOPs (19.2ND) | cost | note |
|---|---|---|---|
| 300M × 6B (recipe validation) | 0.035 EF | **~$400** | |
| **1B × 20B Chinchilla EP** | 0.384 EF | **~$4k** | ≤$8k at the empirical-ceiling rate |
| 1B × 20B BP control | 0.12 EF (6ND) | **~$1.5k** | user's datapoint directly: 10B = $1k |
| **3B × 60B full-Chinchilla EP** | 3.46 EF | **~$30–40k** | inside $50k alone (market); AWS on-demand $45–60k → needs spot or bf16 |
| **7B × 20B (GPT-3-style) EP** | 2.69 EF | **~$25–35k** | inside envelope |
| 7B × 140B full-Chinchilla EP | 18.8 EF | $110–170k | out of scope; dedicated funding |

## Consequences

1. **$50k envelope reaches ONE of {3B full-Chinchilla, 7–8B reduced-token} after the 1B milestone** —
   the scale conversation upgrades from speculative to budgetable.
2. The **calibration day (~$500 on Ben's AWS instances)** now resolves a 2× spread
   (working rate vs empirical ceiling, AWS vs market) — worth more than before.
3. **bf16 mixed-precision validation is the highest-leverage engineering item**: ÷2 across the
   table brings BOTH large options comfortably inside AWS on-demand pricing. Distinct from the
   dead naive-cast: fp32 states + fp32 energy accumulation, bf16 matmuls only, gate = cos(EP,BPTT).
4. Throughput assumptions are conservative (eager, no compile/flash in the nudged phase);
   the speed-profile levers (--holofast, --sdpa ≈ 1.5×) are not priced in.