summaryrefslogtreecommitdiff
path: root/EMAIL_BEN_DRAFT2.md
blob: a78c073e6426df198b7c3c3c2f637a3b3d424587 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
# Draft 2 — reply to Ben (costs + news), 2026-07-11

Subject: Re: costs — numbers you can put in front of accounting, and some news since the draft

Hi Ben,

Great to hear — and yes, AsymEP it is ;)

News first, because it changes what the credits would buy. Since the draft, the scaling line moved to a cascade form: L *distinct* standard transformer blocks, inference is a plain forward pass, EP only in training.

- **A standard 12-block transformer (42.7M) trained through a full epoch with zero backpropagation anywhere in the loop — and it tells coherent stories.** A one-click notebook that samples from it is ready whenever you'd like to play with it.
- As far as we can verify, **it is already the largest fully backprop-free language model trained to date**: the 1.07B optical-DFA result trains 400M of its parameters and keeps BP at the readout; the ES line's from-scratch model is a ~2M integer GRU.
- **At matched tuning, EP is indistinguishable from its tuned backprop twin** (4k steps); at full-epoch the gap is 0.050 nats — and every centinat of it has a named mechanism and a measured dial (estimator SNR vs. nudge amplitude: the same physics an analog chip will meet, rehearsed in fp32).
- **Muon's advantage over AdamW transfers to EP gradients at full magnitude** (−0.13 CE at this scale) — backprop-free training runs on a modern optimizer stack.

On costs, anchored on measured throughput (TF32 validated bit-exact for this trainer; EP = 3.2× BP FLOPs, measured, not estimated):

| run | cost, 1.4× failure buffer included |
|---|---|
| 300M × 6B tok — recipe validation on a real corpus | ~$1k |
| **1B × 20B tok (Chinchilla-optimal)** | **~$12k — inside the cap** |
| 3B × 25B tok | ~$25–30k (two capped, checkpointed segments) |

The buffer is small because the failure model is not a guess: the three instabilities we met this quarter are each mechanistically diagnosed and closed (a β-schedule, a block-architecture change, an amplitude floor), and runs checkpoint every 5k steps — **a failed run costs a segment, not the run.** One day on your instances (~$500) converts these numbers from ±30% to ±10–15%; happy to run that calibration first.

Two questions so we can plan honestly: is the ~$20k a per-run cap inside a larger envelope, or the envelope itself? And does the backprop-twin control (≈1/3 the EP cost, needed for the head-to-head claim) draw from the same pool?

No rush on the draft — it has since been restructured around a control-map spine, so I'd send you the updated PDF before you invest reading time.

Best,
Yuren