summaryrefslogtreecommitdiff
path: root/EMAIL_BEN_DRAFT2.md
diff options
context:
space:
mode:
authorYuren Hao <yurenh2@illinois.edu>2026-07-11 09:27:15 -0500
committerYuren Hao <yurenh2@illinois.edu>2026-07-11 09:27:15 -0500
commit6f4ac86e8e1774cd4c79aabcf71af3eae2c6225d (patch)
tree68321b3e04784bffa3783472b8681b00d6397a10 /EMAIL_BEN_DRAFT2.md
parent635f5a3a974a10dd1b84487f60f2b9a1447967fd (diff)
Ben draft 4: academic register (metaphors stripped per user style correction); register-split rule recorded in style profile
Diffstat (limited to 'EMAIL_BEN_DRAFT2.md')
-rw-r--r--EMAIL_BEN_DRAFT2.md30
1 files changed, 15 insertions, 15 deletions
diff --git a/EMAIL_BEN_DRAFT2.md b/EMAIL_BEN_DRAFT2.md
index 9558ec2..de63afd 100644
--- a/EMAIL_BEN_DRAFT2.md
+++ b/EMAIL_BEN_DRAFT2.md
@@ -1,33 +1,33 @@
-# Draft 3 — reply to Ben (costs + news), 2026-07-11
+# Draft 4 — reply to Ben (academic register), 2026-07-11
-Subject: Re: costs — numbers for accounting, and some news since the draft
+Subject: Re: costs — updated numbers, and some results since the draft
Hi Ben,
Great to hear — and yes, AsymEP it is ;)
-News first, because it changes what the credits buy. Since the draft:
+First, results obtained since the draft, since they bear on what the credits would fund:
-- **A standard 12-layer transformer trained through a full epoch with zero backpropagation anywhere in the training loop — and it tells coherent stories.** Inference is a plain forward pass; a one-click notebook that samples from it is ready whenever you'd like to play.
-- As far as we can verify, **it is already the largest neural network fully trained without backpropagation** — the 1.07B optical-DFA result trains 400M of its parameters and keeps BP at the readout; the ES line's billion-scale results are fine-tuning of BP-pretrained models, its from-scratch model a ~2M integer GRU.
-- **Every previous no-backprop attempt, at any scale, concedes a quality gap to backprop. We measure none.** At matched tuning EP is statistically indistinguishable from its tuned backprop twin; at full-epoch the residual is 0.050 nats — and every centinat of it has a named mechanism and a measured dial (estimator SNR vs. nudge amplitude: the same physics an analog chip will meet, rehearsed in fp32). Muon's advantage over AdamW transfers to EP gradients at full magnitude — this runs on a modern optimizer stack.
+- **A standard 12-layer transformer was trained for a full epoch with no backpropagation at any point in the training loop, and it generates coherent text.** Inference is a standard forward pass. A notebook that samples from the trained model is ready if you would like to try it.
+- To the best of our knowledge, **this is the largest neural network fully trained without backpropagation to date.** The closest claims do not hold under inspection: the 1.07B optical-DFA model trains only 400M of its parameters and retains backpropagation at the readout layer, and the billion-scale evolution-strategies results are fine-tuning of backprop-pretrained models (their from-scratch model is a ~2M-parameter integer GRU).
+- **Previous backpropagation-free methods report a quality gap relative to backpropagation at every scale tested. We observe none.** At matched hyperparameters, EP training is statistically indistinguishable from its backpropagation control at 4k steps; over a full epoch the difference is 0.050 nats, and each contribution to that residual has an identified mechanism and a demonstrated mitigation (estimator signal-to-noise as a function of nudge amplitude — the same quantity that will govern analog implementations). Muon's improvement over AdamW transfers to EP gradients at full magnitude, so the method is compatible with current optimizer practice.
-On costs, anchored on measured throughput (TF32 validated bit-exact; EP = 3.2× BP FLOPs, measured):
+On costs, based on measured throughput (TF32 verified bit-exact for this trainer; EP requires 3.2× the FLOPs of backpropagation, measured rather than estimated):
-| run | cost, 1.4× failure buffer included |
+| run | cost, including a 1.4× failure allowance |
|---|---|
-| 300M × 6B tok — recipe validation on a real corpus | ~$1k |
-| **1B × 20B tok, Chinchilla-optimal** | **~$12k — inside the cap** |
+| 300M × 6B tokens — recipe validation on a standard corpus | ~$1k |
+| **1B × 20B tokens, Chinchilla-optimal** | **~$12k — within the cap** |
-The buffer is small because the failure model is not a guess: the three instabilities we met this quarter are each diagnosed and mechanistically closed, and runs checkpoint every 5k steps — **a failed run costs a segment, not the run.** One day on your instances (~$500) converts ±30% into ±10–15%; happy to run that calibration first.
+The failure allowance is modest because the failure modes are characterized rather than assumed: the three instabilities encountered this quarter were each diagnosed and resolved (a nudge-amplitude schedule, an architecture choice, an amplitude floor), and training checkpoints every 5k steps, so a failure loses only the segment since the last checkpoint. A one-day calibration on your instances (~$500) would reduce the uncertainty of these estimates from ±30% to ±10–15%; we would be glad to run that first.
-The honest question on our minds is the next octave: **is 3B at full Chinchilla thinkable, and is 7–8B thinkable at all?** (~$25–30k and ~$50k at today's measured throughput — beyond a single capped run.) We are not asking for that now. The proposal is to deliver everything above within the cap first, and let those results decide whether the bigger conversation is worth having.
+The question we are weighing is the next scale step: **whether 3B at full Chinchilla, or 7–8B at reduced token count, is feasible** (~$25–30k and ~$50k respectively at current measured throughput — beyond a single capped run). We are not requesting this now; our intention is to complete the runs above within the cap, and to revisit the larger scale on the basis of those results.
-Separately — and this may interest Rain directly — **we are planning a small analog-hardware LM demo, and the recipe was deliberately designed so that every trained operation has an analog implementation**: divisive (current-mode) normalizations, fixed I/Q rotations for position, translinear gates, the classic subthreshold-exp/KCL circuit for attention's softmax, and learning itself as EP's two local phases. A component-by-component hardware map exists; happy to share it.
+Separately, and possibly of direct interest to Rain: **we are planning a small analog-hardware LM demonstration, and the training recipe was designed so that every trained operation has a known analog implementation** — divisive (current-mode) normalization, fixed rotations for the position encoding, translinear multipliers for the gated MLP, the subthreshold-exponential/KCL circuit for attention's softmax, and EP's two relaxation phases for the learning rule itself. A component-by-component hardware mapping is written up; we can share it.
-Two small accounting questions: is the ~$20k a per-run cap inside a larger envelope, or the envelope itself? And does the backprop-twin control (≈1/3 the EP cost, needed for the head-to-head claim) draw from the same pool?
+Two accounting questions: is the ~$20k a per-run cap within a larger total, or the total itself? And would a backpropagation control run (about one third of the EP cost, required for the comparison) draw from the same allocation?
-On the draft — it has since been restructured, so I'd send you the updated PDF before you invest reading time. And as you read, one strategic question we would genuinely value your judgment on. The dynamics story in your hands and the scaling demonstration above are separable: the report stands on its own, and only a slice of it is load-bearing for the LM line. We currently lean toward two papers — the dynamics and its control map now, and the scale demonstration later as a capstone aimed at **the broadest venue the result can carry**, with the dynamics paper as its mechanistic backbone. Merged into one, though, it is a different kind of artifact. You have better instincts than we do about which package travels further.
+On the draft: it has been substantially restructured since the version you have, so we would send the updated PDF before you invest reading time. One question we would value your judgment on as you read: the dynamics results and the scaling demonstration are separable, and only part of the former is required for the latter. Our current inclination is two papers — the dynamics and its taxonomy of stability controls now, and the scale demonstration later, submitted to the most general venue the result supports, with the dynamics paper as its theoretical reference. A single merged paper is the alternative. Your judgment on which is the stronger structure would be valuable.
Best,
Yuren