# Draft 6 — follow-up: Overleaf link + formulation correction + note (2026-07-12) Subject: Re: Overleaf link + the numbers Hi Ben, Here is the Overleaf: [LINK] — comment rights enabled. It is the restructured version: the paper is now organized around a control map (what each protection clamps × when its signal acts), the two systems of yesterday's email are kept explicit throughout (the looped transformer is the paper's non-conservative subject; the scaling line is the layered-energy route, and your ImageNet paper is cited as its closest relative), and the experiments your review asked for are in (the jacreg head-to-head, the no-drive collapse arms, the configuration appendix). First, your correction — you are right, and I withdraw the shortcut. To your direct question: **we use the energy formulation.** Our trainer defines E of your Eq. (10), relaxes the nudged stationarity with a fixed-point scheme of the same family as your mod-PGD, and reads dE/dtheta at the nudged state, single-sided (legal since the free-state read vanishes, as in your Eq. (11)). My "AsymEP reduces to vanilla EP" remark was, as you say, contentless for that case. Your sharper question — whether AsymEP on the directly-defined field F_k = f_k(h_{k-1}) − h_k yields the same algorithm — turned out to have a more interesting answer than "yes": **same gradient, different algorithms.** Composing your group's exactness theorem (arXiv:2602.03670, Eqs. 22–25) with the classical EP theorem gives equality of the gradients in the small-β limit; but the two nudged equilibria differ at first order — the energy formulation displaces by −β[(I−L)ᵀ(I−L)]⁻¹∇C (the Gauss–Newton curvature survives exactly because the free-point errors vanish), the corrected field by −β(I−Lᵀ)⁻¹∇C (your Eq. 24) — and the readout operators and bias structures differ correspondingly. The attached one-page note states this precisely, with a numerical check (predicted resolvents match measured displacements to cos = 1.000000; the two states differ at cos 0.93, norm ratio 1.8; both reads converge to autograd). Consistently, your paper's observation that uncorrected VF trains only the last layer on feedforward structure is the degenerate case: without the correction the nudged equilibrium moves no interior layer. If the note survives your reading, it may make a useful appendix for either paper. And quite right on scope — the report's ensemble evidence is MLP/CNN/RNN at 200 seeds; the looped transformer is its in-vivo instance. The results promised yesterday, concretely: - **A standard 12-layer transformer (42.7M) trained for a full epoch — 59k steps, 361M tokens — with no backpropagation at any point in the training loop; it generates coherent text.** Inference is a standard forward pass. To our knowledge the first transformer language model fully trained this way (the 1.07B optical-DFA model trains 400M of its parameters and keeps BP at the readout; billion-scale ES results are fine-tuning). A notebook that samples from the trained model is attached — one click, no setup. - **The gap ledger.** At matched hyperparameters, 4k steps: EP statistically indistinguishable from its tuned backprop control (three seeds each). Full epoch: 0.050 nats behind, with the mechanism identified — the late-training decline of estimator signal-to-noise as the true gradient shrinks against a fixed precision floor (the quantified version of your β remark) — and a β-schedule that removes most of it. Muon's improvement over AdamW transfers to EP gradients at full magnitude. - **Solver cost.** Our nudged phase runs the full EP step at 3.2× the FLOPs of a backprop step (measured; single-sided, K=3 full-depth sweeps per nudge). For calibration against your Appendix B: that ratio is what turns an 18-day run into a ~4-day one at equal hardware. On costs — calibrated against a measured datapoint rather than list pricing (a 1B × 10B-token backprop run we executed on rented H200s cost ≈$1k all-in), combined with the measured 3.2× EP/BP FLOP ratio: | run | cost at market H100/H200 rates, incl. 1.3× failure allowance | |---|---| | 300M × 6B tokens — recipe validation on a standard corpus | ~$400 | | **1B × 20B tokens, Chinchilla-optimal** | **~$4k** (≤$8k even at the uncalibrated worst case) | | 1B backprop control, same tokens | ~$1.5k | Three qualifications. The numbers assume the TF32 path we have verified bit-exact against fp32; a standard mixed-precision variant (bf16 matrix products, fp32 states and energy accumulation) is not yet validated for EP and would roughly halve them if it passes our gradient-fidelity gate. AWS list pricing runs 1.5–2.5× market rental per GPU-hour, with spot or capacity-block reservations closing most of that gap. And checkpoints every 5k steps bound any failure to the segment since the last one. A one-day calibration on your instances (~$500) would replace these estimates with measurements; we would be glad to run that first. This pricing changes the scale question. Within $50k, the 1B milestone and its control leave room for one of: **3B at full Chinchilla scaling (~$30–40k at market rates), or 7–8B at a GPT-3-style token budget (~$25–35k)** — AWS on-demand pricing adds roughly half on top, and the mixed-precision variant, if it validates, subtracts it back. Our intention is unchanged — complete the 1B result first and let it argue for the larger run — but the envelope does reach one of the two. Separately, and possibly of direct interest to Rain: your paper leaves open whether F_PCN can be realized efficiently in hardware — **that question is the one our recipe was designed around.** Every trained operation was chosen to have a known analog implementation: divisive (current-mode) normalization, fixed rotations for the position encoding, translinear multipliers for the gated MLP, the subthreshold-exponential/KCL circuit for attention's softmax, and EP's two relaxation phases for the learning rule. A component-by-component hardware mapping is written up; we can share it. One question we would value your judgment on as you read: the dynamics results and the scaling demonstration are separable, and only part of the former is required for the latter. Our current inclination is two papers — the dynamics and its taxonomy of stability controls now, and the scale demonstration later, submitted to the most general venue the result supports, with the dynamics paper as its theoretical reference. A single merged paper is the alternative. Your judgment on which is the stronger structure would be valuable. Best, Yuren --- ## 发送前 checklist(内部,不进邮件) 1. [LINK] 换成真实 Overleaf 链接(zip 已备好: overleaf_dynamics_v3.zip,含两对象改写+KHS26引用+floss重调数字) 2. 附 a1_demo.ipynb(token 已内置) 3. ~~gap 数字核对~~ RESOLVED: f3e3cont 终值 1.2883 > 1.2808 → "0.050 nats" 保持不变 4. 3.2× 的对照句("18-day → ~4-day")口吻是否合适你再定——它来自他们自己的 Appendix B,准确但直接 5. 成本表 2026-07-12 已按用户实测锚点重校准(1B×10B BP H200 租赁 ≈$1k → $16.7k/EF 全包上限; 市场价 H100/H200 TF32 路径 $6-9k/EF)。旧表锚 AWS-A100-savings,每 FLOP 贵 2.5-3×,已废。 完整推导见 docs/campaign/COST_MODEL.md