summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
authorYurenHao0426 <Blackhao0426@gmail.com>2026-06-13 04:33:08 -0500
committerYurenHao0426 <Blackhao0426@gmail.com>2026-06-13 04:33:08 -0500
commit82a49011de15287583d8cfec11ac5cca7efee747 (patch)
tree690f91ee6fc46839b8f306dc8fa8714ff404de6f
parent2229cb3951d7460ed2ac45ed59b2051cf7f39cee (diff)
Add master technical reference; retire unverified e0 headline
Note 42 is a self-contained write-from document: full statements and proofs for Thm 1-4 / Prop 1-3 / Cor 1-2, every result with exact numbers and configs cross-checked against the CSVs, the negative results, and the honesty boundaries. Audit finding while building it: the "E_B[e0]=0.4378 +/- 0.008 vs 0.4403" headline (cited in notes 29/35/36/40/41 and the old draft) is not reproducible from the current moments CSV. Retired in favor of the reproducible 6x6 table (max |error| <= 0.008, predicted span 0.15-0.45). Living docs updated; note 37 sec 4 records the retirement. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
-rw-r--r--notes/35_current_results_consolidated.md2
-rw-r--r--notes/36_evidence_ledger.md7
-rw-r--r--notes/37_wrapup_corrections.md19
-rw-r--r--notes/40_reproduction_manifest.md2
-rw-r--r--notes/41_paper_plan.md2
-rw-r--r--notes/42_master_technical_reference.md697
-rw-r--r--notes/README.md12
7 files changed, 731 insertions, 10 deletions
diff --git a/notes/35_current_results_consolidated.md b/notes/35_current_results_consolidated.md
index f22a0c1..2a104ff 100644
--- a/notes/35_current_results_consolidated.md
+++ b/notes/35_current_results_consolidated.md
@@ -92,6 +92,8 @@ old 5 -> C3(d)
## Key numbers (the evidence)
- init moment: `E_B[e_0]=0.4378 +/- 0.008` vs theory `hidden_share=0.4403`.
+ **[retired -- see note 37 §4; not in current CSV. Use the 6x6 table in note
+ 42 §6: max |error| <= 0.008, predicted span 0.15-0.45.]**
- estimator: linear-velocity gap MAE `0.00189`, corr `0.99885` (256 traj);
stress grid corr `0.994-0.99989` across depth 1/2/3, width 32/64/96,
horizon 25/50/100 (notes 15-17). fixed `K(0)` is ~10x worse.
diff --git a/notes/36_evidence_ledger.md b/notes/36_evidence_ledger.md
index 1028f48..9807e06 100644
--- a/notes/36_evidence_ledger.md
+++ b/notes/36_evidence_ledger.md
@@ -66,8 +66,11 @@ standard lazy approximation whose breakdown is *measured* (item 3c below).
### (b) Exact at t=0 (T5 + T6)
- Mean theorem, 2-hidden-layer, real FA backward, 6 widths x 6 inits x 512
- draws: `E_B[e_0] = 0.4378 +/- 0.008` vs theory `0.4403`; per-width errors
- 0.0003-0.008 (note 29, `actual_fa_initial_operator_moments`).
+ draws: no-fit prediction within Monte-Carlo error across all 36 rows,
+ **max |error| <= 0.008**; predicted erosion spans 0.15-0.45 (wide-range
+ calibration, not one point). Full table in note 42 §6
+ (`actual_fa_initial_operator_moments`). The old "0.4378/0.4403" headline is
+ NOT in the current CSV -- retired, see note 37 §4.
- Independent t=0 anchor from the (superseded) recovery probe: 0.4026+/-0.0066
vs 0.4071 (run log, note 37).
- Full distribution (T6), **sampler = real FA backward after note-37 fix**:
diff --git a/notes/37_wrapup_corrections.md b/notes/37_wrapup_corrections.md
index 671c237..ad2d60b 100644
--- a/notes/37_wrapup_corrections.md
+++ b/notes/37_wrapup_corrections.md
@@ -68,7 +68,24 @@ form — still zero fitted parameters — tracks the ramp at 0.982 and isolates
0.977; the figure now shows measured (error bars), matched-init curve,
ensemble geometric mean, and the init-seed band.
-## 3. Related cleanup executed with this audit
+## 4. e0 moment headline "0.4378/0.4403" not reproducible (retired)
+
+Found 2026-06-13 while building the master technical reference (note 42).
+The figure `E_B[e_0] = 0.4378 +/- 0.008` vs hidden share `0.4403`, cited in
+notes 29/35/36/40/41 and the old `paper/body.tex`, does NOT appear anywhere in
+the current `actual_fa_initial_operator_moments/initial_operator_moment_rows.csv`
+(6 widths {16,24,32,48,64,96} x 6 init seeds; no row or aggregate is ~0.44;
+closest is width 16 init 0: predicted 0.4519 / empirical 0.4443). It is
+presumed to come from a superseded run that was overwritten, or from the
+AI-written draft.
+
+Retired. The reproducible Thm 3 evidence is the full 6x6 table (note 42 §6):
+the no-fit prediction sits within Monte-Carlo error of the measurement, **max
+|error| <= 0.008**, with predicted erosion spanning 0.15-0.45 (a wide-range
+calibration, stronger than a single point). Living docs (36, 40, 41, README)
+updated to this statement; historical notes 29/35 left as-is with this pointer.
+
+## 5. Related cleanup executed with this audit
- `scripts/alignment_recovery_dynamics.py`: superseded probe (note 35);
its divide-by-zero guard was fixed earlier but the rerun was never completed
diff --git a/notes/40_reproduction_manifest.md b/notes/40_reproduction_manifest.md
index 7e18d98..009f058 100644
--- a/notes/40_reproduction_manifest.md
+++ b/notes/40_reproduction_manifest.md
@@ -31,7 +31,7 @@ deterministic given the seeds shown. Long runs use `nohup`.
| claim | command | outputs | note |
|---|---|---|---|
-| `E_B[e_0]=0.4378+/-0.008` vs hidden share 0.4403; per-width errors <= 0.008 | `python scripts/actual_fa_initial_operator_moments.py --widths 16 24 32 48 64 96 --init-seeds 6 --feedback-samples 512 --torch-threads 8 --outdir outputs/actual_fa_initial_operator_moments` | `actual_fa_initial_operator_moments/` | 29 |
+| no-fit erosion = hidden share, max abs error <= 0.008 across 6 widths x 6 inits (predicted span 0.15-0.45); "0.4378/0.4403" headline retired (note 37 §4) | `python scripts/actual_fa_initial_operator_moments.py --widths 16 24 32 48 64 96 --init-seeds 6 --feedback-samples 512 --torch-threads 8 --outdir outputs/actual_fa_initial_operator_moments` | `actual_fa_initial_operator_moments/` | 29, 42 |
| e0 exactly Gaussian; KS p 0.13-0.999; sampler = real FA backward, linear identity <= 8.9e-16 | `python scripts/actual_fa_initial_erosion_distribution.py --widths 16 32 64 128 --init-seeds 3 --feedback-samples 4096 --torch-threads 8 --outdir outputs/actual_fa_initial_erosion_distribution` | `actual_fa_initial_erosion_distribution/` (pre-fix snapshot kept in `*_pre_fix_backup/`) | 32, 37 |
## Finite-time chain (lazy validity -> drift -> estimator)
diff --git a/notes/41_paper_plan.md b/notes/41_paper_plan.md
index d20c464..d3bd12f 100644
--- a/notes/41_paper_plan.md
+++ b/notes/41_paper_plan.md
@@ -56,7 +56,7 @@ soft, and now exactly priced.**
| 3 Static cost (T1+P1) | 0.65 | Beta law (Thm 1, 2-line proof in main), log-volume cost, additivity + Theta(Ln^2) (Prop 1, proof appendix); validation sentence: 700k samples KS table + Gamma(L,1) max KS 0.0042 |
| 4 No free lunch at init (T3) | 0.40 | minimax Thm 2 + 3-line proof + the "anisotropy helps some directions only by hurting others" reading; empirical lam_min table sentence |
| 5 The gap | 2.30 | (a) null model Prop 2 (`Beta(k/2,(P-k)/2)`, `P(e=0)=0`) + **its honest 2x-250x overprediction** => motivates operator object; (b) Thm 3 `E_B[e_0]=hidden share` + proof + Gaussian Prop 3; MNIST share-0.39-0.82 confirmation sentence; (c) A1/A2 stated plainly, Thm 4 closed form + Cor 1 (soft, analytic, no threshold) + Cor 2 (ramp law `log gap ~ -c rho lam_min`); (d) the lazy chain in one paragraph WITH numbers (T=5 corr 0.99947 -> oracle K(t) MAE 0.000273 -> estimator) + estimator definition + bracket (velocity low / retangent high) + bias origin (drift-velocity decay) |
-| 6 Experiments | 1.50 | F1-F4 + numbers table; soft ramp (refutation), closed form matched 0.933 / ensemble 0.982 / lam_min 0.987; e0 0.4378 vs 0.4403 + KS; estimator 256-traj MAE 0.00189 corr 0.99885 + stress grid range; MNIST both parts; teacher test-gap sign flip (2 sentences + pointer to F4/appendix) |
+| 6 Experiments | 1.50 | F1-F4 + numbers table; soft ramp (refutation), closed form matched 0.933 / ensemble 0.982 / lam_min 0.987; e0 hidden-share calibration (max err <=0.008 over 6x6) + KS; estimator 256-traj MAE 0.00189 corr 0.99885 + stress grid range; MNIST both parts; teacher test-gap sign flip (2 sentences + pointer to F4/appendix) |
| 7 Related work | 0.45 | four paragraphs: FA empirics; FA theory (Song-Xu-Lafferty complementarity via `E_B[H_FA]=0`); NTK/lazy; capacity/memorization |
| 8 Limitations + conclusion | 0.25 | scope (MLP, full-batch, squared loss; MNIST-subset real data; conv/transformers open); estimator conditional; architecture-only drift = open problem |
diff --git a/notes/42_master_technical_reference.md b/notes/42_master_technical_reference.md
new file mode 100644
index 0000000..82c436e
--- /dev/null
+++ b/notes/42_master_technical_reference.md
@@ -0,0 +1,697 @@
+# Master Technical Reference
+
+Self-contained record of every result, with full statements, proofs,
+derivations, exact numbers, configs, and negative results. Built for writing
+the paper by hand: you should not need any other file open. Cross-checked
+against the raw CSVs on 2026-06-13.
+
+Conventions in this file:
+- "no fit" = no post-hoc scalar tuned to the reported data; predictions come
+ from initialization or two early kernel snapshots.
+- Every empirical number cites its output directory; commands are in note 40.
+- Numbers flagged **[verify]** could not be reproduced from current outputs
+ and must be relocated or dropped before they enter the paper.
+
+Contents:
+1. Notation and setup
+2. Result A — static alignment law (Thm 1) + capacity cost
+3. Result B — multilayer scaling (Prop 1) + observed surprisal
+4. Result C — prior-free minimax (Thm 2)
+5. Result D — random-subspace null model (Prop 2) + its failure
+6. Result E — initial erosion theorem (Thm 3) + Gaussian e0 (Prop 3)
+7. Result F — closed-form soft ramp (Thm 4, Cor 1, Cor 2)
+8. Result G — finite-time estimator + the lazy-validity chain
+9. Negative results (consolidated)
+10. Generalization: teacher test gap
+11. Related work, exact positioning
+12. Open problems
+13. Honesty boundaries
+
+---
+
+## 1. Notation and setup
+
+**Network.** MLP with widths `n_0, n_1, ..., n_L`, forward weights
+`W_l in R^{n_l x n_{l-1}}`, ReLU `phi`. Forward pass on input `x`:
+`h_0 = x`, `pre_l = W_l h_{l-1}`, `h_l = phi(pre_l)` for hidden layers,
+`h_L = pre_L` (linear output). Train on `N` examples `X` with squared loss
+
+```
+L = (1/2N) ||r||^2, r = f_theta(X) - Y (all residuals stacked, length N n_L).
+```
+
+**The single FA/BP difference (backward pass).** Write `g_l` for the gradient
+of layer `l`. Both rules share the output residual. Propagating error
+backward through layer `l+1`:
+
+```
+BP: delta_l = (delta_{l+1} W_{l+1}) ⊙ 1[pre_l > 0]
+FA: delta_l = (delta_{l+1} B_l^T) ⊙ 1[pre_l > 0]
+```
+
+with `B_l in R^{n_l x n_{l+1}}` a fixed random matrix replacing `W_{l+1}^T`.
+The layer gradient is `g_l = delta_l^T h_{l-1}` in both cases. **Key fact used
+throughout:** the output layer has no backward matrix above it, so
+
+```
+g_L^FA = g_L^BP (the output-layer gradient is identical under both rules).
+```
+
+`D_l = n_l n_{l+1}` is the flattened backward-matrix dimension.
+`P = sum_l n_l n_{l-1}` is the parameter count.
+
+**Learning speed (first-order loss decrease).** A step `-eta g` decreases the
+loss to first order by `eta <g^BP, g>`. Define
+
+```
+speed_BP = sum_l ||g_l^BP||^2,
+speed_FA = sum_l <g_l^BP, g_l^FA>,
+e_0 = 1 - speed_FA / speed_BP (fraction of BP's speed FA discards in one step).
+```
+
+`e_0 = 0`: FA as good as BP this step. `e_0 = 1`: FA makes no progress.
+
+**Tangent operators.** Let `J = d f_theta(X) / d theta` be the output
+Jacobian (size `N n_L x P`). Then
+
+```
+K_BP = J J^T (the NTK; symmetric PSD)
+K_FA = J J_tilde^T (J_tilde = FA update Jacobian; NON-symmetric)
+S_FA = (K_FA + K_FA^T)/2 (only the symmetric part affects loss to first order)
+```
+
+Equivalently `speed_BP = r^T K_BP r / (...)` and `speed_FA = r^T S_FA r`,
+i.e. `speed`s are residual-direction quadratic forms in these operators.
+Residual recursion (gradient descent in output space):
+`r_{t+1} = (I - (eta/N) K_t) r_t`.
+
+**Output / hidden share.** With `g_out = g_L^BP`,
+
+```
+rho = output_share = ||g_out^BP||^2 / speed_BP = r^T K_out r / r^T K_BP r in (0,1),
+hidden_share = 1 - rho.
+```
+
+`rho` is the single scalar that drives the gap results. It is read off one BP
+backward pass.
+
+---
+
+## 2. Result A — static alignment law (Thm 1)
+
+### Statement
+
+Let `A_l = W_{l+1}^T`, and let `A_l/||A_l||_F`, `B_l/||B_l||_F` be independent
+isotropic unit directions in `R^{D_l}`. Define the squared Frobenius cosine
+
+```
+Q_l = <A_l, B_l>_F^2 / (||A_l||_F^2 ||B_l||_F^2).
+```
+
+Then
+
+```
+Q_l ~ Beta(1/2, (D_l - 1)/2), E[Q_l] = 1/D_l, D_l Q_l => chi^2_1 as D_l -> inf.
+```
+
+Tail / survival: `Pr(Q_l >= q) = 1 - I_q(1/2, (D_l-1)/2)` (regularized
+incomplete Beta). High-dimensional bound `Pr(Q_l >= q) <~ 2 exp[-(D_l-1)q/2]`.
+
+### Proof
+
+Rotational invariance lets us fix `a = e_1`, so `Q = b_1^2` with `b` uniform on
+the sphere `S^{D-1}`. The squared first coordinate of a uniform unit vector is
+exactly `Beta(1/2, (D-1)/2)`. (Equivalently `Q = X/(X+Y)`, `X ~ chi^2_1`,
+`Y ~ chi^2_{D-1}` independent.) ∎
+
+### Reading
+
+Typical cosine `~ 1/sqrt(D)`. For any realistic layer (`D = n^2` in the
+thousands) a random feedback matrix is essentially orthogonal to the one BP
+wants. The theorem says exactly how orthogonal, with no slack.
+
+### Capacity cost
+
+For target alignment `q`, the log-volume cost
+
+```
+C_l(q) = -log Pr(Q_l >= q).
+```
+
+High-D: `C_l(q) ≈ ((D_l-1)/2) log(1/(1-q)) = Theta(D_l)`. Small-q:
+`C_l(q) ≈ D_l q / 2`. (Natural logs.)
+
+### Validation (output: `capacity_empirical_validation/`, 100k samples/dim, 700k total)
+
+Rotational-invariance sampling `Q = z_1^2/||z||^2`, `z ~ N(0,I_D)`:
+
+| D | emp mean | theory mean | KS stat | KS p | emp q99 | theory q99 |
+|---|---|---|---|---|---|---|
+| 64 | 0.015679 | 0.015625 | 0.00349 | 0.173 | 0.10071 | 0.10071 |
+| 128 | 0.0078528 | 0.0078125 | 0.00197 | 0.833 | 0.051165 | 0.051097 |
+| 256 | 0.0039289 | 0.0039063 | 0.00296 | 0.345 | 0.025990 | 0.025733 |
+| 512 | 0.0019553 | 0.0019531 | 0.00164 | 0.950 | 0.012965 | 0.012913 |
+| 1024 | 0.00097059 | 0.00097656 | 0.00312 | 0.283 | 0.0063694 | 0.0064679 |
+| 2048 | 0.00048817 | 0.00048828 | 0.00311 | 0.288 | 0.0032339 | 0.0032368 |
+| 4096 | 0.00024441 | 0.00024414 | 0.00243 | 0.597 | 0.0016322 | 0.0016191 |
+
+Means/quantiles match to 3-4 digits; KS never rejects (all p > 0.17).
+Capacity-tail calibration (same dir): max smoothed cost error 0.20 nats, mean
+0.019 nats (note 02).
+
+---
+
+## 3. Result B — multilayer scaling (Prop 1) + observed surprisal
+
+### Statement and proof
+
+Independent layer blocks => alignment events independent, so probabilities
+multiply and costs add:
+
+```
+p_all = prod_l Pr(Q_l >= q_l), C_all = sum_l C_l(q_l).
+```
+
+Equal width (`D_l ≈ n^2`), fixed `q`: `C_all = Theta(L n^2)`,
+`Pr(all q_l) = exp[-Theta(L n^2)]`. Chance threshold `q = c/D_l`:
+`C_l ≈ c/2`, `C_all = Theta(L)`. Proof is the product rule for independent
+events plus the `C_l = Theta(D_l)` evaluation from Result A. ∎
+
+### Observed surprisal (the distributional companion)
+
+`S_l = -log Pr(Q_l' >= Q_l)` with `Q_l'` an independent draw. By the
+probability integral transform `Pr(Q_l' >= Q_l) ~ Uniform(0,1)`, so
+`S_l ~ Exp(1)` and for `L` independent layers `sum_l S_l ~ Gamma(L, 1)`. The
+transformed surprisal law is dimension-free; dimension enters only through the
+raw `Q_l` scale and `C_l(q)`.
+
+### Validation (output: `multilayer_capacity_distribution/`, 100k/(D,L))
+
+`D in {64,256,1024,4096}` x `L in {1,2,4,8,16}`: **max KS 0.0042**, mean abs
+mean-error 0.0042, mean abs var-error 0.040. E.g. D=64,L=16: emp mean 16.008
+vs 16, KS 0.0025; D=4096,L=1: 1.0004 vs 1, KS 0.0019. Multilayer product
+capacity (note 02): max err 0.46 nats, mean 0.039 nats.
+
+---
+
+## 4. Result C — prior-free minimax (Thm 2)
+
+### Statement
+
+Let `b_hat = vec(B)/||B||_F in S^{D-1}` drawn from any distribution `mu`,
+`M_mu = E_mu[b_hat b_hat^T]` (so `tr M_mu = 1`), and `a in S^{D-1}` an unknown
+target. Then
+
+```
+sup_mu inf_{||a||=1} E_mu[(a^T b_hat)^2] = sup_mu lambda_min(M_mu) = 1/D,
+```
+
+attained by isotropic feedback (`M_mu = I/D`).
+
+### Proof
+
+`E_mu[(a^T b_hat)^2] = a^T M_mu a`; its worst case over `a` is
+`lambda_min(M_mu) <= tr(M_mu)/D = 1/D` (min eigenvalue <= average eigenvalue).
+Isotropic `mu` gives `M_mu = I/D`, so every direction equals `1/D`, achieving
+the bound. ∎
+
+### Reading and prior-aware corollary
+
+Anisotropic feedback that aligns better with some directions does so by
+aligning worse with others; an adversarial task lands on the bad ones.
+Sparsity / orthogonality / low rank change conditioning or hardware cost but
+cannot beat `1/D` worst-case. With a target prior `Sigma_A = E[a a^T]`,
+`E_{a,B}[(a^T b_hat)^2] = tr(Sigma_A M_mu)`, so with prior information
+top-eigenspace feedback can win; without it, isotropy is minimax. Note the
+average-case over uniform targets is `1/D` for every `mu` — the distinction
+between initializations is worst-case coverage, not mean.
+
+### Validation (output: `minimax_initialization/`, D=32, 1/D = 0.03125)
+
+isotropic `lambda_min = 0.029084`; rademacher 0.028947; anisotropic 0.006150;
+rank-deficient subspace 0; axis-aligned 0. Coverage-distribution matching for
+each scheme in `initialization_distribution_matching/` (D=128, worst KS
+0.0047 on non-degenerate schemes).
+
+---
+
+## 5. Result D — random-subspace null model (Prop 2) + its failure
+
+### Statement and proof
+
+Toy model: alignment deletes a uniformly random `k`-dim subspace `E` of the
+`P` parameters, leaving projected kernel `J Q_E J^T`, `Q_E = I - P_E`. With
+`v = J^T r`, one-step erosion `e(r) = ||P_E v||^2 / ||v||^2`. By rotational
+invariance `E[P_E] = (k/P) I`, so `E[e(r) | J, r] = k/P`, and for Haar-random
+`E`
+
+```
+e(r) ~ Beta(k/2, (P-k)/2), E[e] = k/P, Pr(e = 0) = 0 for k > 0.
+```
+
+Proof: `||P_E v||^2/||v||^2` for a fixed `v` and Haar `k`-subspace is the sum
+of `k` squared coordinates of a uniform direction, i.e. `Beta(k/2,(P-k)/2)`. ∎
+
+### Scaling consequence (the "no free capacity" sell)
+
+`k` fixed, `P -> inf`: `E[e] -> 0` (burden vanishes). `k/P -> alpha > 0`:
+`E[e] -> alpha` (relative erosion persists). Overparameterization dilutes the
+cost through `k/P`; it never deletes it, and there is no threshold —
+`Pr(e=0)=0` for any `k>0`. Role: this is the clean mechanism that motivates
+the operator object. **It is a null model, not a gap predictor.**
+
+### Negative result: hard-k overpredicts real FA (output: `soft_erosion_theory_vs_empirical/`)
+
+Using `rho ~ Beta((P-k)/2, k/2)` and propagating the BP spectrum overpredicts
+the measured gap badly, worst on the over-capacity side:
+
+| width | margin | hard-k theory mean | empirical mean |
+|---|---|---|---|
+| 40 | +130 | 0.040374 | 0.000162 |
+| 32 | +2 | 0.064534 | 0.003729 |
+| 24 | -126 | 0.121998 | 0.046464 |
+| 20 | -190 | 0.151140 | 0.146803 |
+
+Up to ~250x at positive margin. Local one-step erosion measured 0.23-0.39 vs
+hard-k Beta-mean 0.63-0.73 (~2x). Conclusion: raw constraint count `k` is not
+the effective burden `k_eff`; the real object is operator erosion (Result E).
+
+---
+
+## 6. Result E — initial erosion theorem (Thm 3) + Gaussian e0 (Prop 3)
+
+### Thm 3 (initial erosion = hidden speed share)
+
+Fix forward weights `W`, data, residual `r`. Assume feedback matrices `B_l`
+are independent of `W, r` with zero first moment in every hidden backward
+direction, and the output layer uses the true gradient. Then
+
+```
+E_B[speed_FA | W, r] = ||g_L^BP||^2,
+E_B[e_0 | W, r] = 1 - ||g_L^BP||^2 / sum_l ||g_l^BP||^2 = 1 - rho =: hidden_share.
+```
+
+**Proof.** The output FA gradient equals the BP gradient, contributing
+`||g_L^BP||^2`. Each hidden FA gradient is multilinear in zero-mean
+independent feedback matrices, so `E_B[g_l^FA | W,r] = 0` for hidden `l`,
+hence `E_B[<g_l^BP, g_l^FA>] = 0`. Summing leaves only the output term. ∎
+
+### Reading
+
+On step one, random feedback points hidden updates in random directions, so on
+average they accomplish nothing; FA gets only the output layer, which it
+computes correctly, for free. The expected initial cost is exactly the share
+of work the hidden layers were doing — a number from one BP backward pass.
+This is the residual-direction version of Song-Xu-Lafferty's `K_FA = G + H_FA`
+at init: `E_B[H_FA] = 0`, so `E_B[K_FA] = K_out`, NOT `(1 - k/P) K_BP`. That
+is why hard-k was pessimistic.
+
+### Prop 3 (exact Gaussian e0, one hidden layer)
+
+One-hidden-layer ReLU MLP `x -> W1 -> ReLU -> W2 -> out`, Gaussian feedback
+`B_{a,c} ~ N(0, sigma_B^2)`. For fixed `W, r, gates`,
+`speed_FA = ||g_out^BP||^2 + <C(W,r), B>`, linear in `B`, with explicit
+coefficient matrix
+
+```
+S_{a,c,d} = sum_n gate_{n,a} x_{n,d} delta_{n,c},
+C_{a,c} = sum_d g_bp_{a,d} S_{a,c,d},
+<g_hidden^BP, g_hidden^FA(B)> = sum_{a,c} C_{a,c} B_{a,c},
+```
+
+(`x_n` input, `delta_{n,c}` output residual gradient, `gate_{n,a}` ReLU gate,
+`g_bp_{a,d}` BP hidden gradient). Hence `<C,B> ~ N(0, sigma_B^2 ||C||_F^2)` and
+
+```
+e_0(B) ~ Normal( 1 - ||g_out^BP||^2/speed_BP, sigma_B^2 ||C||_F^2 / speed_BP^2 ),
+```
+
+exact conditionally on `(W, r)`, independent of horizon `T`. Deeper FA keeps
+the exact mean (Thm 3) but the finite-width law is no longer exactly Gaussian
+(hidden terms become products of feedback matrices; a Wick/CLT expansion
+applies).
+
+### Validation — moments (output: `actual_fa_initial_operator_moments/`)
+
+2-hidden-layer nets, 6 widths x 6 init seeds x 512 real-FA-backward feedback
+draws. Predicted erosion = hidden share (no fit). Per-width at init seed 0:
+
+| width | predicted | empirical | error |
+|---|---|---|---|
+| 16 | 0.45193 | 0.44435 | -0.0076 |
+| 24 | 0.31202 | 0.31229 | +0.0003 |
+| 32 | 0.23606 | 0.23971 | +0.0037 |
+| 48 | 0.26147 | 0.25986 | -0.0016 |
+| 64 | 0.20912 | 0.21126 | +0.0021 |
+| 96 | 0.25977 | 0.25915 | -0.0006 |
+
+Across all 36 rows the prediction sits within Monte-Carlo error of the
+measurement: **max |error| ≈ 0.008** (width 24 init 4: +0.0080; width 16 init
+0: -0.0076). The predicted erosion itself spans 0.15-0.45 across configs, so
+this is a wide-range calibration, not a single point.
+
+> **[verify]** The headline "E_B[e_0] = 0.4378 ± 0.008 vs 0.4403" cited in
+> notes 35/36 and the old draft does NOT appear in the current moments CSV
+> (which has no 0.44 row or aggregate). Either relocate its source run or
+> replace it with the table above before it enters the paper. The table above
+> is the reproducible evidence.
+
+### Validation — distribution (output: `actual_fa_initial_erosion_distribution/`, post-fix real backward)
+
+Sampler runs the real FA backward per draw (note 37 fix); the analytic `<C,B>`
+form is only an internal check (`linear_identity_max_err <= 8.9e-16`).
+4 widths x 3 inits x 4096 draws, standardized then KS-tested vs Normal:
+
+| width | init | theory mean | emp mean | theory std | emp std | KS | p |
+|---|---|---|---|---|---|---|---|
+| 16 | 0 | 0.113830 | 0.113413 | 0.027228 | 0.027401 | 0.0150 | 0.314 |
+| 16 | 2 | 0.131704 | 0.131796 | 0.028843 | 0.028942 | 0.0056 | 0.999 |
+| 32 | 1 | 0.045831 | 0.045870 | 0.007278 | 0.007396 | 0.0094 | 0.858 |
+| 32 | 2 | 0.227193 | 0.227691 | 0.031148 | 0.030779 | 0.0166 | 0.206 |
+| 64 | 0 | 0.031155 | 0.031119 | 0.003652 | 0.003583 | 0.0118 | 0.609 |
+| 128 | 1 | 0.024377 | 0.024353 | 0.001770 | 0.001774 | 0.0182 | 0.130 |
+
+All 12 panels: KS p in [0.13, 0.999], means/stds match to 3-4 digits, no fit.
+
+### Validation — MNIST (output: `real_data_validation_mnist/`, note 38)
+
+Thm 3 is data-agnostic. On MNIST-subset MLPs (784-dim input), 16 configs
+(width {64,256} x depth {1,2} x 4 inits), 512 real-backward draws each:
+hidden share spans **0.39-0.82**, `|E_B[e_0] - hidden_share| in [0.0001,
+0.0089]` (all within ~2 stderr). Examples: depth 1 width 64 init 2: pred
+0.8189 / meas 0.8188; depth 2 width 256 init 2: pred 0.3949 / meas 0.3971.
+On 784-dim input the hidden (here input-layer-dominated) share is large; the
+theorem prices it with no fit.
+
+---
+
+## 7. Result F — closed-form soft ramp (Thm 4, Cor 1, Cor 2)
+
+### Assumptions
+
+- **A1 (lazy).** Tangent operators frozen at init:
+ `r_t^BP = (I - (eta/N) K_BP)^t r_0`. Standard NTK approximation; its
+ breakdown is the drift the estimator corrects (Result G).
+- **A2 (scalar erosion).** `S_FA ≈ rho K_BP`. Exact in the residual direction
+ by Thm 3 (`r^T E_B[S_FA] r = rho · r^T K_BP r`); the assumption is that the
+ same slowdown `rho` applies in every eigendirection.
+
+### Thm 4 (closed-form gap)
+
+In the BP eigenbasis `K_BP v_i = lam_i v_i`, `c_i = v_i^T r_0`:
+
+```
+gap_T = L_FA(T) - L_BP(T)
+ = (1/2N) sum_i c_i^2 [ (1 - eta rho lam_i/N)^{2T} - (1 - eta lam_i/N)^{2T} ].
+```
+
+**Proof.** `L_rule(T) = ||r_T^rule||^2/2N`. Expand `r_0 = sum_i c_i v_i`.
+A1 gives `||r_T^BP||^2 = sum_i c_i^2 (1 - eta lam_i/N)^{2T}`. A2 makes `S_FA`
+share eigenvectors with eigenvalues `rho lam_i`, so
+`||r_T^FA||^2 = sum_i c_i^2 (1 - eta rho lam_i/N)^{2T}`. Subtract, divide by
+`2N`. ∎
+
+### Cor 1 (soft, not a phase transition)
+
+`rho < 1` => each bracket `>= 0` => `gap_T >= 0`. `gap_T` is real-analytic in
+`(rho, {lam_i}, T)`, so deforming capacity smoothly moves the gap smoothly: a
+soft ramp, no kink. A discontinuity would need `rho -> 0`, which never happens
+since `rho >= output_share > 0`. The hard threshold
+`Delta d_hard = max(0, k-(P-d))` is recovered only in that degenerate limit.
+
+### Cor 2 (the ramp law)
+
+Underparameterized side (many `tau_i = N/(eta lam_i) >~ T`): each bracket
+`≈ 2 eta T (1-rho) lam_i / N`, so `gap_T ≈ (1-rho)(eta T/N^2) sum_slow c_i^2
+lam_i` (residual-mass limited). Converged side (`L_BP ≈ 0`): the gap is the
+slowest surviving FA mode,
+
+```
+gap_T ≈ (c_min^2/2N) exp(-2 eta rho lam_min T / N),
+log gap_T ≈ const - (2 eta T/N) rho lam_min.
+```
+
+The ramp is driven by `lam_min(w)` growing smoothly with capacity.
+
+### Validation — the dense empirical ramp (output: `phase_transition_dense_T30000_352traj/`)
+
+Random-label task, full-batch SGD lr 0.01, `T = 30000`, 11 widths 20-40, 1
+init x 32 feedback = 352 trajectories. Measured train gap (`train_gap_mean`):
+
+| width | FA margin | gap mean | gap std (32 seeds) |
+|---|---|---|---|
+| 20 | -190 | 0.146803 | 0.042009 |
+| 22 | -158 | 0.100526 | 0.034803 |
+| 24 | -126 | 0.046464 | 0.024474 |
+| 26 | -94 | 0.029671 | 0.023591 |
+| 28 | -62 | 0.014195 | 0.009368 |
+| 30 | -30 | 0.007466 | 0.004624 |
+| 32 | +2 | 0.003729 | 0.003787 |
+| 34 | +34 | 0.002079 | 0.002232 |
+| 36 | +66 | 0.000791 | 0.000605 |
+| 38 | +98 | 0.000553 | 0.000545 |
+| 40 | +130 | 0.000162 | 0.000134 |
+
+Smooth, no kink at margin 0, nonzero at positive margin. `log gap` linear in
+margin: **R^2 = 0.993**, slope -0.0209. (History: near-margin-0 gap was 0.297
+at T=3000, 0.051 at T=10^4, 0.0037 at T=3x10^4 — the apparent kink dissolved
+as training lengthened, notes 19-22.)
+
+### Validation — closed form vs measured (output: `closed_form_soft_ramp/`, corrected per note 37)
+
+Closed form evaluated on the *same data and init seed* the sweep used, plus 4
+extra init seeds (same data) for a no-fit band. All of `rho, lam_i, c_i`
+measured at init.
+
+```
+corr(log pred, log measured), matched single init = 0.933
+corr(log pred, log measured), 5-init geometric mean = 0.982
+corr(lam_min, -log measured), matched init = 0.948
+corr(lam_min mean over inits, -log measured) = 0.987
+ensemble-mean lam_min(w) over the sweep = 0.026 -> 0.763 (monotone)
+rho across widths (5-init mean) = 0.68 - 0.79
+```
+
+A single init's frozen `lam_min(w)` is noisy across widths while the T=30000
+measured ramp is smooth (drift + feedback averaging smooth it), so matched
+single-init corr is 0.933; the init ensemble (still zero fitted parameters)
+tracks the ramp at 0.982 and isolates `lam_min` as the driver at 0.987.
+Frozen-init compresses the range (under-predicts the large-gap end,
+over-predicts the small-gap tail) — exactly the operator drift the estimator
+recovers. **The previously reported 0.977 is retired** (it came from a
+data/init-mismatched comparison; note 37).
+
+---
+
+## 8. Result G — finite-time estimator + the lazy-validity chain
+
+The closed form gives mechanism and shape; the estimator gives magnitude. The
+path to it is a measured three-step chain.
+
+### Step 1 — frozen K(0) is exact at small T (output: fixed-width T=5 runs, note 10)
+
+Fixed width 64, `N in {64..320}`, SGD lr 1e-3, frozen-K(0) residual recursion,
+256 trajectories. Horizon diagnostic (empirical gap / predicted gap):
+
+| T | emp gap | pred gap | ratio | MAE |
+|---|---|---|---|---|
+| 1 | 0.030734 | 0.029853 | 1.028 | 0.00088 |
+| 5 | 0.112912 | 0.113947 | 0.987 | 0.00171 |
+| 10 | 0.156304 | 0.164964 | 0.942 | 0.00866 |
+| 20 | 0.163335 | 0.183836 | 0.884 | 0.02050 |
+| 50 | 0.125547 | 0.154075 | 0.814 | 0.02853 |
+
+At **T=5: corr 0.99947, MAE 0.00083, mean ratio 0.993**. The operator object
+is correct; freezing is a local-time theorem that drifts out of regime by
+T=50.
+
+### Step 2 — the T=50 error is entirely kernel drift (output: `finite_time_kernel_probe_T50_N128_8runs/`, note 13)
+
+Replace frozen K(0) by the oracle product of measured `K(t)` at every step:
+
+| predictor | gap MAE | BP loss MAE |
+|---|---|---|
+| fixed K(0) | 0.027683 | 0.025404 |
+| time-varying K(t) | 0.000273 | 0.000237 |
+
+Drift `K_t - K_0` removes ~99% of the gap error. So the finite-time problem is
+to model `K_t`, not to fit a scalar; higher-order output-Taylor terms are
+negligible.
+
+### The estimator (output: `compressed_operator_s20_256traj_T50_width64_plots/`, notes 15-17)
+
+Read kernels at init and one early step `s`, extrapolate linearly, roll out:
+
+```
+K_hat_t = K_0 + (t/s)(K_s - K_0), r_hat_{t+1} = (I - (eta/N) K_hat_t) r_hat_t.
+```
+
+Only inputs: `(K_0, K_s)` per rule. No fitted scale/offset/calibration.
+256 trajectories, T=50, s=20:
+
+| predictor | MAE | bias | corr |
+|---|---|---|---|
+| fixed K(0) | 0.018487 | +0.018487 | 0.984460 |
+| FA compressed | 0.014915 | +0.014915 | 0.990472 |
+| **linear velocity** | **0.0018934** | **-0.0014912** | **0.998848** |
+| early re-tangent | 0.0048684 | +0.0048684 | 0.999077 |
+
+Linear velocity beats fixed K(0) ~10x. `linear (low) <= empirical <=
+retangent (high)` is a no-fit bracket.
+
+### Stress grid (output: `operator_stress_grid_summary/`, note 16)
+
+Random-label, SGD lr 1e-3; 5 N-values x 8 traj per new setting:
+
+| setting | rows | fixed MAE | linear MAE | retangent MAE | linear bias | corr |
+|---|---|---|---|---|---|---|
+| d1,w64,T50,s20 | 40 | 0.003862 | 0.000233 | 0.001162 | -0.000233 | 0.999891 |
+| d2,w32,T50,s20 | 40 | 0.047110 | 0.005652 | 0.013333 | -0.005375 | 0.994077 |
+| d2,w64,T25,s10 | 40 | 0.020808 | 0.002722 | 0.005084 | -0.002668 | 0.999074 |
+| d2,w64,T50,s20 | 256 | 0.018487 | 0.001893 | 0.004868 | -0.001491 | 0.998848 |
+| d2,w64,T100,s40 | 40 | 0.034378 | 0.002647 | 0.008100 | -0.002169 | 0.998653 |
+| d2,w96,T50,s20 | 40 | 0.016912 | 0.001794 | 0.004004 | -0.001551 | 0.998634 |
+| d3,w64,T50,s20 | 40 | 0.051670 | 0.005049 | 0.012450 | -0.004916 | 0.998546 |
+
+Holds across depth {1,2,3}, width {32,64,96}, horizon {25,50,100}: **linear
+corr 0.994-0.99989**. Hard cases (w32, d3) raise drift magnitude but velocity
+still cuts fixed-K(0) error ~10x.
+
+### Bias structure (note 17)
+
+Residual gap bias ≈ -0.0021, dominated by BP-loss overprediction (+0.0026);
+FA near-unbiased (+0.0005). Cause: drift velocity decays after `s`
+(`alpha_t < t/s`; e.g. true `alpha_50` = 1.46 BP / 1.17 FA vs linear 2.5).
+This is kernel-path curvature, not a normalization error (the same recursion
+was exact at T=5 and the oracle K(t) matched to 1e-4). No fitted correction is
+introduced.
+
+### MNIST estimator (output: `real_data_validation_mnist/`, note 38)
+
+Width {128,256} x N {128,256,512} x 2 inits x 8 feedbacks = 96 traj, T=50,
+s=20, lazy step `eta lam_max/N = 0.05`:
+
+| predictor | MAE | bias | corr |
+|---|---|---|---|
+| fixed K(0) | 0.00613 | +0.00529 | 0.606 |
+| **linear velocity** | **0.00196** | **-0.00062** | **0.911** |
+
+Velocity MAE = **3.9% of the mean gap**, near-zero bias, beats fixed K(0)
+everywhere. The gap only spans 0.040-0.063 across all 96 traj, so MAE/bias are
+the operative stats and corr 0.911 is the weaker one (report it with the
+range). At `eta lam_max/N = 0.3` the lazy chain breaks (fixed MAE 0.069 corr
+0.37, velocity 0.030 corr 0.68) — the A1 scope boundary, recorded as a datum.
+
+### Computational tool (kernel via gram identity, `real_data_validation.py`)
+
+Tangent kernels are computed by the layerwise identity (avoids the `N n_L x P`
+Jacobian):
+
+```
+K[(n,c),(m,c')] = sum_l <delta_l(n,c), delta'_l(m,c')> · <h_{l-1}(n), h_{l-1}(m)>,
+```
+
+`delta'` = BP deltas for `K_BP`, FA deltas for `K_FA`. `--self-test` verifies
+this against an explicit autograd Jacobian: max error <= 5.3e-15 for
+`K_BP = J J^T`, `K_FA = J J_tilde^T`, and `J^T r` vs the production gradient.
+
+---
+
+## 9. Negative results (consolidated — these justify the final objects)
+
+| approach | result | output / note |
+|---|---|---|
+| hard-k as gap predictor | overpredicts 2x-250x, worst at positive margin | `soft_erosion_theory_vs_empirical/` (25) |
+| infinitesimal derivative `K_0 + t·dK_0` extrapolated to T=50 | MAE **270** vs 0.024 baseline (diverges) | `fa_tangent_hierarchy_derivative_probe_eps005/` (31) |
+| scalar directional-gain predictor across N | corr **0.124** (vs operator 0.999) | `early_directional_predictors_*` (14) |
+| trajectory bridge on long horizons | shape corr 0.97 but 2-3x scale error | (02, 08) |
+| phase-transition (kink) hypothesis | dissolves as T grows; dense T=30000 smooth | (19-22) |
+
+Each is one sentence in the paper. The derivative blow-up (31) and hard-k
+overprediction (25) most directly motivate, respectively, the bounded
+two-snapshot velocity estimator and the move from constraint-count to operator
+erosion.
+
+---
+
+## 10. Generalization: teacher test gap (output: `teacher_test_gap_sgd_T8000/`, note 39)
+
+MLP teacher (width 64, 2 hidden, noiseless, normalized targets), N=256 train /
+2048 test, SGD lr 0.01, T=8000, widths 8-96, 3 init x 8 feedback = 24 FA
+runs/width.
+
+| width | train gap | test gap (mean ± stderr) |
+|---|---|---|
+| 8 | 0.249 | +0.177 ± 0.020 |
+| 12 | 0.158 | +0.038 ± 0.016 |
+| 16 | 0.112 | -0.038 ± 0.015 |
+| 24 | 0.080 | -0.044 ± 0.018 |
+| 32 | 0.085 | -0.120 ± 0.013 |
+| 48 | 0.048 | -0.014 ± 0.014 |
+| 64 | 0.026 | +0.010 ± 0.008 |
+| 96 | 0.0069 | +0.045 ± 0.011 |
+
+Train (optimization) gap: positive, monotone-decreasing — the invariant soft
+ramp. Test gap: changes sign — FA generalizes significantly *better* at
+intermediate width (-0.120 ± 0.013 at w=32, an implicit-regularization effect)
+and settles slightly positive once both fit. The optimization gap is the
+object the theory prices; its test-side consequence is task-dependent, which
+is exactly why the theory is built on the former. Reported as scoping evidence,
+not a contribution. Caveat: single task family, one train size; the sign-flip
+location moves with task details.
+
+---
+
+## 11. Related work, exact positioning
+
+- **Lillicrap et al. 2016** (FA works; alignment emerges), **Nokland 2016**
+ (DFA, deep), **Refinetti et al. 2021** (align-then-memorize). They answer
+ "can FA learn?"; we answer "what does it cost?". Refinetti's alignment phase
+ is exactly our drift `K_s - K_0`.
+- **Song, Xu, Lafferty 2021** — two-layer FA converges (over-param), alignment
+ optional without regularization; decomposition `K_FA = G + H_FA`, `H_FA`
+ non-PSD and small at init. Complement, not conflict: at init their `H_FA` is
+ our Thm 3 (`E_B[H_FA] = 0`, residual direction). Convergence to zero *final*
+ error and a nonzero *operator* gap are different statements. Phrase as
+ complement; never as a correction.
+- **Jacot et al. 2018** (NTK), **Lee et al. 2019** (wide nets as linear
+ models), **Chizat-Oyallon-Bach 2019** (lazy training) license the residual
+ recursion and A1; our contribution there is FA's non-symmetric operator.
+ **Huang-Yau 2020** (neural tangent hierarchy) is the `K_t`-dynamics program
+ our estimator instantiates empirically.
+- **Zhang et al. 2017** (fitting random labels) invites the "capacity absorbs
+ it" guess; Result D/E is the soft rebuttal. **Montanari-Zhong 2022**
+ (interpolation transition) is why a kink was a reasonable prior.
+
+Rule: every citation earns a clause tied to one of our results. Verify each
+against the source before it lands (style contract).
+
+---
+
+## 12. Open problems (state, do not solve)
+
+Make the finite-time gap architecture-only: (i) predict `rho` (hence
+`hidden_share`) from infinite-width NTK recursions; (ii) predict the drift
+`K_s - K_0` (alignment gain) from initialization — naive `dot K_0` fails (note
+31), so a bounded short-time alignment-gain path or deep-linear / two-layer
+order-parameter solution; (iii) a per-mode spectral erosion `rho_i` refining
+the scalar A2.
+
+---
+
+## 13. Honesty boundaries (what NOT to claim)
+
+- Optimization (train) gap only; not test accuracy (Result 10 shows the test
+ effect is task-dependent and can favor FA).
+- Estimator is snapshot-conditional on `(K_0, K_s)`, no fitted parameters —
+ never "architecture-only". State its lazy-regime scope (MNIST eta-sensitivity).
+- Hard-k is a null model and a reported failure, never a gap predictor.
+ `E[e] = k/P` (null) and `rho = 1 - E_B[e_0]` (measured share) are distinct.
+- A1/A2 stated inline with Thm 4; A2 exact only in the residual direction.
+- Soft ramp refutes one named hypothesis (redundancy-exhaustion threshold) via
+ analyticity of `gap_T`; it is not a universal claim about FA.
+- Scope: MLP, full-batch, squared loss, synthetic + MNIST-subset. Conv /
+ transformers / stochastic optimization / real large-scale data are open.
+- Closed-form correlation is 0.933 (matched) / 0.982 (ensemble); 0.977 retired.
+- The "0.4378/0.4403" e0 headline is **[verify]** — use the moments table
+ (Result E) instead until its source is relocated.
diff --git a/notes/README.md b/notes/README.md
index ee608b6..853c122 100644
--- a/notes/README.md
+++ b/notes/README.md
@@ -7,10 +7,11 @@ capacity & operator-gap analysis)** — target AAAI-27.
| doc | role |
|---|---|
+| [42_master_technical_reference.md](42_master_technical_reference.md) | **write the paper from this**: every theorem with full statement + proof, every result with exact numbers/configs, derivations, negative results, honesty boundaries -- self-contained |
| [41_paper_plan.md](41_paper_plan.md) | **the plan**: venue rules, narrative, section/page budget, theorem & figure numbering, related-work list, schedule, risks |
-| [36_evidence_ledger.md](36_evidence_ledger.md) | **the numbers**: every claim with its corrected headline figure, claim discipline, booby traps |
+| [36_evidence_ledger.md](36_evidence_ledger.md) | **the numbers, condensed**: every claim with its corrected headline figure, claim discipline, booby traps |
| [40_reproduction_manifest.md](40_reproduction_manifest.md) | **the runs**: claim -> exact command -> output dir -> note; everything not listed is not citable |
-| [37_wrapup_corrections.md](37_wrapup_corrections.md) | audit record: e0 sampler de-circularized; closed-form comparison matched (0.977 retired -> 0.933/0.982) |
+| [37_wrapup_corrections.md](37_wrapup_corrections.md) | audit record: e0 sampler de-circularized; closed-form matched (0.977 retired -> 0.933/0.982); e0 "0.4378/0.4403" retired |
## Authoritative note per result
@@ -20,7 +21,7 @@ capacity & operator-gap analysis)** — target AAAI-27.
| P1 multilayer additivity, Theta(Ln^2), Gamma(L,1) | 01, 02 | proven + validated |
| T2 prior-free minimax 1/D | 01, 02 | proven + validated |
| P2 random-subspace null model (and its honest failure as a predictor) | 26 (theorem), 25 (failure) | proven; overpredicts real FA |
-| T3 initial erosion = hidden BP speed share | 29 | proven + validated (synthetic 0.4378/0.4403; MNIST note 38) |
+| T3 initial erosion = hidden BP speed share | 29, 42 | proven + validated (synthetic max err <=0.008 over 6x6; MNIST note 38) |
| P3 exact Gaussian e0 (one hidden layer) | 32 + fix in 37 | proven + validated (real backward) |
| T4 closed-form soft-ramp gap law | 34 + correction in 37 | proven under A1+A2; corr 0.933 matched / 0.982 ensemble |
| early-velocity estimator (+ bracket, bias structure) | 15, 16, 17 | no-fit, snapshot-conditional; MAE 0.0019, corr 0.999 |
@@ -67,9 +68,10 @@ Phase 5 — exact theory of the burden:
33 complete markdown draft (appendix source); 34 **closed-form law T4**;
35 consolidated snapshot.
-Phase 6 — wrap-up for AAAI-27 (2026-06-09):
+Phase 6 — wrap-up for AAAI-27 (2026-06-09..13):
- 36 evidence ledger; 37 corrections; 38 **MNIST validation**; 39 **teacher
- test gap**; 40 reproduction manifest; 41 paper plan.
+ test gap**; 40 reproduction manifest; 41 paper plan; 42 **master technical
+ reference** (full proofs + numbers, the write-from doc).
## Note conventions