diff options
| author | YurenHao0426 <Blackhao0426@gmail.com> | 2026-06-13 04:33:08 -0500 |
|---|---|---|
| committer | YurenHao0426 <Blackhao0426@gmail.com> | 2026-06-13 04:33:08 -0500 |
| commit | 82a49011de15287583d8cfec11ac5cca7efee747 (patch) | |
| tree | 690f91ee6fc46839b8f306dc8fa8714ff404de6f | |
| parent | 2229cb3951d7460ed2ac45ed59b2051cf7f39cee (diff) | |
Add master technical reference; retire unverified e0 headline
Note 42 is a self-contained write-from document: full statements and
proofs for Thm 1-4 / Prop 1-3 / Cor 1-2, every result with exact
numbers and configs cross-checked against the CSVs, the negative
results, and the honesty boundaries.
Audit finding while building it: the "E_B[e0]=0.4378 +/- 0.008 vs
0.4403" headline (cited in notes 29/35/36/40/41 and the old draft) is
not reproducible from the current moments CSV. Retired in favor of the
reproducible 6x6 table (max |error| <= 0.008, predicted span 0.15-0.45).
Living docs updated; note 37 sec 4 records the retirement.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
| -rw-r--r-- | notes/35_current_results_consolidated.md | 2 | ||||
| -rw-r--r-- | notes/36_evidence_ledger.md | 7 | ||||
| -rw-r--r-- | notes/37_wrapup_corrections.md | 19 | ||||
| -rw-r--r-- | notes/40_reproduction_manifest.md | 2 | ||||
| -rw-r--r-- | notes/41_paper_plan.md | 2 | ||||
| -rw-r--r-- | notes/42_master_technical_reference.md | 697 | ||||
| -rw-r--r-- | notes/README.md | 12 |
7 files changed, 731 insertions, 10 deletions
diff --git a/notes/35_current_results_consolidated.md b/notes/35_current_results_consolidated.md index f22a0c1..2a104ff 100644 --- a/notes/35_current_results_consolidated.md +++ b/notes/35_current_results_consolidated.md @@ -92,6 +92,8 @@ old 5 -> C3(d) ## Key numbers (the evidence) - init moment: `E_B[e_0]=0.4378 +/- 0.008` vs theory `hidden_share=0.4403`. + **[retired -- see note 37 §4; not in current CSV. Use the 6x6 table in note + 42 §6: max |error| <= 0.008, predicted span 0.15-0.45.]** - estimator: linear-velocity gap MAE `0.00189`, corr `0.99885` (256 traj); stress grid corr `0.994-0.99989` across depth 1/2/3, width 32/64/96, horizon 25/50/100 (notes 15-17). fixed `K(0)` is ~10x worse. diff --git a/notes/36_evidence_ledger.md b/notes/36_evidence_ledger.md index 1028f48..9807e06 100644 --- a/notes/36_evidence_ledger.md +++ b/notes/36_evidence_ledger.md @@ -66,8 +66,11 @@ standard lazy approximation whose breakdown is *measured* (item 3c below). ### (b) Exact at t=0 (T5 + T6) - Mean theorem, 2-hidden-layer, real FA backward, 6 widths x 6 inits x 512 - draws: `E_B[e_0] = 0.4378 +/- 0.008` vs theory `0.4403`; per-width errors - 0.0003-0.008 (note 29, `actual_fa_initial_operator_moments`). + draws: no-fit prediction within Monte-Carlo error across all 36 rows, + **max |error| <= 0.008**; predicted erosion spans 0.15-0.45 (wide-range + calibration, not one point). Full table in note 42 §6 + (`actual_fa_initial_operator_moments`). The old "0.4378/0.4403" headline is + NOT in the current CSV -- retired, see note 37 §4. - Independent t=0 anchor from the (superseded) recovery probe: 0.4026+/-0.0066 vs 0.4071 (run log, note 37). - Full distribution (T6), **sampler = real FA backward after note-37 fix**: diff --git a/notes/37_wrapup_corrections.md b/notes/37_wrapup_corrections.md index 671c237..ad2d60b 100644 --- a/notes/37_wrapup_corrections.md +++ b/notes/37_wrapup_corrections.md @@ -68,7 +68,24 @@ form — still zero fitted parameters — tracks the ramp at 0.982 and isolates 0.977; the figure now shows measured (error bars), matched-init curve, ensemble geometric mean, and the init-seed band. -## 3. Related cleanup executed with this audit +## 4. e0 moment headline "0.4378/0.4403" not reproducible (retired) + +Found 2026-06-13 while building the master technical reference (note 42). +The figure `E_B[e_0] = 0.4378 +/- 0.008` vs hidden share `0.4403`, cited in +notes 29/35/36/40/41 and the old `paper/body.tex`, does NOT appear anywhere in +the current `actual_fa_initial_operator_moments/initial_operator_moment_rows.csv` +(6 widths {16,24,32,48,64,96} x 6 init seeds; no row or aggregate is ~0.44; +closest is width 16 init 0: predicted 0.4519 / empirical 0.4443). It is +presumed to come from a superseded run that was overwritten, or from the +AI-written draft. + +Retired. The reproducible Thm 3 evidence is the full 6x6 table (note 42 §6): +the no-fit prediction sits within Monte-Carlo error of the measurement, **max +|error| <= 0.008**, with predicted erosion spanning 0.15-0.45 (a wide-range +calibration, stronger than a single point). Living docs (36, 40, 41, README) +updated to this statement; historical notes 29/35 left as-is with this pointer. + +## 5. Related cleanup executed with this audit - `scripts/alignment_recovery_dynamics.py`: superseded probe (note 35); its divide-by-zero guard was fixed earlier but the rerun was never completed diff --git a/notes/40_reproduction_manifest.md b/notes/40_reproduction_manifest.md index 7e18d98..009f058 100644 --- a/notes/40_reproduction_manifest.md +++ b/notes/40_reproduction_manifest.md @@ -31,7 +31,7 @@ deterministic given the seeds shown. Long runs use `nohup`. | claim | command | outputs | note | |---|---|---|---| -| `E_B[e_0]=0.4378+/-0.008` vs hidden share 0.4403; per-width errors <= 0.008 | `python scripts/actual_fa_initial_operator_moments.py --widths 16 24 32 48 64 96 --init-seeds 6 --feedback-samples 512 --torch-threads 8 --outdir outputs/actual_fa_initial_operator_moments` | `actual_fa_initial_operator_moments/` | 29 | +| no-fit erosion = hidden share, max abs error <= 0.008 across 6 widths x 6 inits (predicted span 0.15-0.45); "0.4378/0.4403" headline retired (note 37 §4) | `python scripts/actual_fa_initial_operator_moments.py --widths 16 24 32 48 64 96 --init-seeds 6 --feedback-samples 512 --torch-threads 8 --outdir outputs/actual_fa_initial_operator_moments` | `actual_fa_initial_operator_moments/` | 29, 42 | | e0 exactly Gaussian; KS p 0.13-0.999; sampler = real FA backward, linear identity <= 8.9e-16 | `python scripts/actual_fa_initial_erosion_distribution.py --widths 16 32 64 128 --init-seeds 3 --feedback-samples 4096 --torch-threads 8 --outdir outputs/actual_fa_initial_erosion_distribution` | `actual_fa_initial_erosion_distribution/` (pre-fix snapshot kept in `*_pre_fix_backup/`) | 32, 37 | ## Finite-time chain (lazy validity -> drift -> estimator) diff --git a/notes/41_paper_plan.md b/notes/41_paper_plan.md index d20c464..d3bd12f 100644 --- a/notes/41_paper_plan.md +++ b/notes/41_paper_plan.md @@ -56,7 +56,7 @@ soft, and now exactly priced.** | 3 Static cost (T1+P1) | 0.65 | Beta law (Thm 1, 2-line proof in main), log-volume cost, additivity + Theta(Ln^2) (Prop 1, proof appendix); validation sentence: 700k samples KS table + Gamma(L,1) max KS 0.0042 | | 4 No free lunch at init (T3) | 0.40 | minimax Thm 2 + 3-line proof + the "anisotropy helps some directions only by hurting others" reading; empirical lam_min table sentence | | 5 The gap | 2.30 | (a) null model Prop 2 (`Beta(k/2,(P-k)/2)`, `P(e=0)=0`) + **its honest 2x-250x overprediction** => motivates operator object; (b) Thm 3 `E_B[e_0]=hidden share` + proof + Gaussian Prop 3; MNIST share-0.39-0.82 confirmation sentence; (c) A1/A2 stated plainly, Thm 4 closed form + Cor 1 (soft, analytic, no threshold) + Cor 2 (ramp law `log gap ~ -c rho lam_min`); (d) the lazy chain in one paragraph WITH numbers (T=5 corr 0.99947 -> oracle K(t) MAE 0.000273 -> estimator) + estimator definition + bracket (velocity low / retangent high) + bias origin (drift-velocity decay) | -| 6 Experiments | 1.50 | F1-F4 + numbers table; soft ramp (refutation), closed form matched 0.933 / ensemble 0.982 / lam_min 0.987; e0 0.4378 vs 0.4403 + KS; estimator 256-traj MAE 0.00189 corr 0.99885 + stress grid range; MNIST both parts; teacher test-gap sign flip (2 sentences + pointer to F4/appendix) | +| 6 Experiments | 1.50 | F1-F4 + numbers table; soft ramp (refutation), closed form matched 0.933 / ensemble 0.982 / lam_min 0.987; e0 hidden-share calibration (max err <=0.008 over 6x6) + KS; estimator 256-traj MAE 0.00189 corr 0.99885 + stress grid range; MNIST both parts; teacher test-gap sign flip (2 sentences + pointer to F4/appendix) | | 7 Related work | 0.45 | four paragraphs: FA empirics; FA theory (Song-Xu-Lafferty complementarity via `E_B[H_FA]=0`); NTK/lazy; capacity/memorization | | 8 Limitations + conclusion | 0.25 | scope (MLP, full-batch, squared loss; MNIST-subset real data; conv/transformers open); estimator conditional; architecture-only drift = open problem | diff --git a/notes/42_master_technical_reference.md b/notes/42_master_technical_reference.md new file mode 100644 index 0000000..82c436e --- /dev/null +++ b/notes/42_master_technical_reference.md @@ -0,0 +1,697 @@ +# Master Technical Reference + +Self-contained record of every result, with full statements, proofs, +derivations, exact numbers, configs, and negative results. Built for writing +the paper by hand: you should not need any other file open. Cross-checked +against the raw CSVs on 2026-06-13. + +Conventions in this file: +- "no fit" = no post-hoc scalar tuned to the reported data; predictions come + from initialization or two early kernel snapshots. +- Every empirical number cites its output directory; commands are in note 40. +- Numbers flagged **[verify]** could not be reproduced from current outputs + and must be relocated or dropped before they enter the paper. + +Contents: +1. Notation and setup +2. Result A — static alignment law (Thm 1) + capacity cost +3. Result B — multilayer scaling (Prop 1) + observed surprisal +4. Result C — prior-free minimax (Thm 2) +5. Result D — random-subspace null model (Prop 2) + its failure +6. Result E — initial erosion theorem (Thm 3) + Gaussian e0 (Prop 3) +7. Result F — closed-form soft ramp (Thm 4, Cor 1, Cor 2) +8. Result G — finite-time estimator + the lazy-validity chain +9. Negative results (consolidated) +10. Generalization: teacher test gap +11. Related work, exact positioning +12. Open problems +13. Honesty boundaries + +--- + +## 1. Notation and setup + +**Network.** MLP with widths `n_0, n_1, ..., n_L`, forward weights +`W_l in R^{n_l x n_{l-1}}`, ReLU `phi`. Forward pass on input `x`: +`h_0 = x`, `pre_l = W_l h_{l-1}`, `h_l = phi(pre_l)` for hidden layers, +`h_L = pre_L` (linear output). Train on `N` examples `X` with squared loss + +``` +L = (1/2N) ||r||^2, r = f_theta(X) - Y (all residuals stacked, length N n_L). +``` + +**The single FA/BP difference (backward pass).** Write `g_l` for the gradient +of layer `l`. Both rules share the output residual. Propagating error +backward through layer `l+1`: + +``` +BP: delta_l = (delta_{l+1} W_{l+1}) ⊙ 1[pre_l > 0] +FA: delta_l = (delta_{l+1} B_l^T) ⊙ 1[pre_l > 0] +``` + +with `B_l in R^{n_l x n_{l+1}}` a fixed random matrix replacing `W_{l+1}^T`. +The layer gradient is `g_l = delta_l^T h_{l-1}` in both cases. **Key fact used +throughout:** the output layer has no backward matrix above it, so + +``` +g_L^FA = g_L^BP (the output-layer gradient is identical under both rules). +``` + +`D_l = n_l n_{l+1}` is the flattened backward-matrix dimension. +`P = sum_l n_l n_{l-1}` is the parameter count. + +**Learning speed (first-order loss decrease).** A step `-eta g` decreases the +loss to first order by `eta <g^BP, g>`. Define + +``` +speed_BP = sum_l ||g_l^BP||^2, +speed_FA = sum_l <g_l^BP, g_l^FA>, +e_0 = 1 - speed_FA / speed_BP (fraction of BP's speed FA discards in one step). +``` + +`e_0 = 0`: FA as good as BP this step. `e_0 = 1`: FA makes no progress. + +**Tangent operators.** Let `J = d f_theta(X) / d theta` be the output +Jacobian (size `N n_L x P`). Then + +``` +K_BP = J J^T (the NTK; symmetric PSD) +K_FA = J J_tilde^T (J_tilde = FA update Jacobian; NON-symmetric) +S_FA = (K_FA + K_FA^T)/2 (only the symmetric part affects loss to first order) +``` + +Equivalently `speed_BP = r^T K_BP r / (...)` and `speed_FA = r^T S_FA r`, +i.e. `speed`s are residual-direction quadratic forms in these operators. +Residual recursion (gradient descent in output space): +`r_{t+1} = (I - (eta/N) K_t) r_t`. + +**Output / hidden share.** With `g_out = g_L^BP`, + +``` +rho = output_share = ||g_out^BP||^2 / speed_BP = r^T K_out r / r^T K_BP r in (0,1), +hidden_share = 1 - rho. +``` + +`rho` is the single scalar that drives the gap results. It is read off one BP +backward pass. + +--- + +## 2. Result A — static alignment law (Thm 1) + +### Statement + +Let `A_l = W_{l+1}^T`, and let `A_l/||A_l||_F`, `B_l/||B_l||_F` be independent +isotropic unit directions in `R^{D_l}`. Define the squared Frobenius cosine + +``` +Q_l = <A_l, B_l>_F^2 / (||A_l||_F^2 ||B_l||_F^2). +``` + +Then + +``` +Q_l ~ Beta(1/2, (D_l - 1)/2), E[Q_l] = 1/D_l, D_l Q_l => chi^2_1 as D_l -> inf. +``` + +Tail / survival: `Pr(Q_l >= q) = 1 - I_q(1/2, (D_l-1)/2)` (regularized +incomplete Beta). High-dimensional bound `Pr(Q_l >= q) <~ 2 exp[-(D_l-1)q/2]`. + +### Proof + +Rotational invariance lets us fix `a = e_1`, so `Q = b_1^2` with `b` uniform on +the sphere `S^{D-1}`. The squared first coordinate of a uniform unit vector is +exactly `Beta(1/2, (D-1)/2)`. (Equivalently `Q = X/(X+Y)`, `X ~ chi^2_1`, +`Y ~ chi^2_{D-1}` independent.) ∎ + +### Reading + +Typical cosine `~ 1/sqrt(D)`. For any realistic layer (`D = n^2` in the +thousands) a random feedback matrix is essentially orthogonal to the one BP +wants. The theorem says exactly how orthogonal, with no slack. + +### Capacity cost + +For target alignment `q`, the log-volume cost + +``` +C_l(q) = -log Pr(Q_l >= q). +``` + +High-D: `C_l(q) ≈ ((D_l-1)/2) log(1/(1-q)) = Theta(D_l)`. Small-q: +`C_l(q) ≈ D_l q / 2`. (Natural logs.) + +### Validation (output: `capacity_empirical_validation/`, 100k samples/dim, 700k total) + +Rotational-invariance sampling `Q = z_1^2/||z||^2`, `z ~ N(0,I_D)`: + +| D | emp mean | theory mean | KS stat | KS p | emp q99 | theory q99 | +|---|---|---|---|---|---|---| +| 64 | 0.015679 | 0.015625 | 0.00349 | 0.173 | 0.10071 | 0.10071 | +| 128 | 0.0078528 | 0.0078125 | 0.00197 | 0.833 | 0.051165 | 0.051097 | +| 256 | 0.0039289 | 0.0039063 | 0.00296 | 0.345 | 0.025990 | 0.025733 | +| 512 | 0.0019553 | 0.0019531 | 0.00164 | 0.950 | 0.012965 | 0.012913 | +| 1024 | 0.00097059 | 0.00097656 | 0.00312 | 0.283 | 0.0063694 | 0.0064679 | +| 2048 | 0.00048817 | 0.00048828 | 0.00311 | 0.288 | 0.0032339 | 0.0032368 | +| 4096 | 0.00024441 | 0.00024414 | 0.00243 | 0.597 | 0.0016322 | 0.0016191 | + +Means/quantiles match to 3-4 digits; KS never rejects (all p > 0.17). +Capacity-tail calibration (same dir): max smoothed cost error 0.20 nats, mean +0.019 nats (note 02). + +--- + +## 3. Result B — multilayer scaling (Prop 1) + observed surprisal + +### Statement and proof + +Independent layer blocks => alignment events independent, so probabilities +multiply and costs add: + +``` +p_all = prod_l Pr(Q_l >= q_l), C_all = sum_l C_l(q_l). +``` + +Equal width (`D_l ≈ n^2`), fixed `q`: `C_all = Theta(L n^2)`, +`Pr(all q_l) = exp[-Theta(L n^2)]`. Chance threshold `q = c/D_l`: +`C_l ≈ c/2`, `C_all = Theta(L)`. Proof is the product rule for independent +events plus the `C_l = Theta(D_l)` evaluation from Result A. ∎ + +### Observed surprisal (the distributional companion) + +`S_l = -log Pr(Q_l' >= Q_l)` with `Q_l'` an independent draw. By the +probability integral transform `Pr(Q_l' >= Q_l) ~ Uniform(0,1)`, so +`S_l ~ Exp(1)` and for `L` independent layers `sum_l S_l ~ Gamma(L, 1)`. The +transformed surprisal law is dimension-free; dimension enters only through the +raw `Q_l` scale and `C_l(q)`. + +### Validation (output: `multilayer_capacity_distribution/`, 100k/(D,L)) + +`D in {64,256,1024,4096}` x `L in {1,2,4,8,16}`: **max KS 0.0042**, mean abs +mean-error 0.0042, mean abs var-error 0.040. E.g. D=64,L=16: emp mean 16.008 +vs 16, KS 0.0025; D=4096,L=1: 1.0004 vs 1, KS 0.0019. Multilayer product +capacity (note 02): max err 0.46 nats, mean 0.039 nats. + +--- + +## 4. Result C — prior-free minimax (Thm 2) + +### Statement + +Let `b_hat = vec(B)/||B||_F in S^{D-1}` drawn from any distribution `mu`, +`M_mu = E_mu[b_hat b_hat^T]` (so `tr M_mu = 1`), and `a in S^{D-1}` an unknown +target. Then + +``` +sup_mu inf_{||a||=1} E_mu[(a^T b_hat)^2] = sup_mu lambda_min(M_mu) = 1/D, +``` + +attained by isotropic feedback (`M_mu = I/D`). + +### Proof + +`E_mu[(a^T b_hat)^2] = a^T M_mu a`; its worst case over `a` is +`lambda_min(M_mu) <= tr(M_mu)/D = 1/D` (min eigenvalue <= average eigenvalue). +Isotropic `mu` gives `M_mu = I/D`, so every direction equals `1/D`, achieving +the bound. ∎ + +### Reading and prior-aware corollary + +Anisotropic feedback that aligns better with some directions does so by +aligning worse with others; an adversarial task lands on the bad ones. +Sparsity / orthogonality / low rank change conditioning or hardware cost but +cannot beat `1/D` worst-case. With a target prior `Sigma_A = E[a a^T]`, +`E_{a,B}[(a^T b_hat)^2] = tr(Sigma_A M_mu)`, so with prior information +top-eigenspace feedback can win; without it, isotropy is minimax. Note the +average-case over uniform targets is `1/D` for every `mu` — the distinction +between initializations is worst-case coverage, not mean. + +### Validation (output: `minimax_initialization/`, D=32, 1/D = 0.03125) + +isotropic `lambda_min = 0.029084`; rademacher 0.028947; anisotropic 0.006150; +rank-deficient subspace 0; axis-aligned 0. Coverage-distribution matching for +each scheme in `initialization_distribution_matching/` (D=128, worst KS +0.0047 on non-degenerate schemes). + +--- + +## 5. Result D — random-subspace null model (Prop 2) + its failure + +### Statement and proof + +Toy model: alignment deletes a uniformly random `k`-dim subspace `E` of the +`P` parameters, leaving projected kernel `J Q_E J^T`, `Q_E = I - P_E`. With +`v = J^T r`, one-step erosion `e(r) = ||P_E v||^2 / ||v||^2`. By rotational +invariance `E[P_E] = (k/P) I`, so `E[e(r) | J, r] = k/P`, and for Haar-random +`E` + +``` +e(r) ~ Beta(k/2, (P-k)/2), E[e] = k/P, Pr(e = 0) = 0 for k > 0. +``` + +Proof: `||P_E v||^2/||v||^2` for a fixed `v` and Haar `k`-subspace is the sum +of `k` squared coordinates of a uniform direction, i.e. `Beta(k/2,(P-k)/2)`. ∎ + +### Scaling consequence (the "no free capacity" sell) + +`k` fixed, `P -> inf`: `E[e] -> 0` (burden vanishes). `k/P -> alpha > 0`: +`E[e] -> alpha` (relative erosion persists). Overparameterization dilutes the +cost through `k/P`; it never deletes it, and there is no threshold — +`Pr(e=0)=0` for any `k>0`. Role: this is the clean mechanism that motivates +the operator object. **It is a null model, not a gap predictor.** + +### Negative result: hard-k overpredicts real FA (output: `soft_erosion_theory_vs_empirical/`) + +Using `rho ~ Beta((P-k)/2, k/2)` and propagating the BP spectrum overpredicts +the measured gap badly, worst on the over-capacity side: + +| width | margin | hard-k theory mean | empirical mean | +|---|---|---|---| +| 40 | +130 | 0.040374 | 0.000162 | +| 32 | +2 | 0.064534 | 0.003729 | +| 24 | -126 | 0.121998 | 0.046464 | +| 20 | -190 | 0.151140 | 0.146803 | + +Up to ~250x at positive margin. Local one-step erosion measured 0.23-0.39 vs +hard-k Beta-mean 0.63-0.73 (~2x). Conclusion: raw constraint count `k` is not +the effective burden `k_eff`; the real object is operator erosion (Result E). + +--- + +## 6. Result E — initial erosion theorem (Thm 3) + Gaussian e0 (Prop 3) + +### Thm 3 (initial erosion = hidden speed share) + +Fix forward weights `W`, data, residual `r`. Assume feedback matrices `B_l` +are independent of `W, r` with zero first moment in every hidden backward +direction, and the output layer uses the true gradient. Then + +``` +E_B[speed_FA | W, r] = ||g_L^BP||^2, +E_B[e_0 | W, r] = 1 - ||g_L^BP||^2 / sum_l ||g_l^BP||^2 = 1 - rho =: hidden_share. +``` + +**Proof.** The output FA gradient equals the BP gradient, contributing +`||g_L^BP||^2`. Each hidden FA gradient is multilinear in zero-mean +independent feedback matrices, so `E_B[g_l^FA | W,r] = 0` for hidden `l`, +hence `E_B[<g_l^BP, g_l^FA>] = 0`. Summing leaves only the output term. ∎ + +### Reading + +On step one, random feedback points hidden updates in random directions, so on +average they accomplish nothing; FA gets only the output layer, which it +computes correctly, for free. The expected initial cost is exactly the share +of work the hidden layers were doing — a number from one BP backward pass. +This is the residual-direction version of Song-Xu-Lafferty's `K_FA = G + H_FA` +at init: `E_B[H_FA] = 0`, so `E_B[K_FA] = K_out`, NOT `(1 - k/P) K_BP`. That +is why hard-k was pessimistic. + +### Prop 3 (exact Gaussian e0, one hidden layer) + +One-hidden-layer ReLU MLP `x -> W1 -> ReLU -> W2 -> out`, Gaussian feedback +`B_{a,c} ~ N(0, sigma_B^2)`. For fixed `W, r, gates`, +`speed_FA = ||g_out^BP||^2 + <C(W,r), B>`, linear in `B`, with explicit +coefficient matrix + +``` +S_{a,c,d} = sum_n gate_{n,a} x_{n,d} delta_{n,c}, +C_{a,c} = sum_d g_bp_{a,d} S_{a,c,d}, +<g_hidden^BP, g_hidden^FA(B)> = sum_{a,c} C_{a,c} B_{a,c}, +``` + +(`x_n` input, `delta_{n,c}` output residual gradient, `gate_{n,a}` ReLU gate, +`g_bp_{a,d}` BP hidden gradient). Hence `<C,B> ~ N(0, sigma_B^2 ||C||_F^2)` and + +``` +e_0(B) ~ Normal( 1 - ||g_out^BP||^2/speed_BP, sigma_B^2 ||C||_F^2 / speed_BP^2 ), +``` + +exact conditionally on `(W, r)`, independent of horizon `T`. Deeper FA keeps +the exact mean (Thm 3) but the finite-width law is no longer exactly Gaussian +(hidden terms become products of feedback matrices; a Wick/CLT expansion +applies). + +### Validation — moments (output: `actual_fa_initial_operator_moments/`) + +2-hidden-layer nets, 6 widths x 6 init seeds x 512 real-FA-backward feedback +draws. Predicted erosion = hidden share (no fit). Per-width at init seed 0: + +| width | predicted | empirical | error | +|---|---|---|---| +| 16 | 0.45193 | 0.44435 | -0.0076 | +| 24 | 0.31202 | 0.31229 | +0.0003 | +| 32 | 0.23606 | 0.23971 | +0.0037 | +| 48 | 0.26147 | 0.25986 | -0.0016 | +| 64 | 0.20912 | 0.21126 | +0.0021 | +| 96 | 0.25977 | 0.25915 | -0.0006 | + +Across all 36 rows the prediction sits within Monte-Carlo error of the +measurement: **max |error| ≈ 0.008** (width 24 init 4: +0.0080; width 16 init +0: -0.0076). The predicted erosion itself spans 0.15-0.45 across configs, so +this is a wide-range calibration, not a single point. + +> **[verify]** The headline "E_B[e_0] = 0.4378 ± 0.008 vs 0.4403" cited in +> notes 35/36 and the old draft does NOT appear in the current moments CSV +> (which has no 0.44 row or aggregate). Either relocate its source run or +> replace it with the table above before it enters the paper. The table above +> is the reproducible evidence. + +### Validation — distribution (output: `actual_fa_initial_erosion_distribution/`, post-fix real backward) + +Sampler runs the real FA backward per draw (note 37 fix); the analytic `<C,B>` +form is only an internal check (`linear_identity_max_err <= 8.9e-16`). +4 widths x 3 inits x 4096 draws, standardized then KS-tested vs Normal: + +| width | init | theory mean | emp mean | theory std | emp std | KS | p | +|---|---|---|---|---|---|---|---| +| 16 | 0 | 0.113830 | 0.113413 | 0.027228 | 0.027401 | 0.0150 | 0.314 | +| 16 | 2 | 0.131704 | 0.131796 | 0.028843 | 0.028942 | 0.0056 | 0.999 | +| 32 | 1 | 0.045831 | 0.045870 | 0.007278 | 0.007396 | 0.0094 | 0.858 | +| 32 | 2 | 0.227193 | 0.227691 | 0.031148 | 0.030779 | 0.0166 | 0.206 | +| 64 | 0 | 0.031155 | 0.031119 | 0.003652 | 0.003583 | 0.0118 | 0.609 | +| 128 | 1 | 0.024377 | 0.024353 | 0.001770 | 0.001774 | 0.0182 | 0.130 | + +All 12 panels: KS p in [0.13, 0.999], means/stds match to 3-4 digits, no fit. + +### Validation — MNIST (output: `real_data_validation_mnist/`, note 38) + +Thm 3 is data-agnostic. On MNIST-subset MLPs (784-dim input), 16 configs +(width {64,256} x depth {1,2} x 4 inits), 512 real-backward draws each: +hidden share spans **0.39-0.82**, `|E_B[e_0] - hidden_share| in [0.0001, +0.0089]` (all within ~2 stderr). Examples: depth 1 width 64 init 2: pred +0.8189 / meas 0.8188; depth 2 width 256 init 2: pred 0.3949 / meas 0.3971. +On 784-dim input the hidden (here input-layer-dominated) share is large; the +theorem prices it with no fit. + +--- + +## 7. Result F — closed-form soft ramp (Thm 4, Cor 1, Cor 2) + +### Assumptions + +- **A1 (lazy).** Tangent operators frozen at init: + `r_t^BP = (I - (eta/N) K_BP)^t r_0`. Standard NTK approximation; its + breakdown is the drift the estimator corrects (Result G). +- **A2 (scalar erosion).** `S_FA ≈ rho K_BP`. Exact in the residual direction + by Thm 3 (`r^T E_B[S_FA] r = rho · r^T K_BP r`); the assumption is that the + same slowdown `rho` applies in every eigendirection. + +### Thm 4 (closed-form gap) + +In the BP eigenbasis `K_BP v_i = lam_i v_i`, `c_i = v_i^T r_0`: + +``` +gap_T = L_FA(T) - L_BP(T) + = (1/2N) sum_i c_i^2 [ (1 - eta rho lam_i/N)^{2T} - (1 - eta lam_i/N)^{2T} ]. +``` + +**Proof.** `L_rule(T) = ||r_T^rule||^2/2N`. Expand `r_0 = sum_i c_i v_i`. +A1 gives `||r_T^BP||^2 = sum_i c_i^2 (1 - eta lam_i/N)^{2T}`. A2 makes `S_FA` +share eigenvectors with eigenvalues `rho lam_i`, so +`||r_T^FA||^2 = sum_i c_i^2 (1 - eta rho lam_i/N)^{2T}`. Subtract, divide by +`2N`. ∎ + +### Cor 1 (soft, not a phase transition) + +`rho < 1` => each bracket `>= 0` => `gap_T >= 0`. `gap_T` is real-analytic in +`(rho, {lam_i}, T)`, so deforming capacity smoothly moves the gap smoothly: a +soft ramp, no kink. A discontinuity would need `rho -> 0`, which never happens +since `rho >= output_share > 0`. The hard threshold +`Delta d_hard = max(0, k-(P-d))` is recovered only in that degenerate limit. + +### Cor 2 (the ramp law) + +Underparameterized side (many `tau_i = N/(eta lam_i) >~ T`): each bracket +`≈ 2 eta T (1-rho) lam_i / N`, so `gap_T ≈ (1-rho)(eta T/N^2) sum_slow c_i^2 +lam_i` (residual-mass limited). Converged side (`L_BP ≈ 0`): the gap is the +slowest surviving FA mode, + +``` +gap_T ≈ (c_min^2/2N) exp(-2 eta rho lam_min T / N), +log gap_T ≈ const - (2 eta T/N) rho lam_min. +``` + +The ramp is driven by `lam_min(w)` growing smoothly with capacity. + +### Validation — the dense empirical ramp (output: `phase_transition_dense_T30000_352traj/`) + +Random-label task, full-batch SGD lr 0.01, `T = 30000`, 11 widths 20-40, 1 +init x 32 feedback = 352 trajectories. Measured train gap (`train_gap_mean`): + +| width | FA margin | gap mean | gap std (32 seeds) | +|---|---|---|---| +| 20 | -190 | 0.146803 | 0.042009 | +| 22 | -158 | 0.100526 | 0.034803 | +| 24 | -126 | 0.046464 | 0.024474 | +| 26 | -94 | 0.029671 | 0.023591 | +| 28 | -62 | 0.014195 | 0.009368 | +| 30 | -30 | 0.007466 | 0.004624 | +| 32 | +2 | 0.003729 | 0.003787 | +| 34 | +34 | 0.002079 | 0.002232 | +| 36 | +66 | 0.000791 | 0.000605 | +| 38 | +98 | 0.000553 | 0.000545 | +| 40 | +130 | 0.000162 | 0.000134 | + +Smooth, no kink at margin 0, nonzero at positive margin. `log gap` linear in +margin: **R^2 = 0.993**, slope -0.0209. (History: near-margin-0 gap was 0.297 +at T=3000, 0.051 at T=10^4, 0.0037 at T=3x10^4 — the apparent kink dissolved +as training lengthened, notes 19-22.) + +### Validation — closed form vs measured (output: `closed_form_soft_ramp/`, corrected per note 37) + +Closed form evaluated on the *same data and init seed* the sweep used, plus 4 +extra init seeds (same data) for a no-fit band. All of `rho, lam_i, c_i` +measured at init. + +``` +corr(log pred, log measured), matched single init = 0.933 +corr(log pred, log measured), 5-init geometric mean = 0.982 +corr(lam_min, -log measured), matched init = 0.948 +corr(lam_min mean over inits, -log measured) = 0.987 +ensemble-mean lam_min(w) over the sweep = 0.026 -> 0.763 (monotone) +rho across widths (5-init mean) = 0.68 - 0.79 +``` + +A single init's frozen `lam_min(w)` is noisy across widths while the T=30000 +measured ramp is smooth (drift + feedback averaging smooth it), so matched +single-init corr is 0.933; the init ensemble (still zero fitted parameters) +tracks the ramp at 0.982 and isolates `lam_min` as the driver at 0.987. +Frozen-init compresses the range (under-predicts the large-gap end, +over-predicts the small-gap tail) — exactly the operator drift the estimator +recovers. **The previously reported 0.977 is retired** (it came from a +data/init-mismatched comparison; note 37). + +--- + +## 8. Result G — finite-time estimator + the lazy-validity chain + +The closed form gives mechanism and shape; the estimator gives magnitude. The +path to it is a measured three-step chain. + +### Step 1 — frozen K(0) is exact at small T (output: fixed-width T=5 runs, note 10) + +Fixed width 64, `N in {64..320}`, SGD lr 1e-3, frozen-K(0) residual recursion, +256 trajectories. Horizon diagnostic (empirical gap / predicted gap): + +| T | emp gap | pred gap | ratio | MAE | +|---|---|---|---|---| +| 1 | 0.030734 | 0.029853 | 1.028 | 0.00088 | +| 5 | 0.112912 | 0.113947 | 0.987 | 0.00171 | +| 10 | 0.156304 | 0.164964 | 0.942 | 0.00866 | +| 20 | 0.163335 | 0.183836 | 0.884 | 0.02050 | +| 50 | 0.125547 | 0.154075 | 0.814 | 0.02853 | + +At **T=5: corr 0.99947, MAE 0.00083, mean ratio 0.993**. The operator object +is correct; freezing is a local-time theorem that drifts out of regime by +T=50. + +### Step 2 — the T=50 error is entirely kernel drift (output: `finite_time_kernel_probe_T50_N128_8runs/`, note 13) + +Replace frozen K(0) by the oracle product of measured `K(t)` at every step: + +| predictor | gap MAE | BP loss MAE | +|---|---|---| +| fixed K(0) | 0.027683 | 0.025404 | +| time-varying K(t) | 0.000273 | 0.000237 | + +Drift `K_t - K_0` removes ~99% of the gap error. So the finite-time problem is +to model `K_t`, not to fit a scalar; higher-order output-Taylor terms are +negligible. + +### The estimator (output: `compressed_operator_s20_256traj_T50_width64_plots/`, notes 15-17) + +Read kernels at init and one early step `s`, extrapolate linearly, roll out: + +``` +K_hat_t = K_0 + (t/s)(K_s - K_0), r_hat_{t+1} = (I - (eta/N) K_hat_t) r_hat_t. +``` + +Only inputs: `(K_0, K_s)` per rule. No fitted scale/offset/calibration. +256 trajectories, T=50, s=20: + +| predictor | MAE | bias | corr | +|---|---|---|---| +| fixed K(0) | 0.018487 | +0.018487 | 0.984460 | +| FA compressed | 0.014915 | +0.014915 | 0.990472 | +| **linear velocity** | **0.0018934** | **-0.0014912** | **0.998848** | +| early re-tangent | 0.0048684 | +0.0048684 | 0.999077 | + +Linear velocity beats fixed K(0) ~10x. `linear (low) <= empirical <= +retangent (high)` is a no-fit bracket. + +### Stress grid (output: `operator_stress_grid_summary/`, note 16) + +Random-label, SGD lr 1e-3; 5 N-values x 8 traj per new setting: + +| setting | rows | fixed MAE | linear MAE | retangent MAE | linear bias | corr | +|---|---|---|---|---|---|---| +| d1,w64,T50,s20 | 40 | 0.003862 | 0.000233 | 0.001162 | -0.000233 | 0.999891 | +| d2,w32,T50,s20 | 40 | 0.047110 | 0.005652 | 0.013333 | -0.005375 | 0.994077 | +| d2,w64,T25,s10 | 40 | 0.020808 | 0.002722 | 0.005084 | -0.002668 | 0.999074 | +| d2,w64,T50,s20 | 256 | 0.018487 | 0.001893 | 0.004868 | -0.001491 | 0.998848 | +| d2,w64,T100,s40 | 40 | 0.034378 | 0.002647 | 0.008100 | -0.002169 | 0.998653 | +| d2,w96,T50,s20 | 40 | 0.016912 | 0.001794 | 0.004004 | -0.001551 | 0.998634 | +| d3,w64,T50,s20 | 40 | 0.051670 | 0.005049 | 0.012450 | -0.004916 | 0.998546 | + +Holds across depth {1,2,3}, width {32,64,96}, horizon {25,50,100}: **linear +corr 0.994-0.99989**. Hard cases (w32, d3) raise drift magnitude but velocity +still cuts fixed-K(0) error ~10x. + +### Bias structure (note 17) + +Residual gap bias ≈ -0.0021, dominated by BP-loss overprediction (+0.0026); +FA near-unbiased (+0.0005). Cause: drift velocity decays after `s` +(`alpha_t < t/s`; e.g. true `alpha_50` = 1.46 BP / 1.17 FA vs linear 2.5). +This is kernel-path curvature, not a normalization error (the same recursion +was exact at T=5 and the oracle K(t) matched to 1e-4). No fitted correction is +introduced. + +### MNIST estimator (output: `real_data_validation_mnist/`, note 38) + +Width {128,256} x N {128,256,512} x 2 inits x 8 feedbacks = 96 traj, T=50, +s=20, lazy step `eta lam_max/N = 0.05`: + +| predictor | MAE | bias | corr | +|---|---|---|---| +| fixed K(0) | 0.00613 | +0.00529 | 0.606 | +| **linear velocity** | **0.00196** | **-0.00062** | **0.911** | + +Velocity MAE = **3.9% of the mean gap**, near-zero bias, beats fixed K(0) +everywhere. The gap only spans 0.040-0.063 across all 96 traj, so MAE/bias are +the operative stats and corr 0.911 is the weaker one (report it with the +range). At `eta lam_max/N = 0.3` the lazy chain breaks (fixed MAE 0.069 corr +0.37, velocity 0.030 corr 0.68) — the A1 scope boundary, recorded as a datum. + +### Computational tool (kernel via gram identity, `real_data_validation.py`) + +Tangent kernels are computed by the layerwise identity (avoids the `N n_L x P` +Jacobian): + +``` +K[(n,c),(m,c')] = sum_l <delta_l(n,c), delta'_l(m,c')> · <h_{l-1}(n), h_{l-1}(m)>, +``` + +`delta'` = BP deltas for `K_BP`, FA deltas for `K_FA`. `--self-test` verifies +this against an explicit autograd Jacobian: max error <= 5.3e-15 for +`K_BP = J J^T`, `K_FA = J J_tilde^T`, and `J^T r` vs the production gradient. + +--- + +## 9. Negative results (consolidated — these justify the final objects) + +| approach | result | output / note | +|---|---|---| +| hard-k as gap predictor | overpredicts 2x-250x, worst at positive margin | `soft_erosion_theory_vs_empirical/` (25) | +| infinitesimal derivative `K_0 + t·dK_0` extrapolated to T=50 | MAE **270** vs 0.024 baseline (diverges) | `fa_tangent_hierarchy_derivative_probe_eps005/` (31) | +| scalar directional-gain predictor across N | corr **0.124** (vs operator 0.999) | `early_directional_predictors_*` (14) | +| trajectory bridge on long horizons | shape corr 0.97 but 2-3x scale error | (02, 08) | +| phase-transition (kink) hypothesis | dissolves as T grows; dense T=30000 smooth | (19-22) | + +Each is one sentence in the paper. The derivative blow-up (31) and hard-k +overprediction (25) most directly motivate, respectively, the bounded +two-snapshot velocity estimator and the move from constraint-count to operator +erosion. + +--- + +## 10. Generalization: teacher test gap (output: `teacher_test_gap_sgd_T8000/`, note 39) + +MLP teacher (width 64, 2 hidden, noiseless, normalized targets), N=256 train / +2048 test, SGD lr 0.01, T=8000, widths 8-96, 3 init x 8 feedback = 24 FA +runs/width. + +| width | train gap | test gap (mean ± stderr) | +|---|---|---| +| 8 | 0.249 | +0.177 ± 0.020 | +| 12 | 0.158 | +0.038 ± 0.016 | +| 16 | 0.112 | -0.038 ± 0.015 | +| 24 | 0.080 | -0.044 ± 0.018 | +| 32 | 0.085 | -0.120 ± 0.013 | +| 48 | 0.048 | -0.014 ± 0.014 | +| 64 | 0.026 | +0.010 ± 0.008 | +| 96 | 0.0069 | +0.045 ± 0.011 | + +Train (optimization) gap: positive, monotone-decreasing — the invariant soft +ramp. Test gap: changes sign — FA generalizes significantly *better* at +intermediate width (-0.120 ± 0.013 at w=32, an implicit-regularization effect) +and settles slightly positive once both fit. The optimization gap is the +object the theory prices; its test-side consequence is task-dependent, which +is exactly why the theory is built on the former. Reported as scoping evidence, +not a contribution. Caveat: single task family, one train size; the sign-flip +location moves with task details. + +--- + +## 11. Related work, exact positioning + +- **Lillicrap et al. 2016** (FA works; alignment emerges), **Nokland 2016** + (DFA, deep), **Refinetti et al. 2021** (align-then-memorize). They answer + "can FA learn?"; we answer "what does it cost?". Refinetti's alignment phase + is exactly our drift `K_s - K_0`. +- **Song, Xu, Lafferty 2021** — two-layer FA converges (over-param), alignment + optional without regularization; decomposition `K_FA = G + H_FA`, `H_FA` + non-PSD and small at init. Complement, not conflict: at init their `H_FA` is + our Thm 3 (`E_B[H_FA] = 0`, residual direction). Convergence to zero *final* + error and a nonzero *operator* gap are different statements. Phrase as + complement; never as a correction. +- **Jacot et al. 2018** (NTK), **Lee et al. 2019** (wide nets as linear + models), **Chizat-Oyallon-Bach 2019** (lazy training) license the residual + recursion and A1; our contribution there is FA's non-symmetric operator. + **Huang-Yau 2020** (neural tangent hierarchy) is the `K_t`-dynamics program + our estimator instantiates empirically. +- **Zhang et al. 2017** (fitting random labels) invites the "capacity absorbs + it" guess; Result D/E is the soft rebuttal. **Montanari-Zhong 2022** + (interpolation transition) is why a kink was a reasonable prior. + +Rule: every citation earns a clause tied to one of our results. Verify each +against the source before it lands (style contract). + +--- + +## 12. Open problems (state, do not solve) + +Make the finite-time gap architecture-only: (i) predict `rho` (hence +`hidden_share`) from infinite-width NTK recursions; (ii) predict the drift +`K_s - K_0` (alignment gain) from initialization — naive `dot K_0` fails (note +31), so a bounded short-time alignment-gain path or deep-linear / two-layer +order-parameter solution; (iii) a per-mode spectral erosion `rho_i` refining +the scalar A2. + +--- + +## 13. Honesty boundaries (what NOT to claim) + +- Optimization (train) gap only; not test accuracy (Result 10 shows the test + effect is task-dependent and can favor FA). +- Estimator is snapshot-conditional on `(K_0, K_s)`, no fitted parameters — + never "architecture-only". State its lazy-regime scope (MNIST eta-sensitivity). +- Hard-k is a null model and a reported failure, never a gap predictor. + `E[e] = k/P` (null) and `rho = 1 - E_B[e_0]` (measured share) are distinct. +- A1/A2 stated inline with Thm 4; A2 exact only in the residual direction. +- Soft ramp refutes one named hypothesis (redundancy-exhaustion threshold) via + analyticity of `gap_T`; it is not a universal claim about FA. +- Scope: MLP, full-batch, squared loss, synthetic + MNIST-subset. Conv / + transformers / stochastic optimization / real large-scale data are open. +- Closed-form correlation is 0.933 (matched) / 0.982 (ensemble); 0.977 retired. +- The "0.4378/0.4403" e0 headline is **[verify]** — use the moments table + (Result E) instead until its source is relocated. diff --git a/notes/README.md b/notes/README.md index ee608b6..853c122 100644 --- a/notes/README.md +++ b/notes/README.md @@ -7,10 +7,11 @@ capacity & operator-gap analysis)** — target AAAI-27. | doc | role | |---|---| +| [42_master_technical_reference.md](42_master_technical_reference.md) | **write the paper from this**: every theorem with full statement + proof, every result with exact numbers/configs, derivations, negative results, honesty boundaries -- self-contained | | [41_paper_plan.md](41_paper_plan.md) | **the plan**: venue rules, narrative, section/page budget, theorem & figure numbering, related-work list, schedule, risks | -| [36_evidence_ledger.md](36_evidence_ledger.md) | **the numbers**: every claim with its corrected headline figure, claim discipline, booby traps | +| [36_evidence_ledger.md](36_evidence_ledger.md) | **the numbers, condensed**: every claim with its corrected headline figure, claim discipline, booby traps | | [40_reproduction_manifest.md](40_reproduction_manifest.md) | **the runs**: claim -> exact command -> output dir -> note; everything not listed is not citable | -| [37_wrapup_corrections.md](37_wrapup_corrections.md) | audit record: e0 sampler de-circularized; closed-form comparison matched (0.977 retired -> 0.933/0.982) | +| [37_wrapup_corrections.md](37_wrapup_corrections.md) | audit record: e0 sampler de-circularized; closed-form matched (0.977 retired -> 0.933/0.982); e0 "0.4378/0.4403" retired | ## Authoritative note per result @@ -20,7 +21,7 @@ capacity & operator-gap analysis)** — target AAAI-27. | P1 multilayer additivity, Theta(Ln^2), Gamma(L,1) | 01, 02 | proven + validated | | T2 prior-free minimax 1/D | 01, 02 | proven + validated | | P2 random-subspace null model (and its honest failure as a predictor) | 26 (theorem), 25 (failure) | proven; overpredicts real FA | -| T3 initial erosion = hidden BP speed share | 29 | proven + validated (synthetic 0.4378/0.4403; MNIST note 38) | +| T3 initial erosion = hidden BP speed share | 29, 42 | proven + validated (synthetic max err <=0.008 over 6x6; MNIST note 38) | | P3 exact Gaussian e0 (one hidden layer) | 32 + fix in 37 | proven + validated (real backward) | | T4 closed-form soft-ramp gap law | 34 + correction in 37 | proven under A1+A2; corr 0.933 matched / 0.982 ensemble | | early-velocity estimator (+ bracket, bias structure) | 15, 16, 17 | no-fit, snapshot-conditional; MAE 0.0019, corr 0.999 | @@ -67,9 +68,10 @@ Phase 5 — exact theory of the burden: 33 complete markdown draft (appendix source); 34 **closed-form law T4**; 35 consolidated snapshot. -Phase 6 — wrap-up for AAAI-27 (2026-06-09): +Phase 6 — wrap-up for AAAI-27 (2026-06-09..13): - 36 evidence ledger; 37 corrections; 38 **MNIST validation**; 39 **teacher - test gap**; 40 reproduction manifest; 41 paper plan. + test gap**; 40 reproduction manifest; 41 paper plan; 42 **master technical + reference** (full proofs + numbers, the write-from doc). ## Note conventions |
