diff options
| author | YurenHao0426 <Blackhao0426@gmail.com> | 2026-06-05 15:57:01 -0500 |
|---|---|---|
| committer | YurenHao0426 <Blackhao0426@gmail.com> | 2026-06-05 15:57:01 -0500 |
| commit | a3f6c103678e0dcae3a682945c784a3f88d5039f (patch) | |
| tree | 8658c5715172ba257b69c05fa5ca6b587a842389 | |
| parent | 2df52d3dd7b35053d4a357c3814d7cba6db602ce (diff) | |
Review methods for FA operator dynamics
| -rw-r--r-- | notes/18_contribution_roadmap.md | 28 | ||||
| -rw-r--r-- | notes/30_literature_methods_for_fa_operator_dynamics.md | 634 |
2 files changed, 661 insertions, 1 deletions
diff --git a/notes/18_contribution_roadmap.md b/notes/18_contribution_roadmap.md index aa26caf..89b578b 100644 --- a/notes/18_contribution_roadmap.md +++ b/notes/18_contribution_roadmap.md @@ -230,7 +230,33 @@ This separation avoids overclaiming while still giving a coherent theory chain. ## Next Priority -The weakest current visual was the phase-transition contribution; the dense +The weakest current theory gap is not the static capacity law anymore. It is +the time-dependent operator dynamics: + +```text +e_0 is solved by the actual-FA initial moment theorem +e_t for t > 0 still needs alignment-gain dynamics +``` + +The literature review in `notes/30_literature_methods_for_fa_operator_dynamics.md` +suggests the most principled route is an FA tangent hierarchy: + +```text +K_FA,t = J_t J_tilde_t^T +dK_FA,t/dt = dJ_t/dt J_tilde_t^T + J_t dJ_tilde_t^T/dt +``` + +with a short-time approximation: + +```text +K_FA,t ≈ K_FA,0 + t dot_K_FA,0 +``` + +This formalizes the successful time-varying operator estimator. A lower-risk +full-time theorem can be developed in deep linear FA or in a two-layer +teacher-student order-parameter setting. + +The weakest current visual was the old phase-transition contribution; the dense long-training sweep now shows that the correct phenomenon is a soft ramp, not a hard kink. diff --git a/notes/30_literature_methods_for_fa_operator_dynamics.md b/notes/30_literature_methods_for_fa_operator_dynamics.md new file mode 100644 index 0000000..5c3e75e --- /dev/null +++ b/notes/30_literature_methods_for_fa_operator_dynamics.md @@ -0,0 +1,634 @@ +# Literature Review: Methods for FA Operator Dynamics + +Question: + +```text +Can we move from the exact initial erosion theorem e_0 to a predictive theory +for e_t or finite-time FA/BP gap? +``` + +This note reviews relevant methods and connects them to our estimator evidence. + +## Current State + +We now have an actual-FA initialization theorem: + +```text +speed_BP = sum_l ||g_l^BP||^2 +speed_FA = sum_l <g_l^BP, g_l^FA> + +E_B[speed_FA | W,r] = ||g_output^BP||^2 + +E_B[e_0 | W,r] + = 1 - ||g_output^BP||^2 / sum_l ||g_l^BP||^2 +``` + +This is not a fit. It follows from zero-mean feedback: hidden FA gradients have +zero conditional mean at initialization, while the output-layer FA gradient is +exactly the BP gradient. + +The unresolved part is: + +```text +what happens after W_t becomes dependent on B? +``` + +At later time: + +```text +speed_FA,t = output_speed_t + hidden_alignment_gain_t +``` + +where: + +```text +hidden_alignment_gain_t = sum_hidden <g_l^BP(t), g_l^FA(t)> +``` + +At `t=0`, this hidden gain has zero conditional mean. During training it can +become positive because the forward weights become statistically coupled to the +fixed feedback matrices. + +## Literature Takeaways + +### 1. NTK and frozen-kernel dynamics + +Jacot et al. introduced the Neural Tangent Kernel and showed that, in the +infinite-width limit, the function follows kernel gradient dynamics and the NTK +stays constant during training. + +Source: + +```text +https://papers.nips.cc/paper/8076-neural-tangent-kernel +``` + +Useful for us: + +```text +r_{t+1} ≈ (I - eta K_t / N) r_t +``` + +This justifies our tangent-operator estimator. For BP: + +```text +K_BP = J J^T +``` + +For FA: + +```text +K_FA = J J_tilde^T +``` + +But NTK alone only gives a clean frozen-kernel theorem when the kernel is +approximately constant. + +### 2. Wide networks as linear models / lazy training + +Lee et al. show that wide networks of any depth are governed, in the infinite +width limit, by the first-order Taylor expansion around initialization. They +also report strong finite-width empirical agreement. + +Source: + +```text +https://papers.nips.cc/paper/9063-wide-neural-networks-of-any-depth-evolve-as-linear-models-under-gradient-descent +``` + +Chizat, Oyallon, and Bach frame this as lazy training and provide path-distance +bounds between nonlinear and linearized optimization. + +Source: + +```text +https://papers.neurips.cc/paper_files/paper/2019/hash/ae614c557843b1df326cb29c57225459-Abstract.html +``` + +Useful for us: + +```text +fixed K(0) is a local theorem, not a long-horizon theorem +``` + +This matches our estimator evidence: + +```text +T=1 and T=5: fixed K(0) predicts FA/BP gap almost exactly +T=50 and beyond: fixed K(0) drifts out of regime +``` + +So the correct paper language is: + +```text +local gap distribution: fixed initial FA tangent operator +finite-time gap distribution: time-varying FA tangent operator +``` + +### 3. Neural Tangent Hierarchy for kernel evolution + +Huang and Yau derive a Neural Tangent Hierarchy (NTH), an infinite hierarchy of +ODEs for the evolution of finite-width NTKs. Truncating the hierarchy gives +controlled approximations to NTK dynamics under width/data conditions. + +Source: + +```text +https://proceedings.mlr.press/v119/huang20l.html +``` + +This is the closest existing method for our `t > 0` problem. + +For us, the analogous object is not just NTK: + +```text +K_FA,t = J_t J_tilde_t^T +``` + +Its derivative has the same structure: + +```text +dK_FA,t/dt + = dJ_t/dt J_tilde_t^T + J_t dJ_tilde_t^T/dt +``` + +and both terms depend on higher-order tangent objects. This suggests a direct +method: + +```text +derive an FA tangent hierarchy +truncate at first order: +K_FA,t ≈ K_FA,0 + t dot_K_FA,0 +``` + +This is exactly what our best estimator is doing empirically when it uses early +operator velocity: + +```text +K_hat_t = K_0 + t (K_s - K_0) / s +``` + +The difference is that the estimator measures the velocity, while a theorem +would compute or bound it. + +### 4. Song, Xu, and Lafferty: perturbative FA kernel decomposition + +Song, Xu, and Lafferty analyze two-layer FA and prove overparameterized +convergence. Their key structural decomposition is: + +```text +K_FA = G + H_FA +``` + +where `G` is the positive top-layer kernel and `H_FA` is the hidden +feedback-dependent term. They explicitly treat `H_FA` as non-PSD and +perturbative in the wide regime. + +Source: + +```text +https://papers.nips.cc/paper/2021/file/a576eafbce762079f7d1f77fca1c5cc2-Paper.pdf +``` + +Useful for us: + +Our initial theorem is the residual-direction version of their decomposition: + +```text +E_B[H_FA,0 | W,r] = 0 +``` + +Their proof style suggests a way to bound finite-time FA cost: + +```text +speed_FA,t = r_t^T G_t r_t + r_t^T H_FA,t r_t +``` + +Then: + +```text +operator erosion = 1 - speed_FA,t / speed_BP,t +``` + +If `H_FA,t` remains perturbative, FA can still converge because `G_t` is +positive, but it need not be BP-equivalent. This gives a rigorous way to say: + +```text +convergence does not imply zero BP/FA operator gap +``` + +without contradicting their theorem. + +### 5. Refinetti et al.: align-then-memorise order parameters + +Refinetti et al. derive analytical dynamics for DFA in shallow nonlinear +teacher-student settings and identify an alignment phase followed by a +memorisation phase. They also analyze deep linear alignment matrices and +conditioning. + +Source: + +```text +https://arxiv.org/abs/2011.12428 +``` + +Useful for us: + +They show the right finite-time state variables are low-dimensional alignment +order parameters, not raw parameter counts. Their theory predicts learning +curves by ODEs over overlaps. + +This suggests an architecture-level route: + +```text +derive ODEs for: + residual norm + output BP speed share + hidden alignment gain + layerwise B-W overlaps +``` + +Then: + +```text +e_t = 1 - (output_speed_t + hidden_alignment_gain_t) / speed_BP,t +``` + +This would be more faithful than hard-k capacity exhaustion. + +### 6. Deep linear exact dynamics as a solvable testbed + +Saxe, McClelland, and Ganguli solve learning dynamics in deep linear networks +and show rich nonlinear learning phenomena despite linear input-output maps. + +Source: + +```text +https://arxiv.org/abs/1312.6120 +``` + +Useful for us: + +Deep linear FA may be the cleanest place to get a full time-dependent theorem: + +```text +closed dynamics for W_t and B-induced alignment +closed e_t or gap_T formula in singular modes +``` + +This would not replace nonlinear MLP experiments, but it could give a rigorous +finite-time bridge. + +## What Our Estimator Experience Says + +Our experiments already rule out one bad path: + +```text +hard-k random-subspace model does not predict real FA gap +``` + +It overpredicts positive-margin gaps and treats FA as if it deletes parameter +directions. Real FA instead keeps an output-layer kernel and gradually learns +hidden alignment. + +Our experiments support this path: + +```text +initial actual FA moment theorem + -> e_0 is hidden BP speed share + +early operator velocity + -> predicts finite-time gap much better than fixed K(0) + +time-varying operator product + -> correct object for larger T +``` + +Empirical signs: + +```text +fixed K(0), small T: accurate +fixed K(0), larger T: BP estimated too strong, FA estimated too weak +``` + +Interpretation: + +```text +BP: finite-width feature drift makes initial BP kernel too optimistic +FA: alignment gain makes initial FA operator too pessimistic +``` + +Therefore the theory should model: + +```text +K_BP,t drift +K_FA,t alignment gain +``` + +not just initial random geometry. + +## Candidate Methods + +### Method A: FA tangent hierarchy + +Define: + +```text +K_FA,t = J_t J_tilde_t^T +S_FA,t = (K_FA,t + K_FA,t^T) / 2 +e_t = 1 - (r_t^T S_FA,t r_t) / (r_t^T K_BP,t r_t) +``` + +Derive: + +```text +dK_FA,t/dt + = dJ_t/dt J_tilde_t^T + J_t dJ_tilde_t^T/dt +``` + +Both terms are higher-order tangent objects, exactly like NTH. + +Short-time theorem target: + +```text +K_FA,t = K_FA,0 + t dot_K_FA,0 + error +``` + +and: + +```text +gap_T = fixed-K gap + first-order drift correction + bounded remainder +``` + +This is the most principled extension of our estimator. + +Pros: + +```text +closest to our current successful estimator +no fitted scalar +general MLP-compatible +can become a theorem with smoothness/lazy assumptions +``` + +Cons: + +```text +heavy notation +dot_K_FA,0 is a third-order object +full distribution over B may require Wick/CLT approximations +``` + +### Method B: Song-style `G + H` perturbation + +Decompose: + +```text +K_FA,t = K_output,t + H_FA,t +``` + +where: + +```text +K_output,t = output-layer BP kernel +H_FA,t = hidden feedback-dependent contribution +``` + +At initialization: + +```text +E_B[r_0^T H_FA,0 r_0 | W,r] = 0 +``` + +At time `t`, alignment appears through: + +```text +r_t^T H_FA,t r_t +``` + +The theorem target: + +```text +finite-time FA/BP erosion + = hidden BP speed share + - hidden alignment gain / speed_BP +``` + +Then bound or approximate the hidden alignment gain. + +Pros: + +```text +simple conceptual language +directly compatible with Song/Xu/Lafferty +explains convergence without BP equivalence +``` + +Cons: + +```text +may give bounds more easily than exact distribution +deep nonlinear case still hard +``` + +### Method C: order-parameter ODEs + +Follow Refinetti-style teacher-student analysis. + +Pick a controlled task: + +```text +Gaussian inputs +teacher network labels +two-layer or deep linear student +large input dimension +finite hidden width +``` + +Track overlaps: + +```text +student-teacher overlaps +student-student overlaps +feedback-output overlaps +layerwise alignment matrices +``` + +Then compute: + +```text +e_t and gap_T from ODE trajectories +``` + +Pros: + +```text +architecture-level prediction +can predict full learning curves +natural distribution over B and initialization +``` + +Cons: + +```text +narrower task setting +harder to match our random-label memorization experiments +less general than tangent-operator estimator +``` + +### Method D: deep-linear exact FA dynamics + +Use linear networks as the tractable theorem setting. + +For deep linear networks: + +```text +f(x) = W_L ... W_1 x +``` + +FA replaces backward products by fixed random feedback products. Because gates +are absent, the dynamics can be written as matrix ODEs. + +Target: + +```text +derive exact or perturbative e_t +show e_0 is output-speed share +show alignment gain grows from B-W coupling +compare BP/FA finite-time loss in singular modes +``` + +Pros: + +```text +rigorous and readable +connects to Saxe-style dynamics and Refinetti deep-linear alignment +``` + +Cons: + +```text +not the final nonlinear MLP setting +may be viewed as a sanity theorem unless paired with MLP experiments +``` + +## Recommended Plan + +Use a three-level theory stack. + +### Level 1: exact local theorem + +Keep the current exact theorem: + +```text +E_B[e_0 | W,r] = hidden BP speed share +``` + +Add conditional variance/distribution for the simplest case: + +```text +one-hidden-layer Gaussian feedback: +speed_FA = output_speed + Gaussian hidden term +``` + +This gives a theory-predicted distribution graph, not just a mean. + +### Level 2: short-time FA tangent hierarchy + +Define: + +```text +dot_K_FA,0 = derivative of J J_tilde^T under the FA update +``` + +Then prove a short-time expansion: + +```text +K_FA,t ≈ K_FA,0 + t dot_K_FA,0 +``` + +and propagate residuals with the corresponding product. This formalizes our +successful early-operator estimator. + +If computing `dot_K_FA,0` exactly is too heavy, use `K_s - K_0` as an estimator +and label it correctly: + +```text +conditional tangent-hierarchy estimator +``` + +### Level 3: controlled full-time theory + +For a theorem-only full-time result, use either: + +```text +deep linear FA +``` + +or: + +```text +two-layer teacher-student FA/DFA order-parameter ODE +``` + +This can show explicitly how hidden alignment gain grows and why final +convergence does not imply zero finite-time BP/FA gap. + +## What Not To Do + +Do not return to: + +```text +hard margin predicts exact transition +``` + +The dense T=30000 experiment showed a soft ramp, not a hard phase transition. + +Do not introduce: + +```text +post-hoc scalar correction +``` + +The paper should use either: + +```text +exact distribution / exact moment +``` + +or: + +```text +conditional estimator with observed early operator information +``` + +Do not claim: + +```text +architecture alone predicts all finite-time FA/BP gaps +``` + +unless we derive an ODE or hierarchy closure that predicts `K_FA,t`. + +## Best Current Claim + +The cleanest claim after the literature review is: + +```text +Random feedback has an exact, nonzero initial operator erosion equal to the +hidden BP speed share. At finite time, feedback alignment is precisely the +growth of a hidden alignment-gain term. Existing NTK hierarchy and +align-then-memorise theory suggest two principled ways to model this gain: +short-time tangent-operator hierarchy or low-dimensional order-parameter ODEs. +Our current time-varying operator estimator is the empirical version of the +first route and should be presented as such, not as a fitted curve. +``` + +This makes the theory stack honest: + +```text +e_0 theorem +short-time tangent hierarchy / estimator +controlled ODE or deep-linear full-time theory +large-scale MLP trajectory validation +``` |
