diff options
Diffstat (limited to 'notes/18_contribution_roadmap.md')
| -rw-r--r-- | notes/18_contribution_roadmap.md | 182 |
1 files changed, 182 insertions, 0 deletions
diff --git a/notes/18_contribution_roadmap.md b/notes/18_contribution_roadmap.md new file mode 100644 index 0000000..179bd5d --- /dev/null +++ b/notes/18_contribution_roadmap.md @@ -0,0 +1,182 @@ +# Contribution Roadmap + +Working title: + +`Distributional Capacity Bounds for Feedback Alignment in Multilayer Perceptrons` + +The current paper should use five contributions. The fifth contribution is not +extra decoration: it separates structural capacity theory from finite-time +loss-gap prediction. + +## 1. Capacity Formalization + +We define the capacity burden induced by feedback alignment. + +Layerwise matrix alignment: + +```text +Q_l = cos^2(W_{l+1}^T, B_l) +``` + +For isotropic random feedback in dimension `D_l`: + +```text +Q_l ~ Beta(1/2, (D_l - 1)/2) +``` + +Alignment threshold `q_l` gives a log-volume capacity cost: + +```text +C_l(q_l) = -log P(Q_l >= q_l) +``` + +This formalizes what it means for feedback alignment to consume parameter +direction volume. + +## 2. Scaling Law and Redundancy Exhaustion + +Across independent layers, raw feasible volume multiplies while log-capacity +cost adds: + +```text +p_all = product_l p_l +C_all = sum_l C_l +``` + +For an equal-width MLP with width `n`, depth `L`, and fixed alignment threshold: + +```text +C_all = Θ(L n^2) +p_all = exp[-Θ(L n^2)] +``` + +Functional loss does not need to appear immediately. If total parameter +dimension is `P`, task dimension is `d`, and alignment imposes `k` generic +constraints, then hard functional rank loss is: + +```text +Δd_hard = max(0, k - (P - d)) +``` + +So the FA/BP gap should open when redundant directions are exhausted. + +This is the phase-transition contribution. + +## 3. Prior-Free Minimax Initialization Bound + +Let `b` be the normalized feedback direction in `D` dimensions and `a` be the +unknown normalized target backward direction. + +For any feedback initialization distribution `μ`, define: + +```text +M_μ = E[b b^T] +trace(M_μ) = 1 +``` + +Then: + +```text +inf_a E_μ[(a^T b)^2] = λ_min(M_μ) <= 1/D +``` + +Therefore: + +```text +sup_μ inf_a E_μ[(a^T b)^2] = 1/D +``` + +Isotropic random feedback reaches the bound. Without a prior over `W` or task +directions, no initialization can beat isotropic random feedback in worst-case +expected squared alignment. + +## 4. Tangent-Operator Loss/Gap Estimator + +This is the bridge from capacity regime to finite-time loss gap. + +For BP squared loss, standard tangent-kernel residual dynamics gives: + +```text +r_{t+1}^{BP} ≈ (I - η K_t^{BP}/N) r_t +K_t^{BP} = J_t J_t^T +``` + +For FA, the parameter update uses a surrogate backward Jacobian `J_tilde_t`, +but output change is still measured by the true forward Jacobian `J_t`: + +```text +r_{t+1}^{FA} ≈ (I - η K_t^{FA}/N) r_t +K_t^{FA} = J_t J_tilde_t^T +``` + +`K_t^{FA}` is better called a tangent operator, not a PSD kernel. + +The local estimator freezes `K_0`. The finite-time estimator uses an early +operator velocity: + +```text +K_hat_t = K_0 + t (K_s - K_0) / s +``` + +Then: + +```text +L_hat_T = ||r_hat_T||^2 / (2N) +gap_hat_T = L_hat_T^FA - L_hat_T^BP +``` + +This contribution is conditional on early operator observations `(K_0, K_s)`. +It is not an architecture-only theorem. + +## 5. Distributional Empirical Validation + +Experiments should validate three levels: + +1. static alignment distributions: + ```text + Q_l ~ Beta(1/2, (D_l - 1)/2) + ``` +2. capacity and redundancy transition: + FA/BP gap opens near the hard FA capacity margin crossing; +3. tangent-operator trajectory distributions: + predicted gap distributions overlap empirical trajectory distributions. + +Current strong evidence: + +- 256-trajectory finite-time overlap plot for `d=2,w=64,T=50,s=20`; +- stress grid over depth, width, and horizon; +- systematic low-bias analysis showing residual error is kernel-path curvature, + not normalization. + +## Important Separation + +Do not claim: + +```text +capacity alone predicts exact loss gap +``` + +Correct claim: + +```text +capacity bounds identify the structural regime; +tangent-operator dynamics quantify the finite-time gap within that regime. +``` + +This separation avoids overclaiming while still giving a coherent theory chain. + +## Next Priority + +The weakest current visual is the phase-transition contribution. + +The desired figure should show: + +```text +capacity margin decreases -> redundant directions exhausted -> FA/BP gap opens +and then grows +``` + +The plot should use hard FA capacity margin on the x-axis, place the zero +margin as a vertical reference, and show BP/FA train loss or FA-BP train gap. +The most direct target is a width sweep on a random-label task, because random +labels make the task dimension controllable and force memorization capacity. |
