summaryrefslogtreecommitdiff
path: root/notes/18_contribution_roadmap.md
diff options
context:
space:
mode:
Diffstat (limited to 'notes/18_contribution_roadmap.md')
-rw-r--r--notes/18_contribution_roadmap.md182
1 files changed, 182 insertions, 0 deletions
diff --git a/notes/18_contribution_roadmap.md b/notes/18_contribution_roadmap.md
new file mode 100644
index 0000000..179bd5d
--- /dev/null
+++ b/notes/18_contribution_roadmap.md
@@ -0,0 +1,182 @@
+# Contribution Roadmap
+
+Working title:
+
+`Distributional Capacity Bounds for Feedback Alignment in Multilayer Perceptrons`
+
+The current paper should use five contributions. The fifth contribution is not
+extra decoration: it separates structural capacity theory from finite-time
+loss-gap prediction.
+
+## 1. Capacity Formalization
+
+We define the capacity burden induced by feedback alignment.
+
+Layerwise matrix alignment:
+
+```text
+Q_l = cos^2(W_{l+1}^T, B_l)
+```
+
+For isotropic random feedback in dimension `D_l`:
+
+```text
+Q_l ~ Beta(1/2, (D_l - 1)/2)
+```
+
+Alignment threshold `q_l` gives a log-volume capacity cost:
+
+```text
+C_l(q_l) = -log P(Q_l >= q_l)
+```
+
+This formalizes what it means for feedback alignment to consume parameter
+direction volume.
+
+## 2. Scaling Law and Redundancy Exhaustion
+
+Across independent layers, raw feasible volume multiplies while log-capacity
+cost adds:
+
+```text
+p_all = product_l p_l
+C_all = sum_l C_l
+```
+
+For an equal-width MLP with width `n`, depth `L`, and fixed alignment threshold:
+
+```text
+C_all = Θ(L n^2)
+p_all = exp[-Θ(L n^2)]
+```
+
+Functional loss does not need to appear immediately. If total parameter
+dimension is `P`, task dimension is `d`, and alignment imposes `k` generic
+constraints, then hard functional rank loss is:
+
+```text
+Δd_hard = max(0, k - (P - d))
+```
+
+So the FA/BP gap should open when redundant directions are exhausted.
+
+This is the phase-transition contribution.
+
+## 3. Prior-Free Minimax Initialization Bound
+
+Let `b` be the normalized feedback direction in `D` dimensions and `a` be the
+unknown normalized target backward direction.
+
+For any feedback initialization distribution `μ`, define:
+
+```text
+M_μ = E[b b^T]
+trace(M_μ) = 1
+```
+
+Then:
+
+```text
+inf_a E_μ[(a^T b)^2] = λ_min(M_μ) <= 1/D
+```
+
+Therefore:
+
+```text
+sup_μ inf_a E_μ[(a^T b)^2] = 1/D
+```
+
+Isotropic random feedback reaches the bound. Without a prior over `W` or task
+directions, no initialization can beat isotropic random feedback in worst-case
+expected squared alignment.
+
+## 4. Tangent-Operator Loss/Gap Estimator
+
+This is the bridge from capacity regime to finite-time loss gap.
+
+For BP squared loss, standard tangent-kernel residual dynamics gives:
+
+```text
+r_{t+1}^{BP} ≈ (I - η K_t^{BP}/N) r_t
+K_t^{BP} = J_t J_t^T
+```
+
+For FA, the parameter update uses a surrogate backward Jacobian `J_tilde_t`,
+but output change is still measured by the true forward Jacobian `J_t`:
+
+```text
+r_{t+1}^{FA} ≈ (I - η K_t^{FA}/N) r_t
+K_t^{FA} = J_t J_tilde_t^T
+```
+
+`K_t^{FA}` is better called a tangent operator, not a PSD kernel.
+
+The local estimator freezes `K_0`. The finite-time estimator uses an early
+operator velocity:
+
+```text
+K_hat_t = K_0 + t (K_s - K_0) / s
+```
+
+Then:
+
+```text
+L_hat_T = ||r_hat_T||^2 / (2N)
+gap_hat_T = L_hat_T^FA - L_hat_T^BP
+```
+
+This contribution is conditional on early operator observations `(K_0, K_s)`.
+It is not an architecture-only theorem.
+
+## 5. Distributional Empirical Validation
+
+Experiments should validate three levels:
+
+1. static alignment distributions:
+ ```text
+ Q_l ~ Beta(1/2, (D_l - 1)/2)
+ ```
+2. capacity and redundancy transition:
+ FA/BP gap opens near the hard FA capacity margin crossing;
+3. tangent-operator trajectory distributions:
+ predicted gap distributions overlap empirical trajectory distributions.
+
+Current strong evidence:
+
+- 256-trajectory finite-time overlap plot for `d=2,w=64,T=50,s=20`;
+- stress grid over depth, width, and horizon;
+- systematic low-bias analysis showing residual error is kernel-path curvature,
+ not normalization.
+
+## Important Separation
+
+Do not claim:
+
+```text
+capacity alone predicts exact loss gap
+```
+
+Correct claim:
+
+```text
+capacity bounds identify the structural regime;
+tangent-operator dynamics quantify the finite-time gap within that regime.
+```
+
+This separation avoids overclaiming while still giving a coherent theory chain.
+
+## Next Priority
+
+The weakest current visual is the phase-transition contribution.
+
+The desired figure should show:
+
+```text
+capacity margin decreases -> redundant directions exhausted -> FA/BP gap opens
+and then grows
+```
+
+The plot should use hard FA capacity margin on the x-axis, place the zero
+margin as a vertical reference, and show BP/FA train loss or FA-BP train gap.
+The most direct target is a width sweep on a random-label task, because random
+labels make the task dimension controllable and force memorization capacity.