summaryrefslogtreecommitdiff
path: root/notes
diff options
context:
space:
mode:
authorYurenHao0426 <Blackhao0426@gmail.com>2026-06-05 11:01:35 -0500
committerYurenHao0426 <Blackhao0426@gmail.com>2026-06-05 11:01:35 -0500
commit66fe7f5af856e4d3d2eff2a9d929bcd7ffdb94e8 (patch)
tree68a8aa9bd344db767933bd905bdcdf627e126c08 /notes
parent9ce9158fb1ee767b31454b9c707bc86edd0a9252 (diff)
Add contribution roadmap and phase transition plot
Diffstat (limited to 'notes')
-rw-r--r--notes/18_contribution_roadmap.md182
-rw-r--r--notes/19_phase_transition_visualization.md162
2 files changed, 344 insertions, 0 deletions
diff --git a/notes/18_contribution_roadmap.md b/notes/18_contribution_roadmap.md
new file mode 100644
index 0000000..179bd5d
--- /dev/null
+++ b/notes/18_contribution_roadmap.md
@@ -0,0 +1,182 @@
+# Contribution Roadmap
+
+Working title:
+
+`Distributional Capacity Bounds for Feedback Alignment in Multilayer Perceptrons`
+
+The current paper should use five contributions. The fifth contribution is not
+extra decoration: it separates structural capacity theory from finite-time
+loss-gap prediction.
+
+## 1. Capacity Formalization
+
+We define the capacity burden induced by feedback alignment.
+
+Layerwise matrix alignment:
+
+```text
+Q_l = cos^2(W_{l+1}^T, B_l)
+```
+
+For isotropic random feedback in dimension `D_l`:
+
+```text
+Q_l ~ Beta(1/2, (D_l - 1)/2)
+```
+
+Alignment threshold `q_l` gives a log-volume capacity cost:
+
+```text
+C_l(q_l) = -log P(Q_l >= q_l)
+```
+
+This formalizes what it means for feedback alignment to consume parameter
+direction volume.
+
+## 2. Scaling Law and Redundancy Exhaustion
+
+Across independent layers, raw feasible volume multiplies while log-capacity
+cost adds:
+
+```text
+p_all = product_l p_l
+C_all = sum_l C_l
+```
+
+For an equal-width MLP with width `n`, depth `L`, and fixed alignment threshold:
+
+```text
+C_all = Θ(L n^2)
+p_all = exp[-Θ(L n^2)]
+```
+
+Functional loss does not need to appear immediately. If total parameter
+dimension is `P`, task dimension is `d`, and alignment imposes `k` generic
+constraints, then hard functional rank loss is:
+
+```text
+Δd_hard = max(0, k - (P - d))
+```
+
+So the FA/BP gap should open when redundant directions are exhausted.
+
+This is the phase-transition contribution.
+
+## 3. Prior-Free Minimax Initialization Bound
+
+Let `b` be the normalized feedback direction in `D` dimensions and `a` be the
+unknown normalized target backward direction.
+
+For any feedback initialization distribution `μ`, define:
+
+```text
+M_μ = E[b b^T]
+trace(M_μ) = 1
+```
+
+Then:
+
+```text
+inf_a E_μ[(a^T b)^2] = λ_min(M_μ) <= 1/D
+```
+
+Therefore:
+
+```text
+sup_μ inf_a E_μ[(a^T b)^2] = 1/D
+```
+
+Isotropic random feedback reaches the bound. Without a prior over `W` or task
+directions, no initialization can beat isotropic random feedback in worst-case
+expected squared alignment.
+
+## 4. Tangent-Operator Loss/Gap Estimator
+
+This is the bridge from capacity regime to finite-time loss gap.
+
+For BP squared loss, standard tangent-kernel residual dynamics gives:
+
+```text
+r_{t+1}^{BP} ≈ (I - η K_t^{BP}/N) r_t
+K_t^{BP} = J_t J_t^T
+```
+
+For FA, the parameter update uses a surrogate backward Jacobian `J_tilde_t`,
+but output change is still measured by the true forward Jacobian `J_t`:
+
+```text
+r_{t+1}^{FA} ≈ (I - η K_t^{FA}/N) r_t
+K_t^{FA} = J_t J_tilde_t^T
+```
+
+`K_t^{FA}` is better called a tangent operator, not a PSD kernel.
+
+The local estimator freezes `K_0`. The finite-time estimator uses an early
+operator velocity:
+
+```text
+K_hat_t = K_0 + t (K_s - K_0) / s
+```
+
+Then:
+
+```text
+L_hat_T = ||r_hat_T||^2 / (2N)
+gap_hat_T = L_hat_T^FA - L_hat_T^BP
+```
+
+This contribution is conditional on early operator observations `(K_0, K_s)`.
+It is not an architecture-only theorem.
+
+## 5. Distributional Empirical Validation
+
+Experiments should validate three levels:
+
+1. static alignment distributions:
+ ```text
+ Q_l ~ Beta(1/2, (D_l - 1)/2)
+ ```
+2. capacity and redundancy transition:
+ FA/BP gap opens near the hard FA capacity margin crossing;
+3. tangent-operator trajectory distributions:
+ predicted gap distributions overlap empirical trajectory distributions.
+
+Current strong evidence:
+
+- 256-trajectory finite-time overlap plot for `d=2,w=64,T=50,s=20`;
+- stress grid over depth, width, and horizon;
+- systematic low-bias analysis showing residual error is kernel-path curvature,
+ not normalization.
+
+## Important Separation
+
+Do not claim:
+
+```text
+capacity alone predicts exact loss gap
+```
+
+Correct claim:
+
+```text
+capacity bounds identify the structural regime;
+tangent-operator dynamics quantify the finite-time gap within that regime.
+```
+
+This separation avoids overclaiming while still giving a coherent theory chain.
+
+## Next Priority
+
+The weakest current visual is the phase-transition contribution.
+
+The desired figure should show:
+
+```text
+capacity margin decreases -> redundant directions exhausted -> FA/BP gap opens
+and then grows
+```
+
+The plot should use hard FA capacity margin on the x-axis, place the zero
+margin as a vertical reference, and show BP/FA train loss or FA-BP train gap.
+The most direct target is a width sweep on a random-label task, because random
+labels make the task dimension controllable and force memorization capacity.
diff --git a/notes/19_phase_transition_visualization.md b/notes/19_phase_transition_visualization.md
new file mode 100644
index 0000000..a2672c2
--- /dev/null
+++ b/notes/19_phase_transition_visualization.md
@@ -0,0 +1,162 @@
+# Phase Transition Visualization
+
+This note records the first clean phase-transition visualization for the
+capacity-exhaustion contribution.
+
+## Goal
+
+We want a figure that shows:
+
+```text
+available FA capacity decreases
+-> redundant directions are exhausted
+-> FA/BP train gap opens
+-> gap grows while BP remains capable
+```
+
+The key x-axis is the hard FA capacity margin:
+
+```text
+M_FA = P - K_FA - N*out
+```
+
+where:
+
+- `P` is the parameter count;
+- `K_FA` is the hard feedback-alignment constraint count;
+- `N*out` is the random-label task dimension.
+
+The predicted transition is at:
+
+```text
+M_FA = 0
+```
+
+## Experiment
+
+This run uses SGD, not Adam.
+
+Configuration:
+
+```text
+task: random-label regression
+architecture: 16 -> width -> width -> 4
+widths: 8, 12, 16, 24, 32, 48, 64, 96
+train samples: 128
+optimizer: full-batch SGD
+learning rate: 0.01
+steps: 3000
+init seeds: 4
+feedback seeds per init: 8
+FA trajectories: 256
+```
+
+Output directory:
+
+```text
+outputs/phase_transition_sgd_256_lr001_T3000
+```
+
+Main figures:
+
+```text
+outputs/phase_transition_sgd_256_lr001_T3000/phase_transition_capacity_exhaustion.png
+outputs/phase_transition_sgd_256_lr001_T3000/phase_transition_gap_only.png
+```
+
+## Result
+
+Summary:
+
+| width | FA margin | BP train | FA train | FA-BP train gap |
+|---:|---:|---:|---:|---:|
+| 96 | 1026 | 0.000249 | 0.009436 | 0.009187 |
+| 64 | 514 | 0.003354 | 0.053192 | 0.049838 |
+| 48 | 258 | 0.016732 | 0.146680 | 0.129947 |
+| 32 | 2 | 0.089511 | 0.386777 | 0.297266 |
+| 24 | -126 | 0.201897 | 0.600796 | 0.398899 |
+| 16 | -254 | 0.547408 | 0.993219 | 0.445812 |
+| 12 | -318 | 0.872346 | 1.204686 | 0.332340 |
+| 8 | -382 | 1.260234 | 1.443789 | 0.183555 |
+
+The gap is small in the redundant region:
+
+```text
+M_FA > 0
+```
+
+It rises sharply around the hard FA capacity boundary:
+
+```text
+M_FA ≈ 0
+```
+
+and becomes largest in the FA-deficient but BP-capable region:
+
+```text
+M_FA < 0
+and
+M_BP = P - N*out > 0
+```
+
+At the smallest widths, BP itself also becomes under-capacity:
+
+```text
+M_BP < 0
+```
+
+In that both-deficient regime, the FA-BP gap decreases because both methods
+fail to fit the random labels.
+
+## Interpretation
+
+The visual story is not simply "smaller network means larger gap." The clean
+phase-transition story is:
+
+1. high redundancy:
+ ```text
+ M_FA >> 0
+ ```
+ BP and FA both have enough effective room, so the train gap is small.
+
+2. FA redundancy exhausted:
+ ```text
+ M_FA crosses 0
+ ```
+ FA starts paying the feedback-alignment burden in task-relevant directions,
+ so the gap opens.
+
+3. FA-deficient but BP-capable:
+ ```text
+ M_FA < 0, M_BP > 0
+ ```
+ the gap grows and peaks.
+
+4. both-deficient:
+ ```text
+ M_BP < 0
+ ```
+ BP also cannot memorize, so the difference between FA and BP no longer
+ grows monotonically.
+
+This is the right phase-transition framing for the paper.
+
+## Caveat
+
+The hard margin is intentionally simple. It uses:
+
+```text
+K_FA = (width*width - 1) + (width*out - 1)
+```
+
+for the two-hidden-layer MLP. It predicts the transition location well enough
+to provide a clean regime variable, but the exact peak is shifted because real
+FA alignment is soft, not a hard rank constraint.
+
+Do not claim that hard capacity margin alone predicts the exact final train
+gap. The correct claim is:
+
+```text
+hard FA margin predicts the regime where the gap opens;
+tangent/operator dynamics or empirical trajectories determine the gap size.
+```