From 66fe7f5af856e4d3d2eff2a9d929bcd7ffdb94e8 Mon Sep 17 00:00:00 2001 From: YurenHao0426 Date: Fri, 5 Jun 2026 11:01:35 -0500 Subject: Add contribution roadmap and phase transition plot --- notes/18_contribution_roadmap.md | 182 +++++++++++++++++++++++++++++ notes/19_phase_transition_visualization.md | 162 +++++++++++++++++++++++++ 2 files changed, 344 insertions(+) create mode 100644 notes/18_contribution_roadmap.md create mode 100644 notes/19_phase_transition_visualization.md (limited to 'notes') diff --git a/notes/18_contribution_roadmap.md b/notes/18_contribution_roadmap.md new file mode 100644 index 0000000..179bd5d --- /dev/null +++ b/notes/18_contribution_roadmap.md @@ -0,0 +1,182 @@ +# Contribution Roadmap + +Working title: + +`Distributional Capacity Bounds for Feedback Alignment in Multilayer Perceptrons` + +The current paper should use five contributions. The fifth contribution is not +extra decoration: it separates structural capacity theory from finite-time +loss-gap prediction. + +## 1. Capacity Formalization + +We define the capacity burden induced by feedback alignment. + +Layerwise matrix alignment: + +```text +Q_l = cos^2(W_{l+1}^T, B_l) +``` + +For isotropic random feedback in dimension `D_l`: + +```text +Q_l ~ Beta(1/2, (D_l - 1)/2) +``` + +Alignment threshold `q_l` gives a log-volume capacity cost: + +```text +C_l(q_l) = -log P(Q_l >= q_l) +``` + +This formalizes what it means for feedback alignment to consume parameter +direction volume. + +## 2. Scaling Law and Redundancy Exhaustion + +Across independent layers, raw feasible volume multiplies while log-capacity +cost adds: + +```text +p_all = product_l p_l +C_all = sum_l C_l +``` + +For an equal-width MLP with width `n`, depth `L`, and fixed alignment threshold: + +```text +C_all = Θ(L n^2) +p_all = exp[-Θ(L n^2)] +``` + +Functional loss does not need to appear immediately. If total parameter +dimension is `P`, task dimension is `d`, and alignment imposes `k` generic +constraints, then hard functional rank loss is: + +```text +Δd_hard = max(0, k - (P - d)) +``` + +So the FA/BP gap should open when redundant directions are exhausted. + +This is the phase-transition contribution. + +## 3. Prior-Free Minimax Initialization Bound + +Let `b` be the normalized feedback direction in `D` dimensions and `a` be the +unknown normalized target backward direction. + +For any feedback initialization distribution `μ`, define: + +```text +M_μ = E[b b^T] +trace(M_μ) = 1 +``` + +Then: + +```text +inf_a E_μ[(a^T b)^2] = λ_min(M_μ) <= 1/D +``` + +Therefore: + +```text +sup_μ inf_a E_μ[(a^T b)^2] = 1/D +``` + +Isotropic random feedback reaches the bound. Without a prior over `W` or task +directions, no initialization can beat isotropic random feedback in worst-case +expected squared alignment. + +## 4. Tangent-Operator Loss/Gap Estimator + +This is the bridge from capacity regime to finite-time loss gap. + +For BP squared loss, standard tangent-kernel residual dynamics gives: + +```text +r_{t+1}^{BP} ≈ (I - η K_t^{BP}/N) r_t +K_t^{BP} = J_t J_t^T +``` + +For FA, the parameter update uses a surrogate backward Jacobian `J_tilde_t`, +but output change is still measured by the true forward Jacobian `J_t`: + +```text +r_{t+1}^{FA} ≈ (I - η K_t^{FA}/N) r_t +K_t^{FA} = J_t J_tilde_t^T +``` + +`K_t^{FA}` is better called a tangent operator, not a PSD kernel. + +The local estimator freezes `K_0`. The finite-time estimator uses an early +operator velocity: + +```text +K_hat_t = K_0 + t (K_s - K_0) / s +``` + +Then: + +```text +L_hat_T = ||r_hat_T||^2 / (2N) +gap_hat_T = L_hat_T^FA - L_hat_T^BP +``` + +This contribution is conditional on early operator observations `(K_0, K_s)`. +It is not an architecture-only theorem. + +## 5. Distributional Empirical Validation + +Experiments should validate three levels: + +1. static alignment distributions: + ```text + Q_l ~ Beta(1/2, (D_l - 1)/2) + ``` +2. capacity and redundancy transition: + FA/BP gap opens near the hard FA capacity margin crossing; +3. tangent-operator trajectory distributions: + predicted gap distributions overlap empirical trajectory distributions. + +Current strong evidence: + +- 256-trajectory finite-time overlap plot for `d=2,w=64,T=50,s=20`; +- stress grid over depth, width, and horizon; +- systematic low-bias analysis showing residual error is kernel-path curvature, + not normalization. + +## Important Separation + +Do not claim: + +```text +capacity alone predicts exact loss gap +``` + +Correct claim: + +```text +capacity bounds identify the structural regime; +tangent-operator dynamics quantify the finite-time gap within that regime. +``` + +This separation avoids overclaiming while still giving a coherent theory chain. + +## Next Priority + +The weakest current visual is the phase-transition contribution. + +The desired figure should show: + +```text +capacity margin decreases -> redundant directions exhausted -> FA/BP gap opens +and then grows +``` + +The plot should use hard FA capacity margin on the x-axis, place the zero +margin as a vertical reference, and show BP/FA train loss or FA-BP train gap. +The most direct target is a width sweep on a random-label task, because random +labels make the task dimension controllable and force memorization capacity. diff --git a/notes/19_phase_transition_visualization.md b/notes/19_phase_transition_visualization.md new file mode 100644 index 0000000..a2672c2 --- /dev/null +++ b/notes/19_phase_transition_visualization.md @@ -0,0 +1,162 @@ +# Phase Transition Visualization + +This note records the first clean phase-transition visualization for the +capacity-exhaustion contribution. + +## Goal + +We want a figure that shows: + +```text +available FA capacity decreases +-> redundant directions are exhausted +-> FA/BP train gap opens +-> gap grows while BP remains capable +``` + +The key x-axis is the hard FA capacity margin: + +```text +M_FA = P - K_FA - N*out +``` + +where: + +- `P` is the parameter count; +- `K_FA` is the hard feedback-alignment constraint count; +- `N*out` is the random-label task dimension. + +The predicted transition is at: + +```text +M_FA = 0 +``` + +## Experiment + +This run uses SGD, not Adam. + +Configuration: + +```text +task: random-label regression +architecture: 16 -> width -> width -> 4 +widths: 8, 12, 16, 24, 32, 48, 64, 96 +train samples: 128 +optimizer: full-batch SGD +learning rate: 0.01 +steps: 3000 +init seeds: 4 +feedback seeds per init: 8 +FA trajectories: 256 +``` + +Output directory: + +```text +outputs/phase_transition_sgd_256_lr001_T3000 +``` + +Main figures: + +```text +outputs/phase_transition_sgd_256_lr001_T3000/phase_transition_capacity_exhaustion.png +outputs/phase_transition_sgd_256_lr001_T3000/phase_transition_gap_only.png +``` + +## Result + +Summary: + +| width | FA margin | BP train | FA train | FA-BP train gap | +|---:|---:|---:|---:|---:| +| 96 | 1026 | 0.000249 | 0.009436 | 0.009187 | +| 64 | 514 | 0.003354 | 0.053192 | 0.049838 | +| 48 | 258 | 0.016732 | 0.146680 | 0.129947 | +| 32 | 2 | 0.089511 | 0.386777 | 0.297266 | +| 24 | -126 | 0.201897 | 0.600796 | 0.398899 | +| 16 | -254 | 0.547408 | 0.993219 | 0.445812 | +| 12 | -318 | 0.872346 | 1.204686 | 0.332340 | +| 8 | -382 | 1.260234 | 1.443789 | 0.183555 | + +The gap is small in the redundant region: + +```text +M_FA > 0 +``` + +It rises sharply around the hard FA capacity boundary: + +```text +M_FA ≈ 0 +``` + +and becomes largest in the FA-deficient but BP-capable region: + +```text +M_FA < 0 +and +M_BP = P - N*out > 0 +``` + +At the smallest widths, BP itself also becomes under-capacity: + +```text +M_BP < 0 +``` + +In that both-deficient regime, the FA-BP gap decreases because both methods +fail to fit the random labels. + +## Interpretation + +The visual story is not simply "smaller network means larger gap." The clean +phase-transition story is: + +1. high redundancy: + ```text + M_FA >> 0 + ``` + BP and FA both have enough effective room, so the train gap is small. + +2. FA redundancy exhausted: + ```text + M_FA crosses 0 + ``` + FA starts paying the feedback-alignment burden in task-relevant directions, + so the gap opens. + +3. FA-deficient but BP-capable: + ```text + M_FA < 0, M_BP > 0 + ``` + the gap grows and peaks. + +4. both-deficient: + ```text + M_BP < 0 + ``` + BP also cannot memorize, so the difference between FA and BP no longer + grows monotonically. + +This is the right phase-transition framing for the paper. + +## Caveat + +The hard margin is intentionally simple. It uses: + +```text +K_FA = (width*width - 1) + (width*out - 1) +``` + +for the two-hidden-layer MLP. It predicts the transition location well enough +to provide a clean regime variable, but the exact peak is shifted because real +FA alignment is soft, not a hard rank constraint. + +Do not claim that hard capacity margin alone predicts the exact final train +gap. The correct claim is: + +```text +hard FA margin predicts the regime where the gap opens; +tangent/operator dynamics or empirical trajectories determine the gap size. +``` -- cgit v1.2.3