# Phase Transition Visualization This note records the first clean phase-transition visualization for the capacity-exhaustion contribution. ## Goal We want a figure that shows: ```text available FA capacity decreases -> redundant directions are exhausted -> FA/BP train gap opens -> gap grows while BP remains capable ``` The key x-axis is the hard FA capacity margin: ```text M_FA = P - K_FA - N*out ``` where: - `P` is the parameter count; - `K_FA` is the hard feedback-alignment constraint count; - `N*out` is the random-label task dimension. The predicted transition is at: ```text M_FA = 0 ``` ## Experiment This run uses SGD, not Adam. Configuration: ```text task: random-label regression architecture: 16 -> width -> width -> 4 widths: 8, 12, 16, 24, 32, 48, 64, 96 train samples: 128 optimizer: full-batch SGD learning rate: 0.01 steps: 3000 init seeds: 4 feedback seeds per init: 8 FA trajectories: 256 ``` Output directory: ```text outputs/phase_transition_sgd_256_lr001_T3000 ``` Main figures: ```text outputs/phase_transition_sgd_256_lr001_T3000/phase_transition_capacity_exhaustion.png outputs/phase_transition_sgd_256_lr001_T3000/phase_transition_gap_only.png ``` ## Result Summary: | width | FA margin | BP train | FA train | FA-BP train gap | |---:|---:|---:|---:|---:| | 96 | 1026 | 0.000249 | 0.009436 | 0.009187 | | 64 | 514 | 0.003354 | 0.053192 | 0.049838 | | 48 | 258 | 0.016732 | 0.146680 | 0.129947 | | 32 | 2 | 0.089511 | 0.386777 | 0.297266 | | 24 | -126 | 0.201897 | 0.600796 | 0.398899 | | 16 | -254 | 0.547408 | 0.993219 | 0.445812 | | 12 | -318 | 0.872346 | 1.204686 | 0.332340 | | 8 | -382 | 1.260234 | 1.443789 | 0.183555 | The gap is small in the redundant region: ```text M_FA > 0 ``` It rises sharply around the hard FA capacity boundary: ```text M_FA ≈ 0 ``` and becomes largest in the FA-deficient but BP-capable region: ```text M_FA < 0 and M_BP = P - N*out > 0 ``` At the smallest widths, BP itself also becomes under-capacity: ```text M_BP < 0 ``` In that both-deficient regime, the FA-BP gap decreases because both methods fail to fit the random labels. ## Interpretation The visual story is not simply "smaller network means larger gap." The clean phase-transition story is: 1. high redundancy: ```text M_FA >> 0 ``` BP and FA both have enough effective room, so the train gap is small. 2. FA redundancy exhausted: ```text M_FA crosses 0 ``` FA starts paying the feedback-alignment burden in task-relevant directions, so the gap opens. 3. FA-deficient but BP-capable: ```text M_FA < 0, M_BP > 0 ``` the gap grows and peaks. 4. both-deficient: ```text M_BP < 0 ``` BP also cannot memorize, so the difference between FA and BP no longer grows monotonically. This is the right phase-transition framing for the paper. ## Caveat The hard margin is intentionally simple. It uses: ```text K_FA = (width*width - 1) + (width*out - 1) ``` for the two-hidden-layer MLP. It predicts the transition location well enough to provide a clean regime variable, but the exact peak is shifted because real FA alignment is soft, not a hard rank constraint. Do not claim that hard capacity margin alone predicts the exact final train gap. The correct claim is: ```text hard FA margin predicts the regime where the gap opens; tangent/operator dynamics or empirical trajectories determine the gap size. ```