diff options
Diffstat (limited to 'notes')
| -rw-r--r-- | notes/22_dense_phase_transition_soft_ramp.md | 93 |
1 files changed, 93 insertions, 0 deletions
diff --git a/notes/22_dense_phase_transition_soft_ramp.md b/notes/22_dense_phase_transition_soft_ramp.md new file mode 100644 index 0000000..88f6992 --- /dev/null +++ b/notes/22_dense_phase_transition_soft_ramp.md @@ -0,0 +1,93 @@ +# Dense Long-Training Phase Transition + +The previous `T=30000` transition plot only had four width points. That was not +enough to decide whether the long-time curve has a kink or a soft ramp. + +We ran a denser long-training sweep: + +```text +outputs/phase_transition_dense_T30000_352traj +``` + +Setup: + +```text +task: random-label regression +architecture: 16 -> width -> width -> 4 +optimizer: full-batch SGD +learning rate: 0.01 +train samples: 128 +steps: 30000 +widths: 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40 +init seeds: 1 +feedback seeds: 32 +FA trajectories: 352 total +``` + +Main plot: + +```text +outputs/phase_transition_dense_T30000_352traj/dense_T30000_gap_logscale.png +``` + +## Result + +| width | FA margin | FA trajectories | train gap mean | train gap std | +|---:|---:|---:|---:|---:| +| 20 | -190 | 32 | 0.146803 | 0.042009 | +| 22 | -158 | 32 | 0.100526 | 0.034803 | +| 24 | -126 | 32 | 0.046464 | 0.024474 | +| 26 | -94 | 32 | 0.029671 | 0.023591 | +| 28 | -62 | 32 | 0.014195 | 0.009368 | +| 30 | -30 | 32 | 0.007466 | 0.004624 | +| 32 | 2 | 32 | 0.003729 | 0.003787 | +| 34 | 34 | 32 | 0.002079 | 0.002232 | +| 36 | 66 | 32 | 0.000791 | 0.000605 | +| 38 | 98 | 32 | 0.000553 | 0.000545 | +| 40 | 130 | 32 | 0.000162 | 0.000134 | + +The dense sweep does not show a sharp kink. It shows a smooth, approximately +log-linear soft ramp in the long-training FA/BP train gap. + +## Interpretation + +The earlier `T=10000` curve was partly finite-time undertraining, because +longer training collapses the positive-margin gap from roughly `0.05` to +roughly `0.004` near margin zero. + +However, after adding dense `T=30000` points, the long-training curve is still +not a hard step. The correct conclusion is: + +```text +hard FA margin predicts a capacity-controlled regime variable, +not a discontinuous empirical transition. +``` + +The empirical transition is soft: + +```text +more negative FA margin -> larger train gap; +positive FA margin -> small but nonzero finite-time/soft-alignment tail; +no sharp kink at margin zero. +``` + +## Consequence for the Paper + +Do not write: + +```text +FA and BP are identical until redundant parameters are exhausted, then a sharp +gap appears. +``` + +Use: + +```text +The hard margin gives a conservative redundancy-exhaustion boundary. Empirically, +the FA/BP train-gap distribution changes smoothly with the margin; longer +training removes most positive-margin gap, while negative margins retain a +capacity-controlled gap. +``` + +This is still a useful capacity result, but the contribution should be framed as +a scaling law plus soft capacity transition, not a hard phase transition. |
