summaryrefslogtreecommitdiff
path: root/notes/02_experiment_notes.md
diff options
context:
space:
mode:
Diffstat (limited to 'notes/02_experiment_notes.md')
-rw-r--r--notes/02_experiment_notes.md217
1 files changed, 217 insertions, 0 deletions
diff --git a/notes/02_experiment_notes.md b/notes/02_experiment_notes.md
new file mode 100644
index 0000000..c418fae
--- /dev/null
+++ b/notes/02_experiment_notes.md
@@ -0,0 +1,217 @@
+# Experiment Notes
+
+## Experiment Layer 1: Static Distribution Validation
+
+Purpose: verify the beta law before training dynamics enter.
+
+Procedure:
+
+1. Choose matrix shape \(n_l\times n_{l+1}\), so \(D_l=n_l n_{l+1}\).
+2. Sample \(A_l\) and \(B_l\) independently from isotropic distributions.
+3. Compute:
+
+\[
+Q_l=
+\frac{
+\langle A_l,B_l\rangle_F^2
+}{
+\|A_l\|_F^2\|B_l\|_F^2
+}.
+\]
+
+4. Compare empirical distribution to:
+
+\[
+\mathrm{Beta}\left(\frac12,\frac{D_l-1}{2}\right).
+\]
+
+Metrics:
+
+- Histogram overlay.
+- QQ plot.
+- KS statistic.
+- Tail calibration:
+
+\[
+\Pr(Q_l\ge q).
+\]
+
+Initialization variants:
+
+- Gaussian.
+- Rademacher.
+- Uniform sphere.
+- Orthogonal or semi-orthogonal.
+- Sparse.
+- Low-rank.
+- Block-diagonal.
+
+Expected:
+
+- Dense isotropic variants match beta law.
+- Structured variants deviate in predictable ways.
+
+## Experiment Layer 2: Scaling Validation
+
+Purpose: verify capacity scaling with depth and width.
+
+Compute:
+
+\[
+C_l(q)=
+-\log
+\left[
+1-I_q\left(\frac12,\frac{D_l-1}{2}\right)
+\right].
+\]
+
+Total:
+
+\[
+C_{\mathrm{all}}=\sum_l C_l(q_l).
+\]
+
+Sweeps:
+
+- Width \(n\).
+- Depth \(L\).
+- Threshold \(q\).
+- Threshold regime \(q=c/D_l\).
+- Feedback rank.
+- Feedback sparsity.
+
+Predictions:
+
+Fixed \(q\):
+
+\[
+C_{\mathrm{all}}=\Theta(Ln^2)
+\]
+
+for equal-width MLPs.
+
+Chance-level \(q=c/D_l\):
+
+\[
+C_{\mathrm{all}}=\Theta(L).
+\]
+
+## Experiment Layer 3: Local Gradient Alignment
+
+Purpose: connect static matrix alignment to update direction mismatch.
+
+During training, record:
+
+\[
+\Gamma_t=
+\frac{
+\langle g_t^{\mathrm{BP}},g_t^{\mathrm{FA}}\rangle
+}{
+\|g_t^{\mathrm{BP}}\|\|g_t^{\mathrm{FA}}\|
+}.
+\]
+
+Layerwise:
+
+\[
+\Gamma_{l,t}=
+\frac{
+\langle g_{l,t}^{\mathrm{BP}},g_{l,t}^{\mathrm{FA}}\rangle
+}{
+\|g_{l,t}^{\mathrm{BP}}\|\|g_{l,t}^{\mathrm{FA}}\|
+}.
+\]
+
+Also record:
+
+\[
+Q_l(t)=\cos^2(W_{l+1}(t)^\top,B_l).
+\]
+
+Questions:
+
+- Does \(Q_l(0)\) match the beta baseline?
+- Does \(Q_l(t)\) shift right during the alignment phase?
+- Does gradient alignment improve before memorization or loss reduction?
+
+## Experiment Layer 4: Trajectory Ensemble
+
+Purpose: validate whether capacity proxies explain FA/BP training gaps.
+
+For each architecture and dataset:
+
+1. Fix data seed and model architecture.
+2. Train BP baseline.
+3. Train many FA runs over feedback seeds \(B\).
+4. Record:
+
+\[
+\Delta L_T(B)=L_T^{\mathrm{FA}}(B)-L_T^{\mathrm{BP}},
+\]
+
+\[
+\Delta A_T(B)=A_T^{\mathrm{BP}}-A_T^{\mathrm{FA}},
+\]
+
+\[
+C_{\mathrm{all}}(B,t),
+\quad
+\Gamma_t(B),
+\quad
+Q_l(t).
+\]
+
+Datasets:
+
+- Synthetic Gaussian regression.
+- MNIST MLP.
+- Fashion-MNIST MLP.
+- CIFAR-10 flattened MLP, optional later.
+
+Architectures:
+
+- Equal-width MLPs.
+- Width sweep.
+- Depth sweep.
+- Narrow bottleneck sweep.
+
+Expected:
+
+- Overparameterized regimes: large parameter-volume cost can coexist with small functional gap.
+- Near redundancy exhaustion: FA/BP gap should increase sharply.
+- Poor feedback conditioning can worsen trajectory gap even when angular minimax bound is unchanged.
+
+## Plots
+
+Static:
+
+- Histogram and beta density.
+- QQ plot.
+- Tail probability calibration.
+
+Scaling:
+
+- \(C_{\mathrm{all}}\) vs \(Ln^2\).
+- \(-\log p_{\mathrm{all}}\) vs depth.
+- Scaling collapse for \(D_lQ_l\Rightarrow \chi_1^2\).
+
+Trajectory:
+
+- \(Q_l(t)\) over training.
+- \(\Gamma_t\) over training.
+- \(\Delta L_T\) vs capacity proxy.
+- \(\Delta L_T\) vs conditioning proxy.
+- Phase transition plot against \(k-(P-d)\).
+
+## Implementation Notes
+
+Start with NumPy or PyTorch scripts that do not require full training.
+
+First script target:
+
+- Sample \(A,B\).
+- Compute \(Q\).
+- Save empirical moments and KS statistic.
+- Produce beta overlay plots.
+
+Only after this is clean, add FA/BP training loops.