# Experiment Notes ## Experiment Layer 1: Static Distribution Validation Purpose: verify the beta law before training dynamics enter. Procedure: 1. Choose matrix shape \(n_l\times n_{l+1}\), so \(D_l=n_l n_{l+1}\). 2. Sample \(A_l\) and \(B_l\) independently from isotropic distributions. 3. Compute: \[ Q_l= \frac{ \langle A_l,B_l\rangle_F^2 }{ \|A_l\|_F^2\|B_l\|_F^2 }. \] 4. Compare empirical distribution to: \[ \mathrm{Beta}\left(\frac12,\frac{D_l-1}{2}\right). \] Metrics: - Histogram overlay. - QQ plot. - KS statistic. - Tail calibration: \[ \Pr(Q_l\ge q). \] Initialization variants: - Gaussian. - Rademacher. - Uniform sphere. - Orthogonal or semi-orthogonal. - Sparse. - Low-rank. - Block-diagonal. Expected: - Dense isotropic variants match beta law. - Structured variants deviate in predictable ways. ## Experiment Layer 2: Scaling Validation Purpose: verify capacity scaling with depth and width. Compute: \[ C_l(q)= -\log \left[ 1-I_q\left(\frac12,\frac{D_l-1}{2}\right) \right]. \] Total: \[ C_{\mathrm{all}}=\sum_l C_l(q_l). \] Sweeps: - Width \(n\). - Depth \(L\). - Threshold \(q\). - Threshold regime \(q=c/D_l\). - Feedback rank. - Feedback sparsity. Predictions: Fixed \(q\): \[ C_{\mathrm{all}}=\Theta(Ln^2) \] for equal-width MLPs. Chance-level \(q=c/D_l\): \[ C_{\mathrm{all}}=\Theta(L). \] ## Experiment Layer 3: Local Gradient Alignment Purpose: connect static matrix alignment to update direction mismatch. During training, record: \[ \Gamma_t= \frac{ \langle g_t^{\mathrm{BP}},g_t^{\mathrm{FA}}\rangle }{ \|g_t^{\mathrm{BP}}\|\|g_t^{\mathrm{FA}}\| }. \] Layerwise: \[ \Gamma_{l,t}= \frac{ \langle g_{l,t}^{\mathrm{BP}},g_{l,t}^{\mathrm{FA}}\rangle }{ \|g_{l,t}^{\mathrm{BP}}\|\|g_{l,t}^{\mathrm{FA}}\| }. \] Also record: \[ Q_l(t)=\cos^2(W_{l+1}(t)^\top,B_l). \] Questions: - Does \(Q_l(0)\) match the beta baseline? - Does \(Q_l(t)\) shift right during the alignment phase? - Does gradient alignment improve before memorization or loss reduction? ## Experiment Layer 4: Trajectory Ensemble Purpose: validate whether capacity proxies explain FA/BP training gaps. For each architecture and dataset: 1. Fix data seed and model architecture. 2. Train BP baseline. 3. Train many FA runs over feedback seeds \(B\). 4. Record: \[ \Delta L_T(B)=L_T^{\mathrm{FA}}(B)-L_T^{\mathrm{BP}}, \] \[ \Delta A_T(B)=A_T^{\mathrm{BP}}-A_T^{\mathrm{FA}}, \] \[ C_{\mathrm{all}}(B,t), \quad \Gamma_t(B), \quad Q_l(t). \] Datasets: - Synthetic Gaussian regression. - MNIST MLP. - Fashion-MNIST MLP. - CIFAR-10 flattened MLP, optional later. Architectures: - Equal-width MLPs. - Width sweep. - Depth sweep. - Narrow bottleneck sweep. Expected: - Overparameterized regimes: large parameter-volume cost can coexist with small functional gap. - Near redundancy exhaustion: FA/BP gap should increase sharply. - Poor feedback conditioning can worsen trajectory gap even when angular minimax bound is unchanged. ## Plots Static: - Histogram and beta density. - QQ plot. - Tail probability calibration. Scaling: - \(C_{\mathrm{all}}\) vs \(Ln^2\). - \(-\log p_{\mathrm{all}}\) vs depth. - Scaling collapse for \(D_lQ_l\Rightarrow \chi_1^2\). Trajectory: - \(Q_l(t)\) over training. - \(\Gamma_t\) over training. - \(\Delta L_T\) vs capacity proxy. - \(\Delta L_T\) vs conditioning proxy. - Phase transition plot against \(k-(P-d)\). ## Implementation Notes Start with NumPy or PyTorch scripts that do not require full training. First script target: - Sample \(A,B\). - Compute \(Q\). - Save empirical moments and KS statistic. - Produce beta overlay plots. Only after this is clean, add FA/BP training loops. ## Baseline Run Log Script: ```bash python scripts/static_alignment_beta.py --rows 16 --cols 16 --samples 20000 --seed 7 --plot ``` Result: - \(D=256\) - empirical mean: `0.0039265884` - theoretical mean: `0.00390625` - empirical variance: `3.0311252e-05` - theoretical variance: `3.0162723e-05` - KS statistic: `0.00526931` - KS p-value: `0.633285` This is a clean first-pass validation for the isotropic Gaussian case. ## Scaling Run Log Script: ```bash python scripts/capacity_scaling.py --plot ``` Default sweep: - widths: `16, 32, 64, 128` - feedback-aligned layer counts: `1, 2, 4, 8, 16` - fixed threshold: \(q=0.01\) - chance-level threshold: \(q=1/D\) - log unit: nats Result: - rows written: `40` - fixed-threshold max total cost: `1361.74` nats at width `128`, layers `16` - chance-level max total cost: `18.3652` nats at width `128`, layers `16` This cleanly separates the fixed-threshold regime, where total cost scales like \(Ln^2\), from the chance-level regime, where per-layer cost is nearly width-independent.