summaryrefslogtreecommitdiff
path: root/scripts/README.md
blob: efd420864c497078ae7dc9652a35dcc4a7745b8f (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
# Scripts

## Static Alignment Beta Law

Run:

```bash
python scripts/static_alignment_beta.py --rows 16 --cols 16 --samples 20000 --seed 7 --plot
```

This samples independent matrix pairs \(A,B\), computes:

\[
Q =
\frac{\langle A,B\rangle_F^2}{\|A\|_F^2\|B\|_F^2},
\]

and compares the empirical distribution with:

\[
\mathrm{Beta}\left(\frac12,\frac{D-1}{2}\right),
\qquad
D=\texttt{rows}\times\texttt{cols}.
\]

Outputs are written under `outputs/static_alignment_beta/`, which is ignored by Git.

## Capacity Scaling

Run:

```bash
python scripts/capacity_scaling.py --plot
```

This computes:

\[
C_l(q)
=
-\log\Pr(Q_l\ge q)
\]

for equal-width feedback blocks with \(D=n^2\), then sums over the number of feedback-aligned layers:

\[
C_{\mathrm{all}}=\sum_l C_l(q_l).
\]

The default run compares two regimes:

- `fixed`: \(q=0.01\), where \(C_{\mathrm{all}}\) grows like \(\Theta(Ln^2)\).
- `chance`: \(q=1/D\), where \(C_{\mathrm{all}}\) grows mostly with \(L\).

Outputs are written under `outputs/capacity_scaling/`.

## Large Empirical Capacity Validation

Run:

```bash
python scripts/capacity_empirical_validation.py --dimensions 64 128 256 512 1024 2048 4096 --samples 100000 --batch-size 2048 --seed 123 --plot
```

This uses rotational invariance to fix the target direction and sample random
feedback directions:

\[
Q=\frac{z_1^2}{\|z\|^2},
\qquad
z\sim \mathcal N(0,I_D).
\]

It validates:

- beta-law distribution calibration across dimensions;
- empirical tail probabilities for \(q=c/D\) and fixed \(q\);
- empirical capacity cost \(C(q)=-\log P(Q\ge q)\);
- multilayer product scaling for all-alignment events.

Outputs are written under `outputs/capacity_empirical_validation/`:

- `distribution_summary.csv`
- `capacity_tails.csv`
- `multilayer_capacity.csv`
- diagnostic plots when `--plot` is set.

## Multilayer Observed Capacity Distribution

Run:

```bash
python scripts/multilayer_capacity_distribution.py --dimensions 64 256 1024 4096 --layers 1 2 4 8 16 --samples 100000 --batch-size 8192 --seed 456 --plot
```

This validates the full null distribution of observed capacity surprisal:

\[
S_l=-\log P(Q_l'\ge Q_l).
\]

Theory predicts:

\[
S_l\sim \mathrm{Exp}(1),
\qquad
\sum_{l=1}^L S_l\sim \mathrm{Gamma}(L,1).
\]

The script samples random direction cosines through the exact chi-square
representation \(Q=X/(X+Y)\), with \(X\sim\chi^2_1\) and
\(Y\sim\chi^2_{D-1}\). Outputs are written under
`outputs/multilayer_capacity_distribution/`.

## Minimax Initialization Bound

Run:

```bash
python scripts/minimax_initialization.py --dimension 32 --feedback-samples 20000 --target-samples 10000 --seed 11 --subspace-dim 4 --plot
```

This estimates the feedback second-moment matrix:

\[
M_\mu=\mathbb E_\mu[\hat b\hat b^\top]
\]

for several initialization distributions. The worst-case expected squared alignment is:

\[
\inf_{\|a\|=1}
\mathbb E_\mu[(a^\top \hat b)^2]
=
\lambda_{\min}(M_\mu).
\]

The prior-free minimax theorem says:

\[
\sup_\mu \lambda_{\min}(M_\mu)=\frac1D,
\]

with equality for isotropic feedback. Outputs are written under `outputs/minimax_initialization/`.

## Functional Capacity Overlap

Run:

```bash
python scripts/functional_capacity_overlap.py --parameters 96 --task-rank 24 --constraint-ranks 0 24 48 72 84 96 --trials 100 --seed 5 --plot
```

This samples a task-sensitive subspace \(S\) of dimension \(d\) and an alignment
constraint subspace \(E\) of dimension \(k\) inside \(\mathbb R^P\). It validates:

\[
\Delta d_{\mathrm{hard}}
=
\max(0,k-(P-d))
\]

and:

\[
\mathbb E[\operatorname{tr}(P_E P_S)]
=
\frac{kd}{P}.
\]

Outputs are written under `outputs/functional_capacity_overlap/`.

## Synthetic MLP FA/BP Trajectories

Run:

```bash
python scripts/trajectory_mlp_fa.py --samples 128 --hidden-widths 24 24 --steps 80 --lr 0.02 --eval-every 10 --feedback-runs 3 --data-seed 3 --init-seed 4 --feedback-seed-start 50 --plot
```

This trains one BP baseline and several FA runs from the same initial weights on
a synthetic regression task. At each checkpoint, the script records:

- training loss;
- full-model cosine between the BP gradient and the FA surrogate gradient at the FA weights;
- hidden-layer-only cosine between the BP and FA gradients, excluding the output layer where gradients are identical;
- layerwise \(Q_l=\cos^2(W_{l+1}^{\top},B_l)\).

Outputs are written under `outputs/trajectory_mlp_fa/`:

- `summary.csv`
- `trajectories.csv`
- `layer_metrics.csv`
- diagnostic plots when `--plot` is set.

## Synthetic Trajectory Ensemble

Run:

```bash
python scripts/trajectory_ensemble.py --architectures 16,16 24,24 32,32 48,48 24,24,24 --samples 256 --steps 200 --lr 0.02 --eval-every 20 --feedback-runs 40 --data-seed 20 --init-seed 30 --feedback-seed-start 1000 --plot
```

This runs one BP baseline and many FA feedback seeds for each architecture,
then aggregates:

- final FA/BP loss gap;
- full and hidden-only gradient cosine;
- initial and final \(Q_l\) means;
- initial and final log-volume capacity cost;
- within-architecture and pooled correlations.

Outputs are written under `outputs/trajectory_ensemble/`:

- `run_summary.csv`
- `architecture_summary.csv`
- `correlations.csv`
- `trajectories.csv`
- diagnostic plots when `--plot` is set.