summaryrefslogtreecommitdiff
path: root/notes
diff options
context:
space:
mode:
Diffstat (limited to 'notes')
-rw-r--r--notes/36_evidence_ledger.md9
-rw-r--r--notes/38_mnist_real_data_validation.md99
2 files changed, 106 insertions, 2 deletions
diff --git a/notes/36_evidence_ledger.md b/notes/36_evidence_ledger.md
index 63e797d..14af7df 100644
--- a/notes/36_evidence_ledger.md
+++ b/notes/36_evidence_ledger.md
@@ -100,8 +100,13 @@ standard lazy approximation whose breakdown is *measured* (item 3c below).
overprediction (+0.0026), FA near-unbiased (+0.0005); cause: drift velocity
decays after s (alpha_t < t/s; e.g. true alpha_50 = 1.46 (BP) / 1.17 (FA)
vs linear 2.5). Honest error story, no fitted correction.
-5. **MNIST replication (note 38)**: same protocol, width 256, depth 2,
- N in {128,256,512}, 48 traj: see `real_data_validation_mnist/`.
+5. **MNIST replication (note 38)**: same protocol, widths {128,256} x
+ N {128,256,512} x 2 inits x 8 feedbacks = 96 traj, lazy step
+ (`eta lam_max/N = 0.05`): velocity MAE **0.00196** (= 3.9% of mean gap),
+ bias -0.0006, corr 0.911 (gap range only 0.040-0.063, so MAE/bias are the
+ operative stats); fixed K(0) MAE 0.0061 biased high — same bracket as
+ synthetic. At eta lam_max/N = 0.3 the lazy chain breaks as expected
+ (scope boundary, recorded in note 38).
### (d) The soft ramp + closed form (T7)
diff --git a/notes/38_mnist_real_data_validation.md b/notes/38_mnist_real_data_validation.md
new file mode 100644
index 0000000..011f335
--- /dev/null
+++ b/notes/38_mnist_real_data_validation.md
@@ -0,0 +1,99 @@
+# MNIST Real-Data Validation (e0 Theorem + Estimator)
+
+Script: `scripts/real_data_validation.py`. Outputs:
+`outputs/real_data_validation_mnist/`. Purpose: remove the "synthetic-only"
+objection for the two core no-fit claims by replicating them on MNIST-subset
+MLPs. Run 2026-06-09.
+
+Setup: flattened standardized MNIST (rows scaled by `1/sqrt(784)` so kernel
+scales match the synthetic runs), one-hot regression targets, squared loss
+`L = ||r||^2/(2N)`, float64, full-batch SGD, project init/feedback
+conventions (`downstream_capacity_sweep` gradients are reused directly).
+
+Tangent operators are computed by the layerwise gram identity instead of an
+explicit `(NC) x P` Jacobian; `--self-test` checks the identity against
+autograd Jacobians on a tiny net: max errors `<= 5.3e-15` for
+`K_BP = J J^T`, `K_FA = J Jtilde^T`, and `J^T r` vs the production gradient
+code (machine precision).
+
+## Part A — initial-erosion theorem on real data
+
+`E_B[e_0 | W, r] = hidden BP speed share` (note 29 / T5). For each config the
+prediction is one BP backward pass; the measurement is 512 independent
+feedback draws, each through the **real FA backward pass**.
+
+16 configs: width {64, 256} x hidden depth {1, 2} x 4 init seeds, N=256.
+
+| depth | width | hidden_share range (4 inits) | max abs diff | typical stderr |
+|---|---|---|---|---|
+| 1 | 64 | 0.727 - 0.819 | 0.0067 | 0.005 |
+| 1 | 256 | 0.414 - 0.451 | 0.0018 | 0.0014 |
+| 2 | 64 | 0.684 - 0.792 | 0.0089 | 0.004 |
+| 2 | 256 | 0.395 - 0.531 | 0.0023 | 0.0012 |
+
+All 16 panels: `|E_B[e_0] - hidden_share|` in **0.0001-0.0089**, within ~2
+Monte-Carlo stderr. The predicted share spans **0.39-0.82** across configs and
+the measured mean tracks the diagonal across that whole range
+(`e0_mnist_calibration.png`, `e0_mnist_rows.csv`).
+
+Real-data reading: on 784-dim inputs the hidden share — the predicted initial
+FA burden — is large at small width (~0.8) and decreases as the output layer
+(fan-in = width) grabs more BP energy. The theorem prices this with no fit.
+
+## Part B — early operator-velocity estimator on real data
+
+Protocol mirrors notes 15-17: full-batch SGD to `T=50`, early snapshot at
+`s=20`, predictors rolled through `r_{t+1} = (I - eta K_hat_t/N) r_t` with
+`K_hat_t = K_0` (fixed) or `K_0 + (t/s)(K_s - K_0)` (linear velocity);
+empirical gap = trained `L_FA(T) - L_BP(T)`. No fitted scalar anywhere.
+
+96 trajectories: width {128, 256} x N {128, 256, 512} x 2 inits x 8 feedback
+seeds. Step size set per config as `eta = 0.05 N / lam_max(K_BP_0)` (lazy
+regime; see sensitivity below).
+
+| predictor | MAE | bias | corr |
+|---|---|---|---|
+| fixed K(0) | 0.00613 | +0.00529 | 0.606 |
+| **linear velocity** | **0.00196** | **-0.00062** | **0.911** |
+
+Per-config (16 traj each):
+
+| w | N | gap_emp mean +/- std | velocity MAE | fixed MAE |
+|---|---|---|---|---|
+| 128 | 128 | 0.0570 +/- 0.0038 | 0.0014 | 0.0121 |
+| 128 | 256 | 0.0517 +/- 0.0043 | 0.0016 | 0.0092 |
+| 128 | 512 | 0.0482 +/- 0.0040 | 0.0021 | 0.0075 |
+| 256 | 128 | 0.0533 +/- 0.0026 | 0.0026 | 0.0040 |
+| 256 | 256 | 0.0495 +/- 0.0023 | 0.0020 | 0.0020 |
+| 256 | 512 | 0.0455 +/- 0.0025 | 0.0021 | 0.0020 |
+
+Readings:
+
+- Relative accuracy: velocity MAE is **3.9% of the mean gap** with near-zero
+ bias; it beats fixed K(0) everywhere (4-9x at width 128, parity-to-better at
+ width 256 where the kernel drifts less — wider = lazier, as theory says).
+- Bias signs reproduce the synthetic bracket (fixed biased high +0.0053,
+ velocity slightly low -0.0006; notes 16-17).
+- corr 0.911 vs synthetic 0.999: the MNIST gap only spans 0.040-0.0625 across
+ all 96 trajectories (the sweep moves the gap ~10-20%, vs orders of magnitude
+ in the synthetic margin sweeps), so correlation is a weaker statistic here;
+ MAE/bias are the meaningful numbers. State it this way in the paper.
+
+## Step-size sensitivity (kept as a scope datum)
+
+A first run used `eta = 0.3 N / lam_max` (per-step top-mode decay 30%):
+fixed MAE 0.069 corr 0.37; velocity MAE 0.030 corr 0.68 — both overpredict,
+velocity still 2.3x better. At `0.05` (matching the synthetic regime's
+`eta lam_max/N ~ 0.04`) the chain holds. This is the expected lazy-regime
+boundary of A1, not a new failure mode: the estimator's stated scope is the
+regime where one-step linearization is accurate.
+
+## Paper use
+
+- Part A => extends Fig "e0 calibration" with a real-data panel (F5 in the
+ ledger): theorem holds on MNIST across a 0.39-0.82 share range.
+- Part B => real-data estimator scatter (`estimator_mnist_scatter.png`):
+ no-fit finite-time prediction at ~4% relative error.
+- Limitations sentence upgrade: "synthetic only" becomes "synthetic +
+ MNIST-subset MLPs; conv/transformers and stochastic optimization remain
+ future work."