summaryrefslogtreecommitdiff
path: root/RESULTS.md
blob: 2affc24452bad60dca741a9ace3eb1180f3110fe (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
# SDIL — results log (audited sections report n; timan107 GTX-1080 / ep_pascal)

## 2026-07-21 audit and version boundary

Git was initialized after inheriting the project. The original code/results are preserved at
`215acff`. The following audit findings changed the canonical implementation:

1. `49b86d6`: residual-block local updates now include the branch Jacobian
   `res_alpha=1/sqrt(depth)`. The omission preserved cosine direction but inflated update norms by
   `sqrt(depth)`, confounding cross-depth optimization comparisons.
2. `5255166`: the Harnett baseline is now a per-neuron affine soma-dendrite fit
   `ahat_l = p_l * h_l + b_l`, rather than a dense population map. This follows the paper's
   per-cell least-squares somatic-event to dendritic-event relation.
3. `c2769e6`: causal perturbation targets and predicted residuals are measured on the same
   pre-update network state.
4. `0f107d1` / `e0bfdfd`: simultaneous hidden-layer perturbations reduce calibration from
   quadratic to linear depth work; batching directions makes the reduction visible in GPU wall
   time. `66e3034` fixes diagnostics so raw-apical ablations report the signal actually used.
5. `2f3f44d`: the raw-apical control can be matched per sample to the innovation norm. This
   separates subtraction/direction from a trivial update-magnitude explanation.
6. `6093fe5` / `d6a00aa`: FA, PEPITA, Forward-Forward, and EP were reimplemented against their
   papers and reference code. EP uses persistent free particles; deep PEPITA requires an explicit
   learning rate; FA now logs the signal actually transported through its fixed feedback path.

All older `deepres_*` files predate item 1. Their alignment directions remain informative, but
their accuracy and stability should not be pooled with the corrected runs below.

## Method
SDIL (Somato-Dendritic Innovation Learning), a local non-backprop rule inspired by Harnett 2026.
Per hidden layer l:
- forward: `u_l = W_l h_{l-1}`, `h_l = φ(u_l)`
- apical feedback `a_l = A_l c` (c = output error e = softmax−onehot, broadcast)
- default per-neuron predictor `ĥ_l = p_l ⊙ h_l + b_l`; teaching signal `r_l = a_l − ĥ_l`
- three-factor update `ΔW_l = η (r_l ⊙ φ'(u_l)) h_{l-1}^T`
- `A_l` learned by node perturbation (q_l ≈ −∇_{h_l}L, forward-only, no weight transport),
  following Lansdell, Prakash & Kording's learned synthetic-feedback method
- `P_l` learned on neutral (c=0) periods only, so it strips the soma-predictable nuisance
  without eating the teaching signal.

The scalable perturbation mode injects independent Rademacher interventions at all hidden layers
in the same antithetic forward evaluations. Cross-layer interference adds variance but averages to
zero; it trains all `A_l` using `O(depth)` rather than `O(depth^2)` forward work.

Relative to fixed DFA/FA, the feedback pathway is learned by causal perturbation. That ingredient
is not novel: without traffic and `P`, it is the direct learned-feedback method of
[Lansdell et al.](https://arxiv.org/abs/1906.00889). The candidate SDIL contribution is specifically
the per-cell residual under mixed apical traffic and its neutral-period identification. The full
claim boundary is frozen in `NOVELTY.md`.

## Critical fix
Node-perturbation estimator must divide by σ, not σ² (with δ=σξ, ĝ = δ·ΔL/σ² = ξ·ΔL/σ ⇒ ‖q‖≈‖∇L‖).
The σ² bug made the teaching signal ~100× too large ⇒ effective LR ~100× ⇒ SDIL stalled. Cosine is
magnitude-invariant so alignment looked fine while optimization blew up. Fixing it turned SDIL from
"worse than DFA everywhere" to "matches DFA on MNIST, beats it on FMNIST/CIFAR". (Codex independently
flagged the same σ-scale issue.)

## Verified five-depth scaling (CIFAR-10, width 64, 5 epochs, five seeds)

Mean ± sample SD. Alignment is the mean per-sample cosine over the earliest third of hidden layers
at the final checkpoint. SDIL uses simultaneous perturbations with 16 batched directions.

| depth | method | test acc (%) | early-third alignment | wall time (s) |
|------:|:-------|-------------:|----------------------:|--------------:|
| 5 | BP | 41.546 ± 0.422 | — | 6.9 ± 0.3 |
| 5 | FA | 38.924 ± 0.729 | 0.630 ± 0.022 | 5.7 ± 1.0 |
| 5 | DFA | 37.768 ± 0.654 | 0.514 ± 0.032 | 5.0 ± 0.2 |
| 5 | SDIL | 41.460 ± 0.226 | 0.797 ± 0.018 | 9.8 ± 1.9 |
| 10 | BP | 41.702 ± 0.730 | — | 11.7 ± 0.3 |
| 10 | FA | 39.056 ± 0.557 | 0.610 ± 0.022 | 10.0 ± 2.1 |
| 10 | DFA | 36.716 ± 0.363 | 0.231 ± 0.030 | 9.4 ± 1.7 |
| 10 | SDIL | 41.046 ± 0.222 | 0.496 ± 0.024 | 19.6 ± 4.4 |
| 20 | BP | 41.910 ± 0.560 | — | 20.2 ± 1.1 |
| 20 | FA | 37.448 ± 0.603 | 0.569 ± 0.025 | 18.8 ± 3.6 |
| 20 | DFA | 37.106 ± 0.710 | 0.126 ± 0.023 | 17.9 ± 3.5 |
| 20 | SDIL | 41.288 ± 0.606 | 0.403 ± 0.020 | 31.2 ± 1.5 |
| 30 | BP | 41.908 ± 0.697 | — | 29.8 ± 1.3 |
| 30 | FA | 38.158 ± 0.517 | 0.593 ± 0.014 | 26.9 ± 1.9 |
| 30 | DFA | 36.988 ± 0.524 | 0.087 ± 0.007 | 23.4 ± 0.8 |
| 30 | SDIL | 41.198 ± 0.523 | 0.366 ± 0.007 | 47.8 ± 1.4 |
| 60 | BP | 41.624 ± 0.781 | — | 58.4 ± 1.0 |
| 60 | FA | 38.428 ± 0.387 | 0.577 ± 0.011 | 49.8 ± 1.8 |
| 60 | DFA | 36.898 ± 0.444 | 0.047 ± 0.005 | 47.6 ± 5.6 |
| 60 | SDIL | 41.246 ± 0.323 | 0.336 ± 0.013 | 96.4 ± 12.4 |

Across paired seeds, SDIL changes by `-0.214 ± 0.349` accuracy points from d5 to d60, while the
depth grows 12x and wall time grows 9.87x (geometric mean). Its paired gap to BP is
`0.086 ± 0.346` points at d5 and `0.378 ± 0.927` at d60. The SDIL−DFA advantage stays between
3.69 and 4.35 points at every tested depth. Flattened CIFAR itself is depth-flat, so this establishes
preservation under depth, not a claim that extra depth improves the task.

The mean-based nondominated frontier among local methods is DFA-d5 (`5.01 s`, `37.77%`), FA-d5
(`5.72 s`, `38.92%`), then SDIL-d5 (`9.84 s`, `41.46%`). BP is shown only as a nonlocal reference.
Thus SDIL occupies the high-accuracy end of the audited local-learning frontier; the deeper SDIL
configurations demonstrate stability, but are correctly dominated by its d5 configuration on this
depth-flat task.

SDIL's early-third alignment declines from `0.797` at d5 to `0.336` at d60; it is useful but not
depth-invariant. DFA collapses much further (`0.514` to `0.047`). Residual FA does **not** show that
failure because identity skips transport most of the state gradient exactly (`0.630` to `0.577`),
yet its accuracy stays roughly 2.6–4.5 points below BP. The supported claim is therefore not “all FA
signals lose alignment”: direction, scale, serial transport, and optimization quality must be
separated.

## Harnett-specific residualization (MNIST, d3/w256, 15 epochs, five seeds)

The matched-raw control preserves the raw apical direction but rescales every sample to the
innovation's norm.

| nuisance rho | raw apical acc (%) | matched raw acc (%) | residual acc (%) |
|-------------:|-------------------:|--------------------:|-----------------:|
| 0.00 | 97.450 ± 0.046 | 97.450 ± 0.046 | 97.450 ± 0.046 |
| 0.05 | 96.438 ± 0.077 | 96.980 ± 0.148 | 97.382 ± 0.154 |
| 0.20 | 95.512 ± 0.150 | 95.974 ± 0.144 | 97.370 ± 0.083 |
| 0.50 | 10.380 ± 0.548 | 10.310 ± 0.633 | 97.348 ± 0.133 |

At zero nuisance all three signals are identical. As normal soma-dendrite coupling grows, raw
feedback degrades and eventually reaches chance; residual learning stays at 97.35–97.45%. Matching
the raw signal's norm recovers only part of the low-rho loss and does nothing at rho=0.5. The strong
result therefore depends on subtracting the soma-predictable component, not merely clipping or
normalizing an oversized update. This is the most direct evidence that the Harnett-specific
innovation operation is algorithmically necessary.

## Frozen endogenous mixed-traffic confirmation (MNIST, d3/w256, 15 epochs)

The 95-run C1 panel crossed five paired model seeds, soma and endogenous top-down traffic, two
independently generated projections per traffic family, four signal controls, and a no-traffic
panel. All runs came from clean revision `18c9c1c`. Test accuracy was evaluated once after
training; alignment and traffic diagnostics used a fixed 512-example training prefix. No
intermediate held-out metric or diagnostic entered the trajectories.

| traffic | rho | projection seed | signal | test acc (%) | cos(r,-g) | traffic R2 |
|:--------|----:|----------------:|:-------|-------------:|----------:|-----------:|
| none | 0.0 | 1234 | raw | 97.474 ± 0.122 | +0.443 | — |
| none | 0.0 | 1234 | norm-matched raw | 97.474 ± 0.122 | +0.443 | — |
| none | 0.0 | 1234 | innovation | 97.474 ± 0.122 | +0.443 | — |
| soma | 0.5 | 1234 | raw | 93.546 ± 0.508 | +0.037 | 1.000 |
| soma | 0.5 | 1234 | norm-matched raw | 94.614 ± 0.586 | -0.025 | 1.000 |
| soma | 0.5 | 1234 | innovation | 97.474 ± 0.091 | +0.268 | 1.000 |
| soma | 0.5 | 1234 | task-period predictor | 97.462 ± 0.161 | +0.074 | 1.000 |
| soma | 0.5 | 5678 | raw | 93.026 ± 2.046 | +0.025 | 1.000 |
| soma | 0.5 | 5678 | norm-matched raw | 94.210 ± 0.294 | +0.000 | 1.000 |
| soma | 0.5 | 5678 | innovation | 97.390 ± 0.137 | +0.267 | 1.000 |
| soma | 0.5 | 5678 | task-period predictor | 97.450 ± 0.050 | +0.067 | 1.000 |
| top-down | 0.2 | 1234 | raw | 79.674 ± 9.332 | +0.077 | 1.000 |
| top-down | 0.2 | 1234 | norm-matched raw | 93.258 ± 1.029 | -0.063 | 1.000 |
| top-down | 0.2 | 1234 | innovation | 94.778 ± 0.329 | +0.205 | 0.947 |
| top-down | 0.2 | 1234 | task-period predictor | 94.052 ± 0.220 | +0.105 | 0.905 |
| top-down | 0.2 | 5678 | raw | 76.828 ± 16.956 | +0.121 | 1.000 |
| top-down | 0.2 | 5678 | norm-matched raw | 91.240 ± 2.770 | -0.067 | 1.000 |
| top-down | 0.2 | 5678 | innovation | 94.196 ± 1.106 | +0.217 | 0.963 |
| top-down | 0.2 | 5678 | task-period predictor | 93.984 ± 0.578 | +0.126 | 0.922 |

The innovation-minus-matched gains were `+2.860` and `+3.180` points for the soma projections,
and `+1.520` and `+2.956` points for the top-down projections. Thus three of four realizations
passed all frozen per-realization criteria. The top-down seed-1234 realization passed traffic R2
(`0.947`) and alignment (`+0.205`) but missed the accuracy threshold. No-traffic innovation was
identical to both controls. **The full C1 confirmation gate fails** because it required every
realization to gain at least two points. This is a near miss with a precise failure mode, not a
passed result; the projection and all five finite trajectories remain in the record.

## FashionMNIST C1 recovery screen (validation only)

After the MNIST failure, a 27-run recovery protocol was frozen at clean revision `26407ca` before
any FashionMNIST result was generated. It used one model seed and a fixed 5,000-example stratified
training-only validation split (hash `c5f29347...ddf58`). The test set was never evaluated. All
diagnostics were computed once on the training prefix after training.

| traffic | rho | projection | matched raw val (%) | innovation val (%) | gain (points) | no-traffic gap | alignment | R2 | gate |
|:--------|----:|-----------:|--------------------:|-------------------:|--------------:|---------------:|----------:|---:|:----:|
| soma | 0.5 | 1234 | 76.44 | 87.50 | +11.06 | -0.50 | +0.543 | 1.000 | pass |
| soma | 0.5 | 5678 | 81.60 | 88.82 | +7.22 | +0.82 | +0.532 | 1.000 | pass |
| top-down | 0.2 | 1234 | 79.66 | 80.96 | +1.30 | -7.04 | +0.467 | 0.975 | fail |
| top-down | 0.2 | 5678 | 76.64 | 82.14 | +5.50 | -5.86 | +0.462 | 0.965 | fail |
| top-down | 0.5 | 1234 | 67.80 | 75.14 | +7.34 | -12.86 | +0.449 | 0.998 | fail |
| top-down | 0.5 | 5678 | 67.04 | 77.78 | +10.74 | -10.22 | +0.368 | 0.995 | fail |

The no-traffic raw, norm-matched, and innovation controls were identical at 88.00%. The selection
rule required both projections of a traffic condition to gain at least 2 points, remain within 3
points of no traffic, have R2 at least 0.5, and align at least 0.1. Soma traffic passed, but no
top-down level passed. **The recovery screen therefore fails and FashionMNIST test confirmation
is prohibited.** Innovation substantially mitigates strong contextual traffic, but a high
soma-conditioned R2 and positive gradient alignment do not guarantee that the remaining
innovation is purely instructional. This narrows the algorithmic claim and motivates an explicit
model of context, rather than treating every somatically unpredictable apical component as error.

## Legacy accuracy (pre-audit implementation)
| task                         | BP        | DFA   | SDIL          |
|------------------------------|-----------|-------|---------------|
| MNIST  d3 / d5 / d7 / d10    | 98.3      | 97.3 / 97.2 / 97.1 / 96.8 | 97.4 / 97.3 / 97.1 / 97.1 |
| FashionMNIST d3 / d5 / d7 / d10 | 88–89  | 86.3 / 87.1 / 86.2 / 86.9 | 88.4 / 87.6 / 87.3 / 88.1 |
| CIFAR-10 (no-BN residual) d5/10/20 | 44.2 / 44.8 / 45.5 | 40.6 / 39.9 / 37.4 | 42.4 / 42.6 / 42.4 |

SDIL matches DFA on easy MNIST and **beats DFA on FashionMNIST (+1–2 pts) and CIFAR (+2 to +5 pts)**.
On CIFAR the SDIL−DFA gap GROWS with depth (+1.8 → +2.7 → +5.0 at d5/10/20): DFA degrades with depth,
SDIL holds.

## Endogenous top-down traffic timescale screen (validation only)

This development screen used one model seed, a fixed 5,000-example stratified MNIST validation
split drawn only from training data, and never evaluated the test set. All 18 runs share validation
index hash `c4cec216...2e3d`, clean source revision `17bcd97`, and the same 15-epoch training
protocol. The grid crossed `eta_P in {0.002, 0.01, 0.05}`, neutral warmup steps in
`{0, 200, 1000}`, and endogenous top-down traffic scales in `{0.2, 0.5}`.

| eta_P | warmup steps | rho=0.2 last val (%) | rho=0.5 last val (%) | mean (%) | worst (%) |
|------:|-------------:|---------------------:|---------------------:|---------:|----------:|
| 0.002 | 0 | 91.50 | 77.24 | 84.37 | 77.24 |
| 0.002 | 200 | 91.02 | 89.10 | 90.06 | 89.10 |
| 0.002 | 1000 | 93.10 | 88.82 | 90.96 | 88.82 |
| 0.010 | 0 | 94.94 | 92.40 | 93.67 | 92.40 |
| 0.010 | 200 | 92.88 | 91.12 | 92.00 | 91.12 |
| 0.010 | 1000 | 95.04 | 91.06 | 93.05 | 91.06 |
| 0.050 | 0 | 96.52 | 91.40 | 93.96 | 91.40 |
| **0.050** | **200** | **94.20** | **95.26** | **94.73** | **94.20** |
| 0.050 | 1000 | 94.70 | 93.62 | 94.16 | 93.62 |

The selected development setting is `eta_P=0.05`, 200 warmup steps: it has the highest mean
last-epoch validation accuracy and the highest worst-traffic endpoint. Its final mean traffic R2 is
0.802 at rho=0.2 and 0.999 at rho=0.5. This reverses the initial assumption that `P` must be much
slower than forward learning: once teaching and traffic are separated by neutral periods, fast
tracking is beneficial, while a long warmup is not consistently useful. This is a frozen
validation choice for the multi-seed C1 experiment, not evidence that C1 has passed.

## Causal-calibration query screen (validation only)

This seed-0 development screen used CIFAR-10 d20/w64 for five epochs and a 5,000-example
training-only validation split (hash `8328b206...515b`). All eight runs came from clean revision
`1a22a99`; the test set was never evaluated. Gain retention is
`(candidate - DFA) / (K16/e4 - DFA)` at the final validation endpoint.

| method | directions K | every | val (%) | batch loss queries | query reduction | calibration work reduction | gain retention |
|:-------|-------------:|------:|--------:|-------------------:|----------------:|---------------------------:|---------------:|
| DFA | — | — | 36.66 | 0 | — | — | — |
| learned NP | 16 | 4 | 37.46 | 14,080 | 1.0x | 1.0x | 1.000 |
| **learned NP** | **1** | **4** | **38.08** | **880** | **16.0x** | **11.0x** | **1.775** |
| learned NP | 2 | 8 | 37.32 | 880 | 16.0x | 13.2x | 0.825 |
| learned NP | 4 | 16 | 35.98 | 880 | 16.0x | 14.7x | -0.850 |
| learned NP | 1 | 8 | 36.70 | 440 | 32.0x | 22.0x | 0.050 |
| learned NP | 1 | 16 | 36.52 | 220 | 64.0x | 44.0x | -0.175 |
| learned NP | 2 | 32 | 36.78 | 220 | 64.0x | 52.8x | 0.150 |

`K=1/every=4` was frozen for confirmation: it exceeded the C3 10x-query/90%-retention threshold in
validation and was the only candidate that improved on the reference endpoint. The ablation also
separates directions from update frequency: spending the same 880 queries as `K=2/every=8` or
`K=4/every=16` is worse, so frequent noisy regression targets are more useful than precise sparse
targets. GTX-1080 wall time remains approximately 15–23 seconds across the grid despite the 11x
calibration-work change, demonstrating why the Pareto audit needs hardware-independent work in
addition to wall time.

The frozen full-training/test confirmation used five paired seeds at clean revision `6689994`, with
no intermediate test evaluation or diagnostics:

| method | n | test acc (%) | batch loss queries | total training forward-eq examples | peak allocated (MiB) | wall (s) |
|:-------|--:|-------------:|-------------------:|-----------------------------------:|---------------------:|---------:|
| DFA | 5 | 37.106 ± 0.710 | 0 | 250,000 | 739.8 ± 0.0 | 17.5 ± 2.0 |
| K16/e4 | 5 | 38.490 ± 0.658 | 15,648 | 2,313,952 | 803.5 ± 0.0 | 22.6 ± 4.7 |
| **K1/e4** | **5** | **38.668 ± 0.589** | **978** | **437,632** | **739.8 ± 0.0** | **20.8 ± 0.8** |

The mean paired gain over DFA is 1.384 points for K16/e4 and 1.562 points for K1/e4, giving
112.9% retention. The selected protocol reduces logical perturbation queries 16x, calibration
forward-equivalent work 11x, and total training forward-equivalent work 5.3x. Its peak allocation
falls to the DFA level. **Gate C3 passes.** Wall time improves only modestly because fixed/data/GPU
under-utilization costs dominate this small network, so it is reported but not substituted for the
hardware-independent result.

## Legacy credit assignment / depth utility
Per-layer alignment cos(r_l, −∇_{h_l}L), early-third average, CIFAR no-BN residual:
| depth | DFA early-align | SDIL early-align |
|-------|-----------------|------------------|
| 5     | 0.55            | 0.83             |
| 10    | 0.36            | 0.86             |
| 20    | 0.25            | 0.88             |
| 30    | 0.19            | 0.87             |

DFA's teaching signal to early/deep layers collapses toward random (~0.05–0.19) as depth grows;
SDIL stays ~0.85 and is depth-invariant. This is the "does credit reach early layers" metric that
DFA fails and SDIL passes — the core advantage.

## C2 BP useful-depth screen (validation only)

The frozen candidate uses a two-level Telgarsky tent-map task, width-8 ReLU residual networks,
and three independently generated task/data seeds. Student initialization is fixed at seed 0 for
this shape screen. All 15 runs came from clean revision `76688af`, used stratified validation
subsets drawn only from generated training data, and never evaluated the independent test sets.

| depth | validation acc (%) | task-seed endpoints (%) | final-third lesion drops (points) |
|------:|-------------------:|:------------------------|:---------------------------------|
| 1 | 63.933 ± 10.272 | 75.25, 55.20, 61.35 | — |
| 2 | 60.883 ± 7.243 | 57.85, 69.15, 55.65 | +14.25, +24.75, +15.10 |
| 3 | 97.250 ± 1.097 | 96.30, 97.00, 98.45 | +58.80, +58.75, +60.15 |
| 4 | 96.883 ± 1.457 | 97.35, 98.05, 95.25 | +40.85, +35.05, +34.85 |
| 6 | 90.733 ± 7.743 | 96.35, 93.95, 81.90 | +40.05, +40.35, -2.10 |

For the preselected d1-to-d4 comparison, paired BP gains are `+22.10`, `+42.85`, and `+33.90`
points (mean `+32.95`). The corresponding final-third lesion drops average `36.92` points and are
positive for every task seed. **The BP useful-depth sub-gate passes.** The nonmonotonic curve is
also informative: d2 can optimize worse than d1, and d6 is less stable than d3/d4. The claim is
therefore about a verified useful-depth pair, not monotonic improvement from adding arbitrary
blocks.

## C2 local-rule panel (validation only)

The next frozen panel crossed BP, FA, DFA, and no-traffic SDIL at d1/d4 over three task seeds and
five student seeds (120 runs). SDIL used the C3-selected `K=1/every=4` calibration schedule. All
runs came from clean revision `51301f6`, evaluated validation only after training, and computed
alignment and lesion diagnostics on the fixed training prefix. The independent synthetic test
sets remain unevaluated.

| method | depth | validation acc (%) | d1-to-d4 gain | d4 lesion drop |
|:-------|------:|-------------------:|--------------:|---------------:|
| BP | 1 | 62.757 ± 8.111 | — | — |
| BP | 4 | 92.453 ± 8.834 | +29.697 | +30.157 |
| FA | 1 | 58.600 ± 4.839 | — | — |
| FA | 4 | 77.147 ± 16.082 | +18.547 | +16.853 |
| DFA | 1 | 55.517 ± 11.161 | — | — |
| DFA | 4 | 73.193 ± 13.987 | +17.677 | +22.650 |
| SDIL | 1 | 65.347 ± 6.516 | — | — |
| SDIL | 4 | 71.587 ± 15.307 | +6.240 | +12.787 |

SDIL recovers only `21.0%` of BP's depth gain, versus `62.5%` for FA and `59.5%` for DFA, and its
d4 endpoint trails the stronger comparator by `5.560` points. Its final-third lesion reduces
accuracy in 11/15 runs, so the model often uses the late block even though its learned credit is
poor. **The C2 local validation gate fails and no test confirmation is allowed.** Additional
validation scratch found that K=4 or K=16 perturbations, more frequent events, larger `eta_A`, and
normalized local deltas do not remove the basin variability. This points to a representational
limitation of the fixed linear `A_l c` vectorizer: on a compositional target, the correct hidden
direction depends on neural state as well as the two-dimensional output error.

## C2 bounded context-vectorizer recovery (validation only)

To test the identified representation bottleneck, the apical vectorizer was extended to
`A_l c + G_l vec(c outer tanh(h_top))`. The gated path is zero-initialized, trained by the same
local perturbation regression, and vanishes in neutral periods. After a task-seed-0 development
screen selected `eta_A=0.01`, K1/e4, an independent frozen panel used new task seeds 3--5 and five
student seeds, again crossing BP, FA, DFA, and SDIL at d1/d4. All 120 runs came from clean revision
`b18227f`; independent test sets were never evaluated.

| method | depth | validation acc (%) | d1-to-d4 gain | d4 lesion drop |
|:-------|------:|-------------------:|--------------:|---------------:|
| BP | 1 | 68.633 ± 5.351 | — | — |
| BP | 4 | 97.187 ± 1.532 | +28.553 | +34.170 |
| FA | 1 | 52.640 ± 4.528 | — | — |
| FA | 4 | 80.763 ± 20.592 | +28.123 | +20.313 |
| DFA | 1 | 56.313 ± 8.374 | — | — |
| DFA | 4 | 77.097 ± 17.667 | +20.783 | +25.393 |
| context-SDIL | 1 | 63.977 ± 8.636 | — | — |
| context-SDIL | 4 | 79.120 ± 22.765 | +15.143 | +27.150 |

Context conditioning raises SDIL's recovery of BP depth gain from the linear panel's `21.0%` to
`53.0%`, confirming that state dependence was a real bottleneck. It is still below the frozen
70% requirement, while FA recovers `98.5%`; context-SDIL's d4 mean is `1.643` points below FA.
Exactly 10/15 SDIL pairs gain with depth and 10/15 are hurt by the lesion, but one d4 run ends with
nonfinite loss and chance accuracy. **The independent recovery gate fails and C2 test remains
prohibited.** A richer vectorizer can improve mean credit without curing coupled A/W basin
instability; this is now a documented limitation rather than an oral-level scaling result.

## C2 direct node-perturbation causal diagnosis (validation only)

The failed context-vectorizer panel leaves two distinct explanations: the local eligibility/update
rule may be unable to solve the compositional task, or the learned apical vectorizer may fail to
amortize a causal direction that would work if it were available. To separate them, direct node
perturbation replaces every hidden teaching vector by an unamortized antithetic layerwise estimate
on every update. A task-seed-0 development screen selected `eta=0.03` and K16; the independent
panel then reused the exact untouched validation splits for task seeds 3--5 and crossed d1/d4 with
five model seeds. All 30 runs came from clean revision `e8d698e`; test sets remain unevaluated.

| method | depth | validation acc (%) | d1-to-d4 gain | d4 lesion drop |
|:-------|------:|-------------------:|--------------:|---------------:|
| BP | 1 | 68.633 ± 5.351 | — | — |
| BP | 4 | 97.187 ± 1.532 | +28.553 | +34.170 |
| context-SDIL | 1 | 63.977 ± 8.636 | — | — |
| context-SDIL | 4 | 79.120 ± 22.765 | +15.143 | +27.150 |
| direct NP | 1 | 65.703 ± 8.800 | — | — |
| direct NP | 4 | 96.150 ± 1.719 | +30.447 | +25.743 |

Direct NP recovers `106.6%` of BP's mean depth gain. All 15 paired runs improve with depth, 14/15
are harmed by the frozen final-block lesion, all final losses are finite, and its d4 endpoint beats
context-SDIL by `17.030` points and the stronger FA/DFA endpoint by `15.387` points. **The frozen
causal diagnosis passes: the local eligibility rule and a forward-only causal direction are
sufficient here; learning a reliable amortized apical vectorizer is the present bottleneck.**

This is a diagnosis and upper-cost baseline, not the proposed scalable method. At d4 it uses a mean
`82,560,000` per-example scalar loss evaluations and `43,757,037` forward-equivalent examples,
`68.4x` ordinary training work, with `108.2` seconds mean wall time. The next recovery must retain
the vectorizer and disclose how much of this query burden it avoids.

## Paper-faithful non-backprop baselines

`PETITE` in the initial request is treated as **PEPITA**. Implementations and hyperparameters were
checked against the [FA paper](https://www.nature.com/articles/ncomms13276),
[PEPITA paper](https://proceedings.mlr.press/v162/dellaferrera22a.html) and
[author code](https://github.com/GiorgiaD/PEPITA),
[Forward-Forward paper](https://www.cs.toronto.edu/~hinton/absps/FFXfinal.pdf) and
[reference implementation](https://github.com/mpezeshki/pytorch_forward_forward), and the
[EP paper](https://www.frontiersin.org/journals/computational-neuroscience/articles/10.3389/fncom.2017.00024/full)
and [author code](https://github.com/bscellier/Towards-a-Biologically-Plausible-Backprop).

Native protocol results are not an equal-compute comparison:

| method | architecture | budget | n | last test acc (%) | wall time (s) |
|:-------|:-------------|:-------|--:|------------------:|--------------:|
| PEPITA | d3/w256 | 60 epochs | 5 | 91.094 ± 0.214 | 89.2 ± 3.5 |
| Forward-Forward | d3/w256 | 15 epochs/layer | 5 | 93.558 ± 0.232 | 68.2 ± 0.8 |
| EP | d1/w500 | 25 epochs | 5 | 96.926 ± 0.238 | 800.5 ± 53.3 |
| EP | d2/w500 | 60 epochs | 5 | 89.526 ± 4.461 | 17530.4 ± 613.6 |

PEPITA uses its two forward presentations, He-uniform initialization, shared 10% dropout mask, and
an explicit deep-network learning rate of 0.001. Forward-Forward trains greedy local goodness layers
and has no classifier readout. EP uses the original hard-sigmoid energy dynamics, random beta sign,
and persistent training particles on the fixed first 50k examples. Its published d1 schedule uses
20 free steps, 4 nudged steps, `beta=0.5`, and layer rates `[0.1, 0.05]`; d2 uses 150 free steps,
6 nudged steps, `beta=1`, and rates `[0.4, 0.1, 0.01]`. Evaluation starts fresh zero particles and
runs the full free relaxation. The shallow EP protocol is the strongest native baseline here. The
canonical deep protocol is much more expensive and highly initialization-sensitive in these runs.

### EP exact-architecture comparison

All methods below use the same first 50k MNIST examples and the same epoch budget within each exact
forward architecture: 25 epochs for d1/w500 (397,510 parameters) and 60 for d2/w500 (648,010
parameters). EP retains its native raw-pixel hard-sigmoid energy protocol; BP/DFA/SDIL retain the
shared z-score/tanh feedforward protocol.

| depth | method | n | last test acc (%) | best test acc (%) | wall time (s) |
|:------|:-------|--:|------------------:|------------------:|--------------:|
| d1 | BP | 5 | 98.240 ± 0.073 | 98.272 ± 0.080 | 15.0 ± 0.5 |
| d1 | DFA | 5 | 97.662 ± 0.078 | 97.700 ± 0.074 | 10.6 ± 1.6 |
| d1 | EP | 5 | 96.926 ± 0.238 | 97.198 ± 0.147 | 800.5 ± 53.3 |
| d1 | SDIL | 5 | 97.596 ± 0.089 | 97.638 ± 0.075 | 15.5 ± 0.2 |
| d2 | BP | 5 | 98.288 ± 0.030 | 98.326 ± 0.036 | 45.1 ± 0.8 |
| d2 | DFA | 5 | 97.590 ± 0.046 | 97.622 ± 0.036 | 34.6 ± 2.8 |
| d2 | EP | 5 | 89.526 ± 4.461 | 90.028 ± 4.272 | 17530.4 ± 613.6 |
| d2 | SDIL | 5 | 97.516 ± 0.067 | 97.562 ± 0.058 | 58.5 ± 4.0 |

The paired SDIL−EP last-epoch gap is `0.670 ± 0.253` points at d1 and `7.990 ± 4.459` at d2;
SDIL's geometric-mean wall-time speedup is 51.7x and 299.8x, respectively. The five d2 EP outcomes
are 97.46%, 87.56%, 88.07%, 87.77%, and 86.77%. They are finite completed trajectories, not crashes:
one seed reaches the shallow-level basin while four do not. This is evidence of initialization
sensitivity under the reproduced canonical schedule, not evidence that EP is intrinsically
unstable across implementations. Reproducing the basin split in the original Theano revision or a
strong modern EP variant remains necessary. DFA is also strong on easy MNIST, so the scaling claim
comes from the five-depth CIFAR panel; this comparison establishes that SDIL need not trade away
accuracy to obtain a large cost advantage over canonical EP.

## How to run
`experiments/run.py --mode {bp,fa,dfa,sdil} --dataset {mnist,fmnist,cifar10} --depth D --residual {0,1} --act {tanh,gelu,silu,relu}`
Batteries: `experiments/run_v2.sh <ds> "<depths>" <res> <act> "<seeds>" <ep> <pfx>`.
Five-depth scaling: `experiments/figure_scaling_sweep.sh`; canonical local baselines:
`experiments/baseline_sweep.sh`; original d1 EP:
`experiments/ep_original_sweep.sh`; exact d1 comparison: `experiments/ep_archmatch_sweep.sh`;
canonical d2 comparison: `experiments/ep_depth2_sweep.sh`; near-parameter comparison:
`experiments/ep_matched_sweep.sh`.
C2 direct causal diagnosis: `experiments/c2_nodepert_validation.sh` followed by
`python experiments/analyze_c2_nodepert_validation.py`.
Audited aggregate tables: `experiments/analyze_verified.py`. The claim-locked finalizer is
`experiments/finalize_claims.sh`; it rejects missing/dirty inputs, regenerates both main figures,
their captions, source-hash manifest, and `results/audited_tables.md`, then runs all smoke checks.

## Open items

- Treat learned direct node-perturbation feedback (Lansdell et al.) as the exact backbone baseline,
  and BurstCCN as the strongest dendritic/scaling baseline; do not attribute their ingredients to
  SDIL.
- Compare simultaneous calibration at fixed loss-evaluation budgets and sweep directions
  (4/8/16/32).
- Validate SDIL on a genuinely depth-necessary compositional task and a CNN; flattened CIFAR does
  not gain accuracy with depth.
- Diagnose why residual FA has high directional cosine but lower accuracy: log update-norm ratios,
  parameter-space cosine, and serial feedback latency/cost.
- Reproduce the d2 EP basin sensitivity in the original Theano revision and/or a strong modern EP
  variant without selectively removing seeds.
- Tune SDIL schedules using a held-out validation split, never the test trajectory, then freeze the
  protocol before expanding the EP comparison.
- Extend PEPITA and Forward-Forward beyond shallow MNIST before making any scaling claim about them.