1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
|
# SDIL ICLR 2027 evidence roadmap
This document predeclares the evidence gates used to decide which claims survive into a paper.
It is intentionally stricter than a list of experiments: a failed gate narrows the claim rather
than triggering selective seed removal or post-hoc protocol changes.
Paper-level progress is tracked independently in `REVIEW_SCORECARD.md`. Its ICLR-style score may
change only after a frozen result is complete and audited; partial jobs and development pilots do
not count.
## Order of work
1. **Accept bar:** establish a load-bearing innovation mechanism, useful-depth scaling, fair cost,
theory, and the closest baselines.
2. **Oral bar B:** connect the mechanism to the signatures and new predictions motivated by
Harnett et al.'s somato-dendritic residuals.
3. **Oral bar A:** demonstrate the same mechanism in standard CNN/ResNet families. The runner is
developed early so long jobs can use otherwise idle, explicitly authorized GPUs without
delaying stages 1–2.
The preregistered oral-B task, structural screen, confirmation seeds, signature thresholds, and
phase-specific lesion are specified in `ORAL_B.md`. They were frozen while C4 author baselines
were still running and before any continuous-BCI result was generated. Oral-B confirmation was
sequenced after C4 and the D4 accept gate; the realized original and recovery branches are
reported below.
**Oral-B development status: failed.** The complete frozen 12-candidate by three-task-seed screen
found no eligible mechanism. High-learning-rate variants learned the BCI reliably, and several
innovation/decoder signatures were strong, but every run had the wrong P+/P- error-change
sign-inversion and error magnitude remained more predictive than temporal error change. Per the
preregistered stop rule, environment seeds 10--15 and all confirmation model seeds remain
untouched. This falsifies the current model's claimed emergence of the Harnett derivative signature;
it cannot be repaired by reporting only its successful learning or decoder results.
**Oral-B recovery R0/R1 status: mechanics and development passed.** Reinspection
of Fig. 5 and Extended Data Fig. 13 identifies a structural target conflict:
the failed vectorizer regressed `e_t s_i`, so its role contrast encoded current
error magnitude rather than the sign of error change. `ORAL_B_RECOVERY.md`
factorizes a new candidate into a perturbation-learned causal role and the
within-episode performance innovation `|e_(t-1)|-|e_t|`. Deterministic checks
recover the role at cosine `0.9963` and change the controlled sign index from
negative to `+0.100`. After D4 passes, the frozen two-rate, three-task-seed R1
screen selects `eta=0.1`; all 30 seed-level checks pass and worst-task final
success is `98.05%`. R1 is development evidence and leaves the score at 7.
**Oral-B recovery R2 status: failed; branch closed.** The untouched 6-task by
5-model confirmation retains `99.53%` mean final success, `90.45` points of
learning gain, strong fixed-vectorizer and plasticity-lesion margins, mean
role cosine `0.9395`, positive sign inversion in all 30 records, and a `0.6389`
velocity-over-error advantage. Residual decorrelation and decoder-distance
prediction also pass. The joint biological gate fails because
surrounding-state accuracy is `54.51%` versus the 55% mean threshold, residual
outcome accuracy is only `47.33%`, residuals trail soma outcome decoding by
`4.32` points, and longitudinal prediction is `-0.013`. The score remains 7,
oral-A stays sealed, and no threshold repair or replacement confirmation is
permitted. Because `kappa=0`, neither R1 nor R2 could support online control.
**Oral-B-v2 development path: two failures retained.** `ORAL_B_V2.md` adds an
explicit terminal reward/timeout phase, a local TD critic, and eligibility
traces without changing the old R2. Its complete 24-record development grid
fails from a cold start: the dense performance innovation was scaled to
one-quarter of the validated rule and no candidate learns. A separately frozen
unit-scale recovery restores 100% task performance and passes every mechanism
gate on task seeds 23 and 24; task seed 25 passes 17/18 gates but succeeds on
98.96% of a fixed absolute target ladder. That class-balance failure is also
retained and does not open confirmation.
**Oral-B-v2 calibrated R1/R2 status: passed.** The final branch changes no
learning parameter. Five target levels are fixed by cursor-maximum quantiles
on 512 separate outcome-free calibration trials, then evaluated on independent
trajectories. Fresh development seeds 26--28 pass all 18/18 checks. The
untouched 6-task by 5-model confirmation then passes every learning,
innovation, network-prediction, outcome, lesion, and task-cluster confidence
gate: final success is `100%`, learning gain `98.80` points, fixed-role gap
`99.97` points, role cosine `0.9761`, residual--soma correlation `0.0631`,
surrounding accuracy `54.37%` (lower bound `54.26%`), velocity advantage
`0.6384`, terminal outcome accuracy `99.83%` (lower bound `99.68%`), acute
outcome-lesion separation drop `0.4000` (lower bound `0.3886`), critic
expectedness `0.3187` (lower bound `0.2814`), and 30/30 positive signs. The
formal milestone rises from 7 to 8. The old oral-A gate remains closed; only a
new independently frozen oral-A-v2 protocol may now run.
## Frozen accept claims and gates
### C1. Innovation is necessary under naturally mixed apical traffic
**Status: failed the frozen MNIST confirmation on 2026-07-22.** Three of four traffic
realizations passed, but top-down projection seed 1234 produced a `+1.520`-point innovation gain
over norm-matched raw feedback, below the predeclared `+2`-point threshold. The other top-down
projection gained `+2.956` points and the two soma-traffic projections gained `+2.860` and
`+3.180` points. All four passed the traffic-R2 and alignment requirements, and the no-traffic
control had exactly zero mean change. The failed result is retained; it cannot be repaired by
changing the MNIST threshold or excluding the projection. A new dataset may receive a separately
frozen validation-only selection followed by an untouched test confirmation.
The apical compartment must contain ordinary non-teaching traffic generated by the model or task,
not only an additive nuisance chosen to match the predictor. Candidate traffic includes
back-propagating somatic activity, top-down contextual state, and task-irrelevant feedback shared
between neutral and teaching periods.
Required controls use the same forward architecture, initialization, data order, and update norm:
- raw apical signal;
- raw apical signal matched per sample to the innovation norm;
- somato-dendritic innovation;
- innovation with the predictor trained during task periods rather than neutral periods;
- no-traffic control, where raw and innovation should agree.
**Gate C1:** across at least two independently generated traffic families, innovation must improve
held-out accuracy by at least 2 percentage points over norm-matched raw feedback at a predeclared
traffic level, without harming the no-traffic setting by more than 0.5 points. The predictor must
remove measurable soma-predictable traffic while retaining causal teaching alignment.
The validation-only timescale screen froze `eta_P=0.05` and 200 neutral warmup steps by the highest
mean last-epoch accuracy across top-down traffic scales 0.2 and 0.5 (and the best worst-scale
endpoint). C1 must use this choice without further test-set tuning. Traffic projection seeds remain
separate from model/data seeds and must be recorded explicitly.
The frozen confirmation is MNIST d3/w256 for 15 epochs with the C3-selected low-query calibration
(`K=1`, every four minibatches). It crosses five model seeds with soma traffic at `rho=0.5` and
top-down traffic at `rho=0.2`, each using independently generated projection seeds 1234 and 5678.
Every traffic realization crosses raw, norm-matched raw, neutral-period innovation, and task-period
predictor controls; a no-traffic panel crosses the first three. This is 95 runs total. Test accuracy
is evaluated once after training, while final causal-alignment diagnostics use a fixed 512-example
training prefix. C1 passes only if every one of the four traffic realizations has a mean paired
innovation gain of at least 2 points over norm-matched raw, mean traffic R2 of at least 0.5, and
mean final hidden-layer teaching alignment of at least 0.1. The mean paired no-traffic innovation
change relative to both raw controls must be no worse than -0.5 points.
After the MNIST near-miss, one independent FashionMNIST recovery is frozen before any new result.
The validation screen uses model seed 0, a 5,000-example stratified split from training data
(`split_seed=2027`), d3/w256, 15 epochs, and otherwise the exact MNIST protocol. It crosses soma
traffic at `rho=0.5` and top-down traffic at `rho in {0.2, 0.5}`, each with projection seeds 1234
and 5678 and all four signal controls; a no-traffic panel has three controls (27 runs total).
Only final validation accuracy is observed; diagnostics use the fixed training prefix.
Soma traffic is eligible only if both projections gain at least 2 points over norm-matched raw,
retain accuracy within 3 points of no traffic, have traffic R2 at least 0.5, and alignment at least
0.1. A top-down level is eligible only if both its projections meet the same criteria. Among
eligible levels, select the one with the highest worst-projection gain, breaking an exact tie
toward smaller `rho`. No-traffic harm must remain within 0.5 points. If soma or no top-down level
is eligible, stop without evaluating FashionMNIST test. Otherwise freeze that level for a
five-model-seed, 95-run test confirmation under the original C1 per-realization gate.
**FashionMNIST recovery status: failed validation on 2026-07-22; test remains unevaluated.** Soma
traffic passed strongly in both projections. Neither top-down level was eligible: at `rho=0.2`,
projection 1234 gained only `+1.300` points and both residual endpoints were more than 3 points
below no traffic; at `rho=0.5`, both gains were large but both residual endpoints were more than
10 points below no traffic. This closes the predeclared recovery branch. C1 remains failed, and
the supported mechanism claim is narrowed to removal of soma-predictable ordinary traffic rather
than arbitrary contextual apical traffic.
### C2. SDIL scales when depth is useful
Flattened CIFAR is retained as a preservation result, not as evidence that added depth is used.
The primary controlled task must make BP accuracy increase with depth. Post-training block lesions
and residual-branch statistics must verify that the deep model actually uses its extra blocks.
**Gate C2:** BP must gain at least 5 points from the shallow to the deep configuration. SDIL must
recover at least 70% of that paired gain; the strongest competing local rule must recover no more
than 50%, or SDIL must beat it by at least 2 points at the deep endpoint. Removing the final third
of trained blocks must reduce SDIL performance, ruling out an effective shallow-network solution.
The C2 candidate is frozen from scratch pilots before formal validation: a two-level Telgarsky
tent-map binary task, one informative input, width 8, ReLU residual students, 80 epochs,
batch size 256, `eta=0.03`, and momentum 0.9. Each task has 10,000 generated training examples,
of which a stratified 2,000-example validation split (`split_seed=2027`) is held out; its untouched
test set has 5,000 independently generated examples. Shallow/deep are fixed to d1/d4. A BP shape
screen uses task seeds 0, 1, 2, student seed 0, and depths 1, 2, 3, 4, 6. Contextual depths cannot
replace d4 after results are observed. The BP screen passes only if mean d1-to-d4 gain is at least
5 points, every task seed gains, mean d4 final-third-lesion drop is at least 2 points, and every
task seed is hurt by the lesion. Only final validation metrics are observed.
**BP useful-depth status: passed on 2026-07-22.** Across task seeds 0--2, d1-to-d4 gains were
`+22.10`, `+42.85`, and `+33.90` points (mean `+32.95`). Bypassing the final third of d4 blocks
reduced accuracy by `40.85`, `35.05`, and `34.85` points (mean `36.92`). The contextual d6 curve
was less stable, so the preselected d4 endpoint remains fixed.
If BP passes, the local-rule validation panel freezes d1/d4 for BP, FA, DFA, and no-traffic SDIL
with C3's `K=1/every=4`, crossing the same three task seeds with student seeds 0--4. Its test panel
may run only if the original C2 gate holds on validation and the d4 SDIL lesion loses at least
2 points on average with at least 10 of 15 paired runs harmed. Test confirmation retrains on the
same 8,000-example training subset and evaluates each task seed's untouched 5,000 examples once.
**Local-rule validation status: failed on 2026-07-22; test remains unevaluated.** Across the 15
task/model pairs, BP gained `+29.697` points from d1 to d4 while SDIL gained only `+6.240`
(`21.0%` recovery). FA and DFA recovered `62.5%` and `59.5%`; SDIL's d4 endpoint was `5.560`
points below the stronger comparator. The SDIL lesion criterion itself passed (`+12.787`-point
mean drop; 11/15 harmed), so the failure is credit/optimization rather than unused depth. The
fixed linear output-error vectorizer `A_l c` is now the leading limitation: increasing perturbation
directions/frequency or simple gain normalization did not stabilize scratch validation. No C2
test panel may run unless a separately declared mechanism change passes validation.
The mechanism-recovery candidate is a bounded context-conditioned vectorizer,
`a_l = A_l c + G_l vec(c outer tanh(h_top))`, with both pathways calibrated by the same local
node-perturbation regression. The gate term is zero-initialized and vanishes when `c=0`, preserving
neutral-period innovation identification. An exploratory task-seed-0 validation budget compared
linear, per-cell soma-gated, and context-gated forms; `eta_A in {0.005, 0.01, 0.02}`; and K1/e2,
K1/e4, and K4/e4. Bounded context gating with `eta_A=0.01`, K1/e4 was selected by mean d4
validation accuracy. It recovered 70.2% of BP's task-0 depth gain but retained one bad basin.
Before any C2 test, an independent confirmation is frozen on new task seeds 3, 4, 5 and student
seeds 0--4, crossing BP, FA, DFA, and context-SDIL at d1/d4 (120 validation-only runs). It uses the
same original C2 thresholds, and additionally requires at least 10/15 context-SDIL depth gains to
be positive; its d4 lesion must lose at least 2 points on average and harm at least 10/15 runs.
Failure closes this recovery. Passing freezes a single test confirmation over task seeds 0--5,
student seeds 0--4, all four methods, and d1/d4, retrained on the unchanged 8,000-example subsets.
**Context-vectorizer recovery status: failed validation on 2026-07-22; test remains unevaluated.**
Context-SDIL recovered `53.0%` of BP's depth gain (`+15.143` versus `+28.553` points), below 70%,
and its d4 endpoint trailed FA by `1.643` points. One d4 trajectory had nonfinite final loss.
Positive depth gains and positive lesion effects each occurred in exactly 10/15 runs, so those
secondary conditions passed at the boundary. Context conditioning improves substantially over
linear SDIL's 21.0% recovery but does not make causal-feedback learning reliable enough. The
predeclared mechanism-recovery branch is closed and no C2 test panel will run.
The follow-up causal diagnosis replaces the learned vectorizer with direct layerwise antithetic
node-perturbation targets on every update. Task seed 0 and model seeds 0--2 were development-only:
an estimator/LR screen selected `eta=0.03`, and a predeclared smaller-K-within-one-point rule chose
K16 over K4. The frozen diagnosis then used task seeds 3--5, model seeds 0--4, d1/d4, the same
training-only validation splits, and no intermediate held-out metrics. It diagnoses an
amortization bottleneck only if it recovers at least 70% of BP's depth gain, has at least 10/15
positive paired gains, beats context-SDIL d4 by at least 2 points, has a mean lesion drop of at
least 2 points with at least 10/15 positive lesions, and has no nonfinite final loss.
**Direct causal diagnosis status: passed on 2026-07-22.** Direct NP recovered `106.6%` of BP's
depth gain (`+30.447` versus `+28.553` points), reached `96.150 ± 1.719%` at d4, improved in 15/15
depth pairs, and had positive lesions in 14/15 pairs. It beat context-SDIL d4 by `17.030` points;
all losses were finite. The result rules out the local eligibility rule and causal direction as
the primary C2 failure, but costs `68.4x` ordinary forward-equivalent work and is not a scalable
solution.
One calibration-quality recovery is now allowed on development task seed 0 only. It keeps the
bounded context vectorizer, `eta=0.03`, `eta_A=0.01`, and every-four-step calibration fixed, and
compares layerwise K4 against K16 at d1/d4 over model seeds 0--2. Select the higher mean d4
validation accuracy, choosing K4 if it lies within one point of K16; any nonfinite candidate is
ineligible. No result from task seeds 3--5 may enter this selection. A selected protocol must be
frozen before a new independent panel on task seeds 6--8, and that panel must satisfy the original
C2 gain/comparator/lesion conditions plus at least 10/15 positive depth gains and finite losses.
**Calibration-quality recovery status: failed development on 2026-07-22.** K4/e4 and K16/e4
reached only `65.633 ± 27.078%` and `65.567 ± 26.962%` at d4, with mean depth gains of `-0.583`
and `-0.350` points. Each had only 1/3 positive depth gains and lesions; K16 also produced one
nonfinite d4 loss. The frozen rule mechanically selects finite K4, but its negative depth gain
prohibits spending untouched task seeds 6--8 on confirmation. Increasing target quality while A
and W learn concurrently is therefore closed as a recovery branch.
One feedback-first timescale screen is allowed next on development task seed 0. It will use an
A-only calibration prefix with W and the output readout frozen, followed by the original
context-gated simultaneous K1/e4 joint phase. The prefix uses layerwise K4 targets and the same
`eta_A=0.01`; its length grid and selection rule must be committed before any run. This tests the
specific coupled-basin diagnosis without increasing vectorizer capacity or changing the task.
The frozen prefix-length grid is `{0, 100, 400}` A-only minibatches, crossed with d1/d4 and model
seeds 0--2. Warmup consumes an isolated RNG stream and restores the loader generator, so every
candidate's subsequent joint phase receives the same minibatch order and perturbation randomness.
A prefix is development-eligible only if all six final losses are finite, mean d4 validation is at
least 90%, at least 2/3 paired depth gains are positive, and the d4 lesion loses at least 2 points
on average with at least 2/3 positive lesions. Among eligible prefixes, select the highest mean d4
endpoint, choosing the shortest prefix within one point. If none is eligible, the feedback-first
branch closes without using task seeds 6--8.
**Feedback-first status: failed development on 2026-07-22.** The 100- and 400-step prefixes raised
mean d4 validation from `78.633 ± 24.866%` to `81.867 ± 8.779%` and `82.450 ± 14.226%`; both had
3/3 positive paired depth gains and lesions, versus 2/3 without warmup. Their mean gains were about
18 points and all losses were finite. Neither met the frozen 90% endpoint threshold, so no prefix
advances and task seeds 6--8 remain unused. Timescale separation reduces basin variance but does
not make plain SGD regression of the moving apical map accurate enough.
One final task-seed-0 regression-mechanism screen is allowed before stopping C2 development.
Normalized LMS divides each local A update by the squared norm of its complete instructional
feature `[c, c outer tanh(h_top)]`, addressing the angle/gain separation in C5 without enlarging
the vectorizer. It fixes the A-only prefix to 100 layerwise-K4 steps, keeps the joint phase at
simultaneous K1/e4, and compares only `eta_A in {0.01, 0.05}` over d1/d4 and model seeds 0--2.
Eligibility is identical to the feedback-first screen: all losses finite, d4 mean at least 90%, at
least 2/3 positive depth gains, and lesion mean at least 2 points with at least 2/3 positive. Among
eligible candidates select the highest d4 mean, preferring `eta_A=0.01` within one point. Failure
closes further tent-map tuning; success permits one frozen confirmation on task seeds 6--8.
**Normalized-regression status: failed development on 2026-07-22.** NLMS with `eta_A=0.01` and
`0.05` reached only `75.150 ± 23.011%` and `73.350 ± 23.008%` at d4, recovering mean depth gains
of `+4.433` and `+0.350` points. Each had only 2/3 positive gains and lesions. Both are ineligible;
the tent-map recovery program is closed, no task seed 6--8 result was generated, and C2 remains a
negative result for the amortized method. The passed direct-NP diagnosis and all failed recovery
mechanisms remain part of the evidence rather than being hidden by another search branch.
### C3. Causal calibration is genuinely amortized
**Status: passed on 2026-07-22.** The frozen five-seed confirmation retained 112.9% of the K16/e4
gain over DFA while using 16x fewer perturbation loss queries, 11x less calibration
forward-equivalent work, and 5.3x less total training forward-equivalent work. Peak allocated
memory fell from 803.5 MiB to 739.8 MiB. These numbers apply to the audited CIFAR-10 d20/w64
protocol; extrapolation to standard CNNs remains an oral-A question.
Report wall time, forward-equivalent work, number of scalar loss queries, peak device memory, and
the perturbation batch expansion. Sweep perturbation frequency and number of directions under both
fixed epoch and fixed query/FLOP budgets.
**Gate C3:** a protocol using at least 10x fewer perturbation loss queries than the current
16-directions-every-4-steps setting must preserve at least 90% of its gain over DFA. Pareto claims
must hold in a hardware-independent cost coordinate as well as GTX-1080 wall time.
The validation screen is frozen to CIFAR-10 d20/w64, five epochs, seed 0, and compares the current
`K=16/every=4` reference with `K/every` pairs `1/4`, `2/8`, `4/16`, `1/8`, `1/16`, and `2/32`, plus
DFA. Selection uses final validation accuracy and logical batch loss queries; the test split is not
evaluated. `K=1/every=4` was selected: it used 16x fewer batch loss queries and retained 177.5% of
the reference gain over DFA. This selection is now frozen; the protocol receives a five-seed test
confirmation with no further tuning before C3 can be declared passed.
### C4. The comparison set contains the closest alternatives
The main comparison must include exact BP, FA, DFA, direct/unamortized node perturbation, learned
direct node-perturbation feedback (Lansdell et al. 2020), Dual Prop, and BurstCCN. The Lansdell
method is the exact no-traffic/P=0 backbone of the current implementation, not merely a related
baseline. BurstCCN is already demonstrated on CIFAR-10 and ImageNet and already models the Harnett
BCI signatures, so it is the strongest dendritic/scaling comparator. EP remains an informative
relaxation-based baseline. Forward-Forward and PEPITA are appendix context unless they become
competitive under the frozen protocol.
Each method receives a documented validation budget. Native-protocol and exact-architecture
comparisons are labelled separately; neither substitutes for a matched compute/query comparison.
**Status: passed on 2026-07-22.** In-repository BP/FA/DFA, direct NP, learned NP,
PEPITA, Forward-Forward, and canonical EP are complete. Paper-faithful BurstCCN
seed 0 reaches `80.07%` final / `80.10%` validation-selected / `80.25%`
test-selected accuracy in `15451.3 s`, versus published `82.97 +/- 0.21%`.
Author-code Dual Prop VGG16 seed 1988 reaches `92.46%` from the best-validation
checkpoint in `23119.8 s`, versus published `92.41 +/- 0.07%`, with one test
evaluation. Both records pass artifact, source, environment, dataset,
completeness, finiteness, selection, and cost-definition checks in
`finalize_accept.sh`. These one-seed method-native reproductions remain
separate from matched-compute comparisons and from published multi-seed
uncertainty.
### C5. Theory predicts the observed regimes
The minimum theory package contains:
1. bias and variance of simultaneous multi-layer perturbation, including cross-layer interference;
2. a smooth-loss descent bound separating angular alignment, gain calibration, and curvature;
3. an innovation/SNR result for subtracting soma-predictable apical traffic;
4. an explicit complexity table covering learning phases, transport, loss queries, FLOPs, and
memory.
Numerical simulations must test the predicted scaling with depth, width, directions, perturbation
scale, and predictor timescale.
**Status: passed on 2026-07-22.** `THEORY.md` gives the exact simultaneous-Rademacher MSE with its
cross-layer term, the Gaussian layerwise comparison, an `O(sigma^2)` antithetic bias bound, the
smooth-loss angle/gain/curvature descent condition, the neutral innovation/SNR theorem and
task-fit absorption result, and a phase/query/work/memory table. `experiments/verify_theory.py`
checks depth, width, K, sigma, and predictor timescale deterministically; all assertions pass.
## Protocol discipline
- Hyperparameters are selected on a validation split created only from the training set.
- The test set is evaluated for frozen candidate protocols, never used to choose schedules.
- Main results use five model seeds. Synthetic task claims additionally vary the task/teacher seed.
- Every completed finite trajectory is retained; crashes and excluded runs are reported with a
reason fixed before inspecting accuracy.
- Pilot results remain versioned but cannot be pooled with frozen runs.
- Every result records source revision, dirty state, protocol identifier, data split hash, logical
loss queries, forward-equivalent work, peak memory, and wall time.
## Oral bar B: biological bridge
**Status: development gate failed and confirmation is closed.** The complete
36-run preregistered screen is reported in `ORAL_B.md` and `RESULTS.md`. Local
learning, residual decorrelation, outcome decoding, and the plasticity lesion
were positive, but all causal-role sign-inversion indices were negative and
acute online-control lesions had little effect. Seeds reserved for confirmation
were not touched.
After C1–C5 pass, test whether the model reproduces the qualitative Harnett signatures:
- dendritic activity contains information absent after conditioning on somatic activity;
- the residual decodes outcome/reward-related events;
- residual sign follows the neuron's causal contribution to the objective;
- residual predicts subsequent activity change or desired velocity;
- selectively disrupting the residual impairs learning more than matched nonspecific disruption.
The paper needs at least one prediction not used to construct SDIL. A preferred route is to request
the original data/analysis from the authors and test the prediction out of sample; otherwise the
claim is explicitly computational rather than a fit to cortical data.
## Oral bar A: standard deep architectures
**Status: A3 failed; A4 was not opened and all confirmation test seeds remain untouched.**
`ORAL_A.md` froze a validation-only ResNet-20 development funnel. A1 passed with
`91.62%` BP validation accuracy. A2b selected channel-gated SDIL at `41.98%`,
ahead of tuned DFA at `37.16%`, on the 10k/20-epoch screen. In full A3, DFA
ended finite at `33.06%`; SDIL became nonfinite at epoch 89 and ended at
`10.00%`, failing alignment, accuracy, and finiteness gates. The MAC gate alone
passed. Per the stop rule, the five-seed ResNet-20/32/56 test panel was never
run. The exact BatchNorm/local-gradient mechanics remain verified, but the
current training recipe does not establish standard-network scaling.
A separate dynamic-innovation recovery is now frozen in
`ORAL_A_RECOVERY.md` before any new endpoint. It does not reopen failed A3/A4:
only complete D4 and oral-B R2 passes can unlock a 60-cell
BP/DFA/clean-KP/dynamic ResNet-20/32/56 panel. D4 supplies the ten depth-20
KP/dynamic cells verbatim; exactly 50 new cells are permitted. A full pass is
the sole 8-to-9 score gate and must show a positive paired depth benefit, not
merely survival on another depth-flat task.
Post-failure diagnosis is recorded separately in
`results/oral_a_failure_diagnosis.json`: prediction--target cosine remains
below `8.51e-5` in magnitude, calibration MSE is indistinguishable from target
power, and target power grows `2.13e9x` before nonfiniteness. The next
development branch must therefore reduce causal-estimator variance in the
representable channel-gated subspace; changing only the v1 threshold is not an
allowed response.
**Post-failure v2 status: causal-capture gate failed.** The structured
representable-subspace estimator raised matched frozen-forward early/all-layer alignment from
`0.001105/0.011664` to `0.007209/0.052740`. It passed the all-layer threshold
but missed both the preregistered `0.01` early-layer threshold and the `0.01`
absolute early-layer advantage. Per `ORAL_A_V2.md`, the full 200-epoch v2 run
was not launched; test and confirmation seeds remain untouched. This localizes
the remaining bottleneck to early-layer credit rather than merely spatial
projection variance.
The subsequent no-training representation oracle further decomposes that
bottleneck. The existing spatial fields reach `0.05494` early alignment with
unconstrained per-example coefficients; the actual output-error-conditioned
family reaches `0.02397` on a disjoint 32-example batch, while learned v2
reaches `0.00721`. Naive local-average/channel-context fields fall slightly to
`0.02246`. The next algorithmic branch should therefore address both causal
sample efficiency and feedback context; simply lowering `eta_A`, adding the
tested fields, or reopening V2-2 is not justified.
**Vectorizer-space V3 status: causal-capture gate failed.** Directly estimating
the A/G matrix target passed exact mechanics and reduced matched synthetic
one-query MSE to `0.03381x`, but its selected real early alignment was
`0.007139` versus the matched V2 reference's `0.007209`. All-layer alignment
rose from `0.052740` to `0.062579`; three early/oracle checks failed and only
the all-layer check passed. V3 full training was not launched, and test access
remains sealed. This closes learning-rate tuning of the same output-error-only
channel-gated family; a further branch must make a substantive feedback-context
or cross-layer-noise change.
The fixed post-failure hierarchical oracle identifies the substantive next
change. Ordinary downstream activation context moves early alignment only from
`0.031123` to `0.032069`, while held-out local maps from the actual residual-DAG
child error fields reach `0.440525` (1x1), `0.819917` (3x3), and `0.999767`
(3x3 plus the locally available child ReLU gate). This is not task evidence
and uses exact child gradients as oracle inputs. It establishes that future
development should learn spatial hierarchical feedback without weight copying,
with random/fixed hierarchical FA and BurstCCN as mandatory baselines.
**Fixed hierarchical-FA baseline status: short gate failed.** The audited
ResNet-20 screen reaches `39.64%`, `41.58%`, and `43.52%` at hidden rates 0.01,
0.03, and 0.1. The selected run improves on matched DFA by 6.36 points and has
early alignment `0.040406`, but misses its frozen 50% full-run threshold.
HFA-S2 is therefore closed. This validates hierarchical feedback as a stronger
baseline while ruling out fixed random hierarchy as the complete repair.
The next mechanics branch perturbs every hierarchical feedback-parameter
subspace in one antithetic pair and subtracts locally computed prediction
moments. Its causal JVP and exact symmetric-limit delta rule pass, but it has no
task claim until a frozen-forward alignment gate is preregistered and audited.
**Hierarchical task-scalar V4 status: causal-capture gate failed.** Stable
etaA 0.1 changes early/all-layer alignment from `-0.000263/0.011279` to only
`-0.000180/0.017336`. EtaA 1 reaches early `0.000348` but produces a 58.48x
feedback norm ratio; etaA 10 becomes nonfinite. All substantive gates fail,
so no accuracy run opens. This closes simultaneous global scalar calibration
of the full 267,904-parameter feedback path at 400 events. The next branch must
use richer local child-response information and include learned FA/weight
mirroring as an explicit inherited baseline before adding innovation.
**Normalized response-mirror WM-1 status: passed.** Twenty batch-1 local
observations at selected etaM 0.1 raise early/all alignment from
`-0.000263/0.011279` to `0.446141/0.542159`, with feedback/forward cosine
`0.915480`, norm ratios `[0.889,1.157]`, zero task-loss queries, and only
`1.6305e9` MACs. All five gates pass. The bounded two-rate WM-2 accuracy screen
is opened. This is inherited baseline performance; the paper score changes
only if innovation later becomes load-bearing on top of the scalable mirror.
**Normalized response-mirror WM-2 status: short accuracy gate failed
narrowly.** Selected hidden LR 0.1 reaches `64.04%`, versus `43.52%` fixed HFA
and `74.94%` BP, with early/all alignment `0.9393/0.9162`, zero task-loss
queries, and `0.9968x` BP MACs. It misses the 65% threshold by 0.96 points and
the within-10-point BP gate by 0.90 points, so WM-3 is closed. The next branch
cannot tune this mirror recipe; it must make a substantive rule or information
change and must not treat high alignment alone as task equivalence.
**Residual response-mirror RRM-1 status: capture gate passed.** Selected
etaM 0.1 reaches early/all alignment `0.665070/0.730253` and feedback/forward
cosine `0.953680` after 20 local observations, with norm ratios
`[0.884,1.018]`, zero task-loss queries, and `2.4458e9` MACs. All frozen gates
pass. EtaM 0.3 is visibly unstable (maximum norm ratio 30.13), so the selected
rate is not widened or retuned. The frozen two-rate RRM-2 short accuracy screen
is opened. This is inherited baseline progress and cannot change the paper
score until a load-bearing innovation experiment uses the substrate.
**Residual response-mirror RRM-2 status: short accuracy gate passed
narrowly.** Selected hidden LR 0.1 reaches `65.08%`, 9.86 points below matched
BP at `74.94%`, with early/all alignment `0.999128/0.999377`, feedback/forward
cosine `0.999902`, zero task-loss queries, and `0.9969x` BP MACs. It clears the
two accuracy margins by only 0.08 and 0.14 points; the LR-0.03 run reaches
62.32%. Every frozen check nevertheless passes, so exactly one RRM-3 full
validation run opens at the selected settings. The result remains baseline
evidence and the reviewer score stays 5/10.
**Residual response-mirror RRM-3 status: failed decisively.** The sole frozen
full record ends at 10.00% accuracy and NaN validation loss. Training loss
peaks at `1.5808e16` in epoch 92 and remains `8.9503e13` at epoch 200. Yet the
decayed-LR endpoint reports early/all-layer alignment `0.8774/0.8487`, Q/W
cosine `0.999998`, and norm ratios `[1.0002,1.0024]`. This closes RRM without
cadence/rate rescue and demonstrates why endpoint cosine cannot substitute for
a training-trajectory audit. No innovation or test panel opens from RRM.
**Modified Kolen--Pollack baseline protocol: mechanics and both task gates
passed.** The separate local reciprocal correlations, post-observation W/Q
independence, and exact symmetric-limit momentum updates all pass at zero
numerical error. `KP_BASELINE.md` freezes one 20-epoch full-development screen
with no LR/decay grid and requires both validation accuracy and training-period
feedback tracking. No KP task endpoint had been generated when this protocol
was committed. KP is inherited Akrout et al. machinery and cannot raise the
paper score; only a later load-bearing innovation ablation can do that.
**Modified Kolen--Pollack KP-1 status: passed.** The sole frozen record reaches
`82.66%` validation accuracy versus matched BP's epoch-20 `81.02%`, with early
alignment `0.885583`, final feedback/forward cosine `0.901034`, and epoch-11--20
mean cosine `0.854594`. It is finite, uses zero task-loss queries, and costs
`1.3261x` BP after charging reciprocal correlations. KP-2 is opened at exactly
the frozen settings. This strongly repairs the feedback-substrate engineering
problem but remains inherited Akrout et al. evidence, so the paper score stays
5/10.
**Modified Kolen--Pollack KP-2 status: passed.** The sole frozen 200-epoch
record reaches `91.26%` validation accuracy versus matched BP's `91.62%`.
Early alignment is `0.999397`; final and epoch-151--200 mean feedback/forward
cosines are both `0.999663`. The complete trajectory is finite, feedback uses
zero task-loss queries, and cost remains `1.3261x` BP. This opens MT-1 without
raising the reviewer score because the result is entirely inherited KP
evidence.
**KP mixed-traffic status: MT-0 passed; MT-1 failed.** `MIXED_TRAFFIC.md` fixes
a four-to-one, initialization-calibrated
soma-predictable traffic intervention and crosses raw apical activity,
per-example norm-matched raw activity, and neutral-period innovation on the
same reciprocal KP substrate. The zero-traffic limit, exact-predictor limit,
gain calibration, norm/direction control, reciprocal local correlations, and
parameter independence pass at zero or float64 machine error. A realistic
20-step neutral warmup leaves `0.1360` of traffic RMS versus the frozen `0.25`
ceiling. Elementwise traffic/predictor work is reported separately from affine
MACs, and every control pays the predictor schedule. The synthetic batch-128
matched path is finite on a GTX 1080 and peaks at 0.857 GB allocated after
reset; no task endpoint was used for this hardware check.
The complete MT-1 panel then becomes nonfinite in epoch 1 under all three
signals and ends at 10.00% validation accuracy. Calibration error remains only
`4.77e-7`, predictor warmup reaches residual ratio `0.139942`, cost is
`1.3271x` BP, and feedback uses zero task-loss queries, so those mechanical
checks pass while all performance, finite, alignment, and tracking checks
fail. Per the frozen stop rule, MT-2 and MT-3 remain untouched, no weaker
traffic ratio is allowed, and this standard-ResNet accept path closes at 5/10.
The training-only post-failure audit localizes the common instability to stem
`W[0]` and its momentum at steps 23/69/74 for raw/matched/innovation,
respectively, with no validation or test evaluation. This narrows any future
independent method to operator-level gain control or dissipativity; more
endpoint tuning of the same residual-RMS gate is not justified.
The follow-up S0 one-sided stability branch also closes before held-out access.
Its frozen four-margin training-prefix grid has no eligible candidate: small
margins still overflow BN state and grow weights to `1e17`, while larger
margins retain finite tensors but violate loss, signal-ratio, weight, and
momentum envelopes by orders of magnitude. No validation protocol opens. This
rules out fixed sign-only predictor bias as the missing gain-control mechanism.
**Dynamic neutral projection D1: passed.** `DYNAMIC_INNOVATION.md` freezes a
single fast local controller rather than another margin grid. Every task batch
provides a paired instruction-off soma/apical observation; the fast affine fit
removes the current neutral residual coupling without reading task instruction
or changing the slow predictor. All 352 training-only steps and states remain
finite, maximum loss is `2.9281`, and the controller reduces a growing
`0.005833` neutral/traffic ratio to at most `3.03e-8`. No held-out evaluation
occurs, so this mechanics pass opens D2 without changing the score.
**Dynamic neutral projection D2: passed.** The sole frozen 20-epoch record
reaches `83.58%` validation accuracy, above clean KP's `82.66%` and 73.58
points above both failed MT-1 controls. Early alignment is `0.883830`, final
feedback cosine is `0.904056`, all projection certificates pass, cost is
`1.3261x` BP MACs plus explicitly reported elementwise work, and test remains
untouched. Because this is one short seed, the score stays 5/10. The
predeclared 200-epoch D3 validation endpoint is now open; only a complete D3
pass can make the 5-to-6 score update eligible and open independent test
confirmation. Before observing D3, `DYNAMIC_INNOVATION.md` and executable
scripts freeze D4 as a five-seed paired clean-KP/dynamic test panel. D4 is
hard-gated on a complete D3 pass, evaluates test exactly once per run, and can
raise 6-to-7 only through its predeclared accuracy, paired noninferiority,
alignment, leakage, query, and cost checks.
**Dynamic neutral projection D3: passed.** The sole frozen 200-epoch record
passes all 19 gates at source revision `d945c42`: `91.18%` validation versus
BP's `91.62%` and clean KP's `91.26%`, `0.999353` early alignment, `0.999663`
late feedback cosine, zero task-loss queries, and `1.3261x` BP MACs. Exactly
one validation and zero test evaluations occur. Per the frozen rule, the
strict reviewer score rises from 5 to 6 and the already committed D4 paired
five-seed test panel is now open.
**Dynamic neutral projection D4: passed.** All ten untouched test records at
source revision `0008f2c` pass the frozen audit. Across seeds 10--14, dynamic
innovation reaches `91.584%` mean test accuracy versus clean KP's `91.388%`;
the mean paired clean-minus-dynamic deficit is `-0.196` points and its
one-sided 95% upper bound is `0.131` points. Every dynamic seed reaches at
least `91.51%`, mean early alignment is `0.999687`, and all finite-trajectory,
feedback-tracking, projection, leakage, query, cost, memory, hardware,
provenance, split, and exactly-once test checks pass. Per the frozen rule the
strict reviewer score rises from 6 to 7, establishing the accept bar and
opening oral-B recovery R1. Oral-A remains sealed until the separately frozen
oral-B R2 confirmation passes.
R1 subsequently passes and selects `eta=0.1`, but the complete untouched R2
gate fails the outcome-vectorization and longitudinal signatures. Oral-A
therefore remains sealed permanently under this frozen sequence.
Prepare convolutional local-update primitives and ResNet-20/32/56 protocols early. Queue frozen
runs opportunistically on authorized idle GPUs. Because BurstCCN already reports CIFAR-10 and
ImageNet scaling, dataset scale alone is not novel. The oral-level target is a memorable joint
result: near-BP accuracy in a standard architecture, a win over or materially simpler/lower-cost
regime than BurstCCN, and a nondominated accuracy–hardware-independent-cost point. Scale alone is
also insufficient if C1 does not show that somato-dendritic innovation is load-bearing.
## Stop conditions
- Do not extend flattened-CIFAR MLP depth merely to obtain a larger depth number.
- Do not add MNIST EP seeds unless a protocol bug invalidates the completed five-seed panel.
- Do not promote a wall-time frontier that disappears under loss-query or FLOP accounting.
- Do not describe the algorithm as a cortical implementation if simultaneous perturbation or
supervision assumptions remain biologically unsupported.
|