# SDIL ICLR 2027 evidence roadmap This document predeclares the evidence gates used to decide which claims survive into a paper. It is intentionally stricter than a list of experiments: a failed gate narrows the claim rather than triggering selective seed removal or post-hoc protocol changes. Paper-level progress is tracked independently in `REVIEW_SCORECARD.md`. Its ICLR-style score may change only after a frozen result is complete and audited; partial jobs and development pilots do not count. ## Order of work 1. **Accept bar:** establish a load-bearing innovation mechanism, useful-depth scaling, fair cost, theory, and the closest baselines. 2. **Oral bar B:** connect the mechanism to the signatures and new predictions motivated by Harnett et al.'s somato-dendritic residuals. 3. **Oral bar A:** demonstrate the same mechanism in standard CNN/ResNet families. The runner is developed early so long jobs can use otherwise idle, explicitly authorized GPUs without delaying stages 1–2. The preregistered oral-B task, structural screen, confirmation seeds, signature thresholds, and phase-specific lesion are specified in `ORAL_B.md`. They were frozen while C4 author baselines were still running and before any continuous-BCI result was generated. Oral-B confirmation was sequenced after C4 and the D4 accept gate; the realized original and recovery branches are reported below. **Oral-B development status: failed.** The complete frozen 12-candidate by three-task-seed screen found no eligible mechanism. High-learning-rate variants learned the BCI reliably, and several innovation/decoder signatures were strong, but every run had the wrong P+/P- error-change sign-inversion and error magnitude remained more predictive than temporal error change. Per the preregistered stop rule, environment seeds 10--15 and all confirmation model seeds remain untouched. This falsifies the current model's claimed emergence of the Harnett derivative signature; it cannot be repaired by reporting only its successful learning or decoder results. **Oral-B recovery R0/R1 status: mechanics and development passed.** Reinspection of Fig. 5 and Extended Data Fig. 13 identifies a structural target conflict: the failed vectorizer regressed `e_t s_i`, so its role contrast encoded current error magnitude rather than the sign of error change. `ORAL_B_RECOVERY.md` factorizes a new candidate into a perturbation-learned causal role and the within-episode performance innovation `|e_(t-1)|-|e_t|`. Deterministic checks recover the role at cosine `0.9963` and change the controlled sign index from negative to `+0.100`. After D4 passes, the frozen two-rate, three-task-seed R1 screen selects `eta=0.1`; all 30 seed-level checks pass and worst-task final success is `98.05%`. R1 is development evidence and leaves the score at 7. **Oral-B recovery R2 status: failed; branch closed.** The untouched 6-task by 5-model confirmation retains `99.53%` mean final success, `90.45` points of learning gain, strong fixed-vectorizer and plasticity-lesion margins, mean role cosine `0.9395`, positive sign inversion in all 30 records, and a `0.6389` velocity-over-error advantage. Residual decorrelation and decoder-distance prediction also pass. The joint biological gate fails because surrounding-state accuracy is `54.51%` versus the 55% mean threshold, residual outcome accuracy is only `47.33%`, residuals trail soma outcome decoding by `4.32` points, and longitudinal prediction is `-0.013`. The score remains 7, oral-A stays sealed, and no threshold repair or replacement confirmation is permitted. Because `kappa=0`, neither R1 nor R2 could support online control. **Oral-B-v2 development path: two failures retained.** `ORAL_B_V2.md` adds an explicit terminal reward/timeout phase, a local TD critic, and eligibility traces without changing the old R2. Its complete 24-record development grid fails from a cold start: the dense performance innovation was scaled to one-quarter of the validated rule and no candidate learns. A separately frozen unit-scale recovery restores 100% task performance and passes every mechanism gate on task seeds 23 and 24; task seed 25 passes 17/18 gates but succeeds on 98.96% of a fixed absolute target ladder. That class-balance failure is also retained and does not open confirmation. **Oral-B-v2 calibrated R1/R2 status: passed.** The final branch changes no learning parameter. Five target levels are fixed by cursor-maximum quantiles on 512 separate outcome-free calibration trials, then evaluated on independent trajectories. Fresh development seeds 26--28 pass all 18/18 checks. The untouched 6-task by 5-model confirmation then passes every learning, innovation, network-prediction, outcome, lesion, and task-cluster confidence gate: final success is `100%`, learning gain `98.80` points, fixed-role gap `99.97` points, role cosine `0.9761`, residual--soma correlation `0.0631`, surrounding accuracy `54.37%` (lower bound `54.26%`), velocity advantage `0.6384`, terminal outcome accuracy `99.83%` (lower bound `99.68%`), acute outcome-lesion separation drop `0.4000` (lower bound `0.3886`), critic expectedness `0.3187` (lower bound `0.2814`), and 30/30 positive signs. The formal milestone rises from 7 to 8. The old oral-A gate remains closed; only a new independently frozen oral-A-v2 protocol may now run. ## Frozen accept claims and gates ### C1. Innovation is necessary under naturally mixed apical traffic **Status: failed the frozen MNIST confirmation on 2026-07-22.** Three of four traffic realizations passed, but top-down projection seed 1234 produced a `+1.520`-point innovation gain over norm-matched raw feedback, below the predeclared `+2`-point threshold. The other top-down projection gained `+2.956` points and the two soma-traffic projections gained `+2.860` and `+3.180` points. All four passed the traffic-R2 and alignment requirements, and the no-traffic control had exactly zero mean change. The failed result is retained; it cannot be repaired by changing the MNIST threshold or excluding the projection. A new dataset may receive a separately frozen validation-only selection followed by an untouched test confirmation. The apical compartment must contain ordinary non-teaching traffic generated by the model or task, not only an additive nuisance chosen to match the predictor. Candidate traffic includes back-propagating somatic activity, top-down contextual state, and task-irrelevant feedback shared between neutral and teaching periods. Required controls use the same forward architecture, initialization, data order, and update norm: - raw apical signal; - raw apical signal matched per sample to the innovation norm; - somato-dendritic innovation; - innovation with the predictor trained during task periods rather than neutral periods; - no-traffic control, where raw and innovation should agree. **Gate C1:** across at least two independently generated traffic families, innovation must improve held-out accuracy by at least 2 percentage points over norm-matched raw feedback at a predeclared traffic level, without harming the no-traffic setting by more than 0.5 points. The predictor must remove measurable soma-predictable traffic while retaining causal teaching alignment. The validation-only timescale screen froze `eta_P=0.05` and 200 neutral warmup steps by the highest mean last-epoch accuracy across top-down traffic scales 0.2 and 0.5 (and the best worst-scale endpoint). C1 must use this choice without further test-set tuning. Traffic projection seeds remain separate from model/data seeds and must be recorded explicitly. The frozen confirmation is MNIST d3/w256 for 15 epochs with the C3-selected low-query calibration (`K=1`, every four minibatches). It crosses five model seeds with soma traffic at `rho=0.5` and top-down traffic at `rho=0.2`, each using independently generated projection seeds 1234 and 5678. Every traffic realization crosses raw, norm-matched raw, neutral-period innovation, and task-period predictor controls; a no-traffic panel crosses the first three. This is 95 runs total. Test accuracy is evaluated once after training, while final causal-alignment diagnostics use a fixed 512-example training prefix. C1 passes only if every one of the four traffic realizations has a mean paired innovation gain of at least 2 points over norm-matched raw, mean traffic R2 of at least 0.5, and mean final hidden-layer teaching alignment of at least 0.1. The mean paired no-traffic innovation change relative to both raw controls must be no worse than -0.5 points. After the MNIST near-miss, one independent FashionMNIST recovery is frozen before any new result. The validation screen uses model seed 0, a 5,000-example stratified split from training data (`split_seed=2027`), d3/w256, 15 epochs, and otherwise the exact MNIST protocol. It crosses soma traffic at `rho=0.5` and top-down traffic at `rho in {0.2, 0.5}`, each with projection seeds 1234 and 5678 and all four signal controls; a no-traffic panel has three controls (27 runs total). Only final validation accuracy is observed; diagnostics use the fixed training prefix. Soma traffic is eligible only if both projections gain at least 2 points over norm-matched raw, retain accuracy within 3 points of no traffic, have traffic R2 at least 0.5, and alignment at least 0.1. A top-down level is eligible only if both its projections meet the same criteria. Among eligible levels, select the one with the highest worst-projection gain, breaking an exact tie toward smaller `rho`. No-traffic harm must remain within 0.5 points. If soma or no top-down level is eligible, stop without evaluating FashionMNIST test. Otherwise freeze that level for a five-model-seed, 95-run test confirmation under the original C1 per-realization gate. **FashionMNIST recovery status: failed validation on 2026-07-22; test remains unevaluated.** Soma traffic passed strongly in both projections. Neither top-down level was eligible: at `rho=0.2`, projection 1234 gained only `+1.300` points and both residual endpoints were more than 3 points below no traffic; at `rho=0.5`, both gains were large but both residual endpoints were more than 10 points below no traffic. This closes the predeclared recovery branch. C1 remains failed, and the supported mechanism claim is narrowed to removal of soma-predictable ordinary traffic rather than arbitrary contextual apical traffic. ### C2. SDIL scales when depth is useful Flattened CIFAR is retained as a preservation result, not as evidence that added depth is used. The primary controlled task must make BP accuracy increase with depth. Post-training block lesions and residual-branch statistics must verify that the deep model actually uses its extra blocks. **Gate C2:** BP must gain at least 5 points from the shallow to the deep configuration. SDIL must recover at least 70% of that paired gain; the strongest competing local rule must recover no more than 50%, or SDIL must beat it by at least 2 points at the deep endpoint. Removing the final third of trained blocks must reduce SDIL performance, ruling out an effective shallow-network solution. The C2 candidate is frozen from scratch pilots before formal validation: a two-level Telgarsky tent-map binary task, one informative input, width 8, ReLU residual students, 80 epochs, batch size 256, `eta=0.03`, and momentum 0.9. Each task has 10,000 generated training examples, of which a stratified 2,000-example validation split (`split_seed=2027`) is held out; its untouched test set has 5,000 independently generated examples. Shallow/deep are fixed to d1/d4. A BP shape screen uses task seeds 0, 1, 2, student seed 0, and depths 1, 2, 3, 4, 6. Contextual depths cannot replace d4 after results are observed. The BP screen passes only if mean d1-to-d4 gain is at least 5 points, every task seed gains, mean d4 final-third-lesion drop is at least 2 points, and every task seed is hurt by the lesion. Only final validation metrics are observed. **BP useful-depth status: passed on 2026-07-22.** Across task seeds 0--2, d1-to-d4 gains were `+22.10`, `+42.85`, and `+33.90` points (mean `+32.95`). Bypassing the final third of d4 blocks reduced accuracy by `40.85`, `35.05`, and `34.85` points (mean `36.92`). The contextual d6 curve was less stable, so the preselected d4 endpoint remains fixed. If BP passes, the local-rule validation panel freezes d1/d4 for BP, FA, DFA, and no-traffic SDIL with C3's `K=1/every=4`, crossing the same three task seeds with student seeds 0--4. Its test panel may run only if the original C2 gate holds on validation and the d4 SDIL lesion loses at least 2 points on average with at least 10 of 15 paired runs harmed. Test confirmation retrains on the same 8,000-example training subset and evaluates each task seed's untouched 5,000 examples once. **Local-rule validation status: failed on 2026-07-22; test remains unevaluated.** Across the 15 task/model pairs, BP gained `+29.697` points from d1 to d4 while SDIL gained only `+6.240` (`21.0%` recovery). FA and DFA recovered `62.5%` and `59.5%`; SDIL's d4 endpoint was `5.560` points below the stronger comparator. The SDIL lesion criterion itself passed (`+12.787`-point mean drop; 11/15 harmed), so the failure is credit/optimization rather than unused depth. The fixed linear output-error vectorizer `A_l c` is now the leading limitation: increasing perturbation directions/frequency or simple gain normalization did not stabilize scratch validation. No C2 test panel may run unless a separately declared mechanism change passes validation. The mechanism-recovery candidate is a bounded context-conditioned vectorizer, `a_l = A_l c + G_l vec(c outer tanh(h_top))`, with both pathways calibrated by the same local node-perturbation regression. The gate term is zero-initialized and vanishes when `c=0`, preserving neutral-period innovation identification. An exploratory task-seed-0 validation budget compared linear, per-cell soma-gated, and context-gated forms; `eta_A in {0.005, 0.01, 0.02}`; and K1/e2, K1/e4, and K4/e4. Bounded context gating with `eta_A=0.01`, K1/e4 was selected by mean d4 validation accuracy. It recovered 70.2% of BP's task-0 depth gain but retained one bad basin. Before any C2 test, an independent confirmation is frozen on new task seeds 3, 4, 5 and student seeds 0--4, crossing BP, FA, DFA, and context-SDIL at d1/d4 (120 validation-only runs). It uses the same original C2 thresholds, and additionally requires at least 10/15 context-SDIL depth gains to be positive; its d4 lesion must lose at least 2 points on average and harm at least 10/15 runs. Failure closes this recovery. Passing freezes a single test confirmation over task seeds 0--5, student seeds 0--4, all four methods, and d1/d4, retrained on the unchanged 8,000-example subsets. **Context-vectorizer recovery status: failed validation on 2026-07-22; test remains unevaluated.** Context-SDIL recovered `53.0%` of BP's depth gain (`+15.143` versus `+28.553` points), below 70%, and its d4 endpoint trailed FA by `1.643` points. One d4 trajectory had nonfinite final loss. Positive depth gains and positive lesion effects each occurred in exactly 10/15 runs, so those secondary conditions passed at the boundary. Context conditioning improves substantially over linear SDIL's 21.0% recovery but does not make causal-feedback learning reliable enough. The predeclared mechanism-recovery branch is closed and no C2 test panel will run. The follow-up causal diagnosis replaces the learned vectorizer with direct layerwise antithetic node-perturbation targets on every update. Task seed 0 and model seeds 0--2 were development-only: an estimator/LR screen selected `eta=0.03`, and a predeclared smaller-K-within-one-point rule chose K16 over K4. The frozen diagnosis then used task seeds 3--5, model seeds 0--4, d1/d4, the same training-only validation splits, and no intermediate held-out metrics. It diagnoses an amortization bottleneck only if it recovers at least 70% of BP's depth gain, has at least 10/15 positive paired gains, beats context-SDIL d4 by at least 2 points, has a mean lesion drop of at least 2 points with at least 10/15 positive lesions, and has no nonfinite final loss. **Direct causal diagnosis status: passed on 2026-07-22.** Direct NP recovered `106.6%` of BP's depth gain (`+30.447` versus `+28.553` points), reached `96.150 ± 1.719%` at d4, improved in 15/15 depth pairs, and had positive lesions in 14/15 pairs. It beat context-SDIL d4 by `17.030` points; all losses were finite. The result rules out the local eligibility rule and causal direction as the primary C2 failure, but costs `68.4x` ordinary forward-equivalent work and is not a scalable solution. One calibration-quality recovery is now allowed on development task seed 0 only. It keeps the bounded context vectorizer, `eta=0.03`, `eta_A=0.01`, and every-four-step calibration fixed, and compares layerwise K4 against K16 at d1/d4 over model seeds 0--2. Select the higher mean d4 validation accuracy, choosing K4 if it lies within one point of K16; any nonfinite candidate is ineligible. No result from task seeds 3--5 may enter this selection. A selected protocol must be frozen before a new independent panel on task seeds 6--8, and that panel must satisfy the original C2 gain/comparator/lesion conditions plus at least 10/15 positive depth gains and finite losses. **Calibration-quality recovery status: failed development on 2026-07-22.** K4/e4 and K16/e4 reached only `65.633 ± 27.078%` and `65.567 ± 26.962%` at d4, with mean depth gains of `-0.583` and `-0.350` points. Each had only 1/3 positive depth gains and lesions; K16 also produced one nonfinite d4 loss. The frozen rule mechanically selects finite K4, but its negative depth gain prohibits spending untouched task seeds 6--8 on confirmation. Increasing target quality while A and W learn concurrently is therefore closed as a recovery branch. One feedback-first timescale screen is allowed next on development task seed 0. It will use an A-only calibration prefix with W and the output readout frozen, followed by the original context-gated simultaneous K1/e4 joint phase. The prefix uses layerwise K4 targets and the same `eta_A=0.01`; its length grid and selection rule must be committed before any run. This tests the specific coupled-basin diagnosis without increasing vectorizer capacity or changing the task. The frozen prefix-length grid is `{0, 100, 400}` A-only minibatches, crossed with d1/d4 and model seeds 0--2. Warmup consumes an isolated RNG stream and restores the loader generator, so every candidate's subsequent joint phase receives the same minibatch order and perturbation randomness. A prefix is development-eligible only if all six final losses are finite, mean d4 validation is at least 90%, at least 2/3 paired depth gains are positive, and the d4 lesion loses at least 2 points on average with at least 2/3 positive lesions. Among eligible prefixes, select the highest mean d4 endpoint, choosing the shortest prefix within one point. If none is eligible, the feedback-first branch closes without using task seeds 6--8. **Feedback-first status: failed development on 2026-07-22.** The 100- and 400-step prefixes raised mean d4 validation from `78.633 ± 24.866%` to `81.867 ± 8.779%` and `82.450 ± 14.226%`; both had 3/3 positive paired depth gains and lesions, versus 2/3 without warmup. Their mean gains were about 18 points and all losses were finite. Neither met the frozen 90% endpoint threshold, so no prefix advances and task seeds 6--8 remain unused. Timescale separation reduces basin variance but does not make plain SGD regression of the moving apical map accurate enough. One final task-seed-0 regression-mechanism screen is allowed before stopping C2 development. Normalized LMS divides each local A update by the squared norm of its complete instructional feature `[c, c outer tanh(h_top)]`, addressing the angle/gain separation in C5 without enlarging the vectorizer. It fixes the A-only prefix to 100 layerwise-K4 steps, keeps the joint phase at simultaneous K1/e4, and compares only `eta_A in {0.01, 0.05}` over d1/d4 and model seeds 0--2. Eligibility is identical to the feedback-first screen: all losses finite, d4 mean at least 90%, at least 2/3 positive depth gains, and lesion mean at least 2 points with at least 2/3 positive. Among eligible candidates select the highest d4 mean, preferring `eta_A=0.01` within one point. Failure closes further tent-map tuning; success permits one frozen confirmation on task seeds 6--8. **Normalized-regression status: failed development on 2026-07-22.** NLMS with `eta_A=0.01` and `0.05` reached only `75.150 ± 23.011%` and `73.350 ± 23.008%` at d4, recovering mean depth gains of `+4.433` and `+0.350` points. Each had only 2/3 positive gains and lesions. Both are ineligible; the tent-map recovery program is closed, no task seed 6--8 result was generated, and C2 remains a negative result for the amortized method. The passed direct-NP diagnosis and all failed recovery mechanisms remain part of the evidence rather than being hidden by another search branch. ### C3. Causal calibration is genuinely amortized **Status: passed on 2026-07-22.** The frozen five-seed confirmation retained 112.9% of the K16/e4 gain over DFA while using 16x fewer perturbation loss queries, 11x less calibration forward-equivalent work, and 5.3x less total training forward-equivalent work. Peak allocated memory fell from 803.5 MiB to 739.8 MiB. These numbers apply to the audited CIFAR-10 d20/w64 protocol; extrapolation to standard CNNs remains an oral-A question. Report wall time, forward-equivalent work, number of scalar loss queries, peak device memory, and the perturbation batch expansion. Sweep perturbation frequency and number of directions under both fixed epoch and fixed query/FLOP budgets. **Gate C3:** a protocol using at least 10x fewer perturbation loss queries than the current 16-directions-every-4-steps setting must preserve at least 90% of its gain over DFA. Pareto claims must hold in a hardware-independent cost coordinate as well as GTX-1080 wall time. The validation screen is frozen to CIFAR-10 d20/w64, five epochs, seed 0, and compares the current `K=16/every=4` reference with `K/every` pairs `1/4`, `2/8`, `4/16`, `1/8`, `1/16`, and `2/32`, plus DFA. Selection uses final validation accuracy and logical batch loss queries; the test split is not evaluated. `K=1/every=4` was selected: it used 16x fewer batch loss queries and retained 177.5% of the reference gain over DFA. This selection is now frozen; the protocol receives a five-seed test confirmation with no further tuning before C3 can be declared passed. ### C4. The comparison set contains the closest alternatives The main comparison must include exact BP, FA, DFA, direct/unamortized node perturbation, learned direct node-perturbation feedback (Lansdell et al. 2020), Dual Prop, and BurstCCN. The Lansdell method is the exact no-traffic/P=0 backbone of the current implementation, not merely a related baseline. BurstCCN is already demonstrated on CIFAR-10 and ImageNet and already models the Harnett BCI signatures, so it is the strongest dendritic/scaling comparator. EP remains an informative relaxation-based baseline. Forward-Forward and PEPITA are appendix context unless they become competitive under the frozen protocol. Each method receives a documented validation budget. Native-protocol and exact-architecture comparisons are labelled separately; neither substitutes for a matched compute/query comparison. **Status: passed on 2026-07-22.** In-repository BP/FA/DFA, direct NP, learned NP, PEPITA, Forward-Forward, and canonical EP are complete. Paper-faithful BurstCCN seed 0 reaches `80.07%` final / `80.10%` validation-selected / `80.25%` test-selected accuracy in `15451.3 s`, versus published `82.97 +/- 0.21%`. Author-code Dual Prop VGG16 seed 1988 reaches `92.46%` from the best-validation checkpoint in `23119.8 s`, versus published `92.41 +/- 0.07%`, with one test evaluation. Both records pass artifact, source, environment, dataset, completeness, finiteness, selection, and cost-definition checks in `finalize_accept.sh`. These one-seed method-native reproductions remain separate from matched-compute comparisons and from published multi-seed uncertainty. ### C5. Theory predicts the observed regimes The minimum theory package contains: 1. bias and variance of simultaneous multi-layer perturbation, including cross-layer interference; 2. a smooth-loss descent bound separating angular alignment, gain calibration, and curvature; 3. an innovation/SNR result for subtracting soma-predictable apical traffic; 4. an explicit complexity table covering learning phases, transport, loss queries, FLOPs, and memory. Numerical simulations must test the predicted scaling with depth, width, directions, perturbation scale, and predictor timescale. **Status: passed on 2026-07-22.** `THEORY.md` gives the exact simultaneous-Rademacher MSE with its cross-layer term, the Gaussian layerwise comparison, an `O(sigma^2)` antithetic bias bound, the smooth-loss angle/gain/curvature descent condition, the neutral innovation/SNR theorem and task-fit absorption result, and a phase/query/work/memory table. `experiments/verify_theory.py` checks depth, width, K, sigma, and predictor timescale deterministically; all assertions pass. ## Protocol discipline - Hyperparameters are selected on a validation split created only from the training set. - The test set is evaluated for frozen candidate protocols, never used to choose schedules. - Main results use five model seeds. Synthetic task claims additionally vary the task/teacher seed. - Every completed finite trajectory is retained; crashes and excluded runs are reported with a reason fixed before inspecting accuracy. - Pilot results remain versioned but cannot be pooled with frozen runs. - Every result records source revision, dirty state, protocol identifier, data split hash, logical loss queries, forward-equivalent work, peak memory, and wall time. ## Oral bar B: biological bridge **Status: development gate failed and confirmation is closed.** The complete 36-run preregistered screen is reported in `ORAL_B.md` and `RESULTS.md`. Local learning, residual decorrelation, outcome decoding, and the plasticity lesion were positive, but all causal-role sign-inversion indices were negative and acute online-control lesions had little effect. Seeds reserved for confirmation were not touched. After C1–C5 pass, test whether the model reproduces the qualitative Harnett signatures: - dendritic activity contains information absent after conditioning on somatic activity; - the residual decodes outcome/reward-related events; - residual sign follows the neuron's causal contribution to the objective; - residual predicts subsequent activity change or desired velocity; - selectively disrupting the residual impairs learning more than matched nonspecific disruption. The paper needs at least one prediction not used to construct SDIL. A preferred route is to request the original data/analysis from the authors and test the prediction out of sample; otherwise the claim is explicitly computational rather than a fit to cortical data. ## Oral bar A: standard deep architectures **Status: A3 failed; A4 was not opened and all confirmation test seeds remain untouched.** `ORAL_A.md` froze a validation-only ResNet-20 development funnel. A1 passed with `91.62%` BP validation accuracy. A2b selected channel-gated SDIL at `41.98%`, ahead of tuned DFA at `37.16%`, on the 10k/20-epoch screen. In full A3, DFA ended finite at `33.06%`; SDIL became nonfinite at epoch 89 and ended at `10.00%`, failing alignment, accuracy, and finiteness gates. The MAC gate alone passed. Per the stop rule, the five-seed ResNet-20/32/56 test panel was never run. The exact BatchNorm/local-gradient mechanics remain verified, but the original channel-gated training recipe does not establish standard-network scaling. A separate dynamic-innovation recovery is now frozen in `ORAL_A_RECOVERY.md` before any new endpoint. It does not reopen failed A3/A4: only complete D4 and oral-B R2 passes can unlock a 60-cell BP/DFA/clean-KP/dynamic ResNet-20/32/56 panel. D4 supplies the ten depth-20 KP/dynamic cells verbatim; exactly 50 new cells are permitted. A full pass is the sole 8-to-9 score gate and must show a positive paired depth benefit, not merely survival on another depth-flat task. **Independent oral-A-v2 status: passed.** After the calibrated oral-B-v2 confirmation separately satisfied the prerequisite, `ORAL_A_RECOVERY_V2.md` froze a new protocol without reopening the old panel. All 60 BP/DFA/clean-KP/dynamic ResNet-20/32/56 records pass. Dynamic innovation rises `91.584% -> 92.254% -> 92.760%`; every depth-20-to-56 seed pair improves, with a `1.176`-point mean gain, and mean early-third alignment remains `0.999687/0.999613/0.999423`. The internal milestone is therefore 9/10. The next evidence gap is a full matched strong-baseline and architecture crossover, not positive standard-depth utility. Post-failure diagnosis is recorded separately in `results/oral_a_failure_diagnosis.json`: prediction--target cosine remains below `8.51e-5` in magnitude, calibration MSE is indistinguishable from target power, and target power grows `2.13e9x` before nonfiniteness. The next development branch must therefore reduce causal-estimator variance in the representable channel-gated subspace; changing only the v1 threshold is not an allowed response. **Post-failure v2 status: causal-capture gate failed.** The structured representable-subspace estimator raised matched frozen-forward early/all-layer alignment from `0.001105/0.011664` to `0.007209/0.052740`. It passed the all-layer threshold but missed both the preregistered `0.01` early-layer threshold and the `0.01` absolute early-layer advantage. Per `ORAL_A_V2.md`, the full 200-epoch v2 run was not launched; test and confirmation seeds remain untouched. This localizes the remaining bottleneck to early-layer credit rather than merely spatial projection variance. The subsequent no-training representation oracle further decomposes that bottleneck. The existing spatial fields reach `0.05494` early alignment with unconstrained per-example coefficients; the actual output-error-conditioned family reaches `0.02397` on a disjoint 32-example batch, while learned v2 reaches `0.00721`. Naive local-average/channel-context fields fall slightly to `0.02246`. The next algorithmic branch should therefore address both causal sample efficiency and feedback context; simply lowering `eta_A`, adding the tested fields, or reopening V2-2 is not justified. **Vectorizer-space V3 status: causal-capture gate failed.** Directly estimating the A/G matrix target passed exact mechanics and reduced matched synthetic one-query MSE to `0.03381x`, but its selected real early alignment was `0.007139` versus the matched V2 reference's `0.007209`. All-layer alignment rose from `0.052740` to `0.062579`; three early/oracle checks failed and only the all-layer check passed. V3 full training was not launched, and test access remains sealed. This closes learning-rate tuning of the same output-error-only channel-gated family; a further branch must make a substantive feedback-context or cross-layer-noise change. The fixed post-failure hierarchical oracle identifies the substantive next change. Ordinary downstream activation context moves early alignment only from `0.031123` to `0.032069`, while held-out local maps from the actual residual-DAG child error fields reach `0.440525` (1x1), `0.819917` (3x3), and `0.999767` (3x3 plus the locally available child ReLU gate). This is not task evidence and uses exact child gradients as oracle inputs. It establishes that future development should learn spatial hierarchical feedback without weight copying, with random/fixed hierarchical FA and BurstCCN as mandatory baselines. **Fixed hierarchical-FA baseline status: short gate failed.** The audited ResNet-20 screen reaches `39.64%`, `41.58%`, and `43.52%` at hidden rates 0.01, 0.03, and 0.1. The selected run improves on matched DFA by 6.36 points and has early alignment `0.040406`, but misses its frozen 50% full-run threshold. HFA-S2 is therefore closed. This validates hierarchical feedback as a stronger baseline while ruling out fixed random hierarchy as the complete repair. The next mechanics branch perturbs every hierarchical feedback-parameter subspace in one antithetic pair and subtracts locally computed prediction moments. Its causal JVP and exact symmetric-limit delta rule pass, but it has no task claim until a frozen-forward alignment gate is preregistered and audited. **Hierarchical task-scalar V4 status: causal-capture gate failed.** Stable etaA 0.1 changes early/all-layer alignment from `-0.000263/0.011279` to only `-0.000180/0.017336`. EtaA 1 reaches early `0.000348` but produces a 58.48x feedback norm ratio; etaA 10 becomes nonfinite. All substantive gates fail, so no accuracy run opens. This closes simultaneous global scalar calibration of the full 267,904-parameter feedback path at 400 events. The next branch must use richer local child-response information and include learned FA/weight mirroring as an explicit inherited baseline before adding innovation. **Normalized response-mirror WM-1 status: passed.** Twenty batch-1 local observations at selected etaM 0.1 raise early/all alignment from `-0.000263/0.011279` to `0.446141/0.542159`, with feedback/forward cosine `0.915480`, norm ratios `[0.889,1.157]`, zero task-loss queries, and only `1.6305e9` MACs. All five gates pass. The bounded two-rate WM-2 accuracy screen is opened. This is inherited baseline performance; the paper score changes only if innovation later becomes load-bearing on top of the scalable mirror. **Normalized response-mirror WM-2 status: short accuracy gate failed narrowly.** Selected hidden LR 0.1 reaches `64.04%`, versus `43.52%` fixed HFA and `74.94%` BP, with early/all alignment `0.9393/0.9162`, zero task-loss queries, and `0.9968x` BP MACs. It misses the 65% threshold by 0.96 points and the within-10-point BP gate by 0.90 points, so WM-3 is closed. The next branch cannot tune this mirror recipe; it must make a substantive rule or information change and must not treat high alignment alone as task equivalence. **Residual response-mirror RRM-1 status: capture gate passed.** Selected etaM 0.1 reaches early/all alignment `0.665070/0.730253` and feedback/forward cosine `0.953680` after 20 local observations, with norm ratios `[0.884,1.018]`, zero task-loss queries, and `2.4458e9` MACs. All frozen gates pass. EtaM 0.3 is visibly unstable (maximum norm ratio 30.13), so the selected rate is not widened or retuned. The frozen two-rate RRM-2 short accuracy screen is opened. This is inherited baseline progress and cannot change the paper score until a load-bearing innovation experiment uses the substrate. **Residual response-mirror RRM-2 status: short accuracy gate passed narrowly.** Selected hidden LR 0.1 reaches `65.08%`, 9.86 points below matched BP at `74.94%`, with early/all alignment `0.999128/0.999377`, feedback/forward cosine `0.999902`, zero task-loss queries, and `0.9969x` BP MACs. It clears the two accuracy margins by only 0.08 and 0.14 points; the LR-0.03 run reaches 62.32%. Every frozen check nevertheless passes, so exactly one RRM-3 full validation run opens at the selected settings. The result remains baseline evidence and the reviewer score stays 5/10. **Residual response-mirror RRM-3 status: failed decisively.** The sole frozen full record ends at 10.00% accuracy and NaN validation loss. Training loss peaks at `1.5808e16` in epoch 92 and remains `8.9503e13` at epoch 200. Yet the decayed-LR endpoint reports early/all-layer alignment `0.8774/0.8487`, Q/W cosine `0.999998`, and norm ratios `[1.0002,1.0024]`. This closes RRM without cadence/rate rescue and demonstrates why endpoint cosine cannot substitute for a training-trajectory audit. No innovation or test panel opens from RRM. **Modified Kolen--Pollack baseline protocol: mechanics and both task gates passed.** The separate local reciprocal correlations, post-observation W/Q independence, and exact symmetric-limit momentum updates all pass at zero numerical error. `KP_BASELINE.md` freezes one 20-epoch full-development screen with no LR/decay grid and requires both validation accuracy and training-period feedback tracking. No KP task endpoint had been generated when this protocol was committed. KP is inherited Akrout et al. machinery and cannot raise the paper score; only a later load-bearing innovation ablation can do that. **Modified Kolen--Pollack KP-1 status: passed.** The sole frozen record reaches `82.66%` validation accuracy versus matched BP's epoch-20 `81.02%`, with early alignment `0.885583`, final feedback/forward cosine `0.901034`, and epoch-11--20 mean cosine `0.854594`. It is finite, uses zero task-loss queries, and costs `1.3261x` BP after charging reciprocal correlations. KP-2 is opened at exactly the frozen settings. This strongly repairs the feedback-substrate engineering problem but remains inherited Akrout et al. evidence, so the paper score stays 5/10. **Modified Kolen--Pollack KP-2 status: passed.** The sole frozen 200-epoch record reaches `91.26%` validation accuracy versus matched BP's `91.62%`. Early alignment is `0.999397`; final and epoch-151--200 mean feedback/forward cosines are both `0.999663`. The complete trajectory is finite, feedback uses zero task-loss queries, and cost remains `1.3261x` BP. This opens MT-1 without raising the reviewer score because the result is entirely inherited KP evidence. **KP mixed-traffic status: MT-0 passed; MT-1 failed.** `MIXED_TRAFFIC.md` fixes a four-to-one, initialization-calibrated soma-predictable traffic intervention and crosses raw apical activity, per-example norm-matched raw activity, and neutral-period innovation on the same reciprocal KP substrate. The zero-traffic limit, exact-predictor limit, gain calibration, norm/direction control, reciprocal local correlations, and parameter independence pass at zero or float64 machine error. A realistic 20-step neutral warmup leaves `0.1360` of traffic RMS versus the frozen `0.25` ceiling. Elementwise traffic/predictor work is reported separately from affine MACs, and every control pays the predictor schedule. The synthetic batch-128 matched path is finite on a GTX 1080 and peaks at 0.857 GB allocated after reset; no task endpoint was used for this hardware check. The complete MT-1 panel then becomes nonfinite in epoch 1 under all three signals and ends at 10.00% validation accuracy. Calibration error remains only `4.77e-7`, predictor warmup reaches residual ratio `0.139942`, cost is `1.3271x` BP, and feedback uses zero task-loss queries, so those mechanical checks pass while all performance, finite, alignment, and tracking checks fail. Per the frozen stop rule, MT-2 and MT-3 remain untouched, no weaker traffic ratio is allowed, and this standard-ResNet accept path closes at 5/10. The training-only post-failure audit localizes the common instability to stem `W[0]` and its momentum at steps 23/69/74 for raw/matched/innovation, respectively, with no validation or test evaluation. This narrows any future independent method to operator-level gain control or dissipativity; more endpoint tuning of the same residual-RMS gate is not justified. The follow-up S0 one-sided stability branch also closes before held-out access. Its frozen four-margin training-prefix grid has no eligible candidate: small margins still overflow BN state and grow weights to `1e17`, while larger margins retain finite tensors but violate loss, signal-ratio, weight, and momentum envelopes by orders of magnitude. No validation protocol opens. This rules out fixed sign-only predictor bias as the missing gain-control mechanism. **Dynamic neutral projection D1: passed.** `DYNAMIC_INNOVATION.md` freezes a single fast local controller rather than another margin grid. Every task batch provides a paired instruction-off soma/apical observation; the fast affine fit removes the current neutral residual coupling without reading task instruction or changing the slow predictor. All 352 training-only steps and states remain finite, maximum loss is `2.9281`, and the controller reduces a growing `0.005833` neutral/traffic ratio to at most `3.03e-8`. No held-out evaluation occurs, so this mechanics pass opens D2 without changing the score. **Dynamic neutral projection D2: passed.** The sole frozen 20-epoch record reaches `83.58%` validation accuracy, above clean KP's `82.66%` and 73.58 points above both failed MT-1 controls. Early alignment is `0.883830`, final feedback cosine is `0.904056`, all projection certificates pass, cost is `1.3261x` BP MACs plus explicitly reported elementwise work, and test remains untouched. Because this is one short seed, the score stays 5/10. The predeclared 200-epoch D3 validation endpoint is now open; only a complete D3 pass can make the 5-to-6 score update eligible and open independent test confirmation. Before observing D3, `DYNAMIC_INNOVATION.md` and executable scripts freeze D4 as a five-seed paired clean-KP/dynamic test panel. D4 is hard-gated on a complete D3 pass, evaluates test exactly once per run, and can raise 6-to-7 only through its predeclared accuracy, paired noninferiority, alignment, leakage, query, and cost checks. **Dynamic neutral projection D3: passed.** The sole frozen 200-epoch record passes all 19 gates at source revision `d945c42`: `91.18%` validation versus BP's `91.62%` and clean KP's `91.26%`, `0.999353` early alignment, `0.999663` late feedback cosine, zero task-loss queries, and `1.3261x` BP MACs. Exactly one validation and zero test evaluations occur. Per the frozen rule, the strict reviewer score rises from 5 to 6 and the already committed D4 paired five-seed test panel is now open. **Dynamic neutral projection D4: passed.** All ten untouched test records at source revision `0008f2c` pass the frozen audit. Across seeds 10--14, dynamic innovation reaches `91.584%` mean test accuracy versus clean KP's `91.388%`; the mean paired clean-minus-dynamic deficit is `-0.196` points and its one-sided 95% upper bound is `0.131` points. Every dynamic seed reaches at least `91.51%`, mean early alignment is `0.999687`, and all finite-trajectory, feedback-tracking, projection, leakage, query, cost, memory, hardware, provenance, split, and exactly-once test checks pass. Per the frozen rule the strict reviewer score rises from 6 to 7, establishing the accept bar and opening oral-B recovery R1. Oral-A remains sealed until the separately frozen oral-B R2 confirmation passes. R1 subsequently passes and selects `eta=0.1`, but the complete untouched R2 gate fails the outcome-vectorization and longitudinal signatures. Oral-A therefore remains sealed permanently under this frozen sequence. Prepare convolutional local-update primitives and ResNet-20/32/56 protocols early. Queue frozen runs opportunistically on authorized idle GPUs. Because BurstCCN already reports CIFAR-10 and ImageNet scaling, dataset scale alone is not novel. The oral-level target is a memorable joint result: near-BP accuracy in a standard architecture, a win over or materially simpler/lower-cost regime than BurstCCN, and a nondominated accuracy–hardware-independent-cost point. Scale alone is also insufficient if C1 does not show that somato-dendritic innovation is load-bearing. ## Stop conditions - Do not extend flattened-CIFAR MLP depth merely to obtain a larger depth number. - Do not add MNIST EP seeds unless a protocol bug invalidates the completed five-seed panel. - Do not promote a wall-time frontier that disappears under loss-query or FLOP accounting. - Do not describe the algorithm as a cortical implementation if simultaneous perturbation or supervision assumptions remain biologically unsupported.