# Oral-B recovery: temporal-difference somato-dendritic innovation This branch follows the complete negative result in `ORAL_B.md`; it does not reinterpret or overwrite that preregistration. It is motivated by Fig. 5 and Extended Data Fig. 13 of Francioni et al. (2026): P+ versus P- SD residuals separate epochs by the sign of recent error change, whereas absolute error magnitude alone does not separate the populations. ## Structural diagnosis The failed model always calibrated its vectorizer against the instantaneous squared-loss descent direction `e_t s_i`. Its P+ minus P- residual contrast is therefore proportional to current error magnitude. Conditional on lower current error during improving epochs, the preregistered sign-inversion index must be negative. Supplying error velocity as a second regressor cannot alter the target being regressed. All 36 negative signs are consistent with this structural mismatch. ## Candidate mechanism The recovery factorizes the instruction into two locally obtainable terms: 1. a slowly learned per-cell causal-role coefficient `m_i`, estimated from sparse antithetic perturbations and the scalar BCI cursor difference; 2. the within-episode performance innovation `delta_t = |e_(t-1)| - |e_t|`, reset at episode boundaries. The task instruction is `m_i delta_t`. Ordinary soma-predictable apical traffic is added before the same per-cell neutral predictor, so the plasticity signal remains the somato-dendritic innovation rather than a directly supplied role label. Forward updates use only current presynaptic context, local postsynaptic gain, and that innovation. The first recovery is plasticity-only (`kappa=0`): the empirical residual points along observed performance change, not necessarily a corrective online-control direction. For perturbation vector `xi`, the role target is ```text q_i = [(z(h + sigma xi) - z(h - sigma xi)) / (2 sigma)] xi_i, ``` whose expectation is the unknown causal role `s_i`. This consumes scalar cursor observations rather than gradients or reverse-mode differentiation. ## R0 mechanics gate R0 generates no task endpoint and touches no development or confirmation environment. It passes only if deterministic checks establish all of: - the mean perturbation role target has cosine above 0.99 with the analytic role used only by the diagnostic; - the local vectorizer update moves toward that role without reading it; - the neutral predictor exactly removes affine soma traffic in the controlled synthetic check; - episode-initial performance velocity is zero; - perfect role identification gives a strictly positive P+/P- error-change sign-inversion index; and - all original BCI pairing, lesion-mask, decoder, and finite checks still pass. R0 is implemented by `performance_velocity` in `sdil/bci.py` and audited by `experiments/bci_smoke.py` plus `experiments/verify_theory.py`. Passing R0 only permits a separately committed training-only structural screen after the D4 accept confirmation closes. No success-rate endpoint, hyperparameter selection, or oral-B score change is authorized by R0. ## R1 frozen development gate R1 is hard-gated in code on a complete D4 pass with reviewer score 7. It reuses only the already-development task seeds 0--2 and model seed 0; confirmation seeds 10--15 remain untouched. It tests exactly two forward rates, `{0.03, 0.1}`, with vectorizer rate `0.03`, 14 days, 64 episodes per day, 28 steps per episode, one role perturbation event every four temporal steps, `kappa=0`, and all other dynamics copied from the failed original screen. Before task learning, every condition receives 100 instruction-off batches of 64 neural-state probes. The per-cell neutral predictor uses these batches; two scalar antithetic cursor observations per example calibrate the learned-role conditions. Fixed-vectorizer ignores the role targets but consumes the same observations and random stream. The oracle-role condition explicitly reads the synthetic environment map only as a labelled diagnostic ceiling and is ineligible for selection. Neutral warmup must leave maximum predictor error at most `1e-5`; all warmup and online scalar observations are reported. The runner additionally requires the passed D4 gate itself to be tracked and the worktree to be clean. Every R1 record binds both the frozen protocol and that exact D4 gate by SHA-256; the complete-grid analyzer rejects any digest drift. This is an audit-only constraint and does not alter the frozen rates, seeds, conditions, metrics, or thresholds. For each rate and task seed, run four paired conditions on identical context, noise, perturbations, and evaluation trajectories: 1. intact learned-role temporal-difference innovation; 2. fixed random vectorizer; 3. plasticity lesion with role learning intact; and 4. an exact-role diagnostic oracle. Every one of the three task seeds must pass every requirement: - early-to-late success gain at least 10 points and final development success at least 70%; - final success at least 20 points above fixed vectorizer and no more than 10 points below the exact-role oracle; - causal-role sign-inversion index at least `0.01` and absolute CV correlation advantage of performance velocity over error magnitude at least `0.05`; - mean absolute innovation-soma correlation at most `0.10`, with raw minus innovation correlation at least `0.20`; - the plasticity lesion retains at most half the intact learning gain; and - learned role cosine at least `0.80`, with all trajectories and audited values finite. Select the eligible rate with the largest worst-task final success, breaking a tie by the smaller rate. If neither rate is eligible, the recovery closes and no confirmation is run. R1 is development evidence and cannot change the reviewer score. A sign that is positive merely because `delta_t` was inserted is therefore insufficient: it must coexist with actual learning, vectorizer necessity, innovation identification, and a causal plasticity lesion. The executable runner, complete-grid analyzer, and shell entry point are `experiments/bci_td_run.py`, `experiments/analyze_bci_td_development.py`, and `experiments/bci_td_development_screen.sh`. None may run before D4 closes. Once the deterministic R1 gate exists, later invocations may only verify exact equality and cannot overwrite it. ## R2 frozen untouched confirmation R2 is committed before any R1 task endpoint and is executable only if the complete R1 gate selects one eligible rate. It runs that rate without further selection over untouched task seeds 10--15 and model seeds 0--4: exactly 30 paired records, with no deletion or replacement. Each record repeats intact, fixed-vectorizer, plasticity-lesion, and exact-role diagnostic conditions on identical trajectories. Its separate 256-episode evaluation uses task seed `300000 + training_task_seed`. The R2 runner requires clean tracked code, protocol, D4 gate, and R1 gate and binds all of them by SHA-256. The task seed, rather than each of the five models sharing its environment, is the independent unit for uncertainty. Every aggregate confidence check first averages model seeds within each of the six task seeds and then uses a one-sided 95% Student-t bound with five degrees of freedom. R2 requires all of the following learning/plasticity checks: - mean intact final success at least 70%, every task mean at least 60%, and the lower confidence bound at least 60%; - mean early-to-late learning gain at least 10 points and its lower bound at least 5 points; - mean paired final gap over fixed vectorizer at least 20 points and its lower bound at least 10 points; - mean deficit to the exact-role diagnostic at most 10 points and its upper bound at most 20 points; - the lower bound of `0.5 * intact_gain - plasticity_lesion_gain` is nonnegative; and - mean learned-role cosine at least 0.80 and every one of 30 cosines at least 0.70. The biological checks retain the thresholds frozen for the original B1/B2 programme and add only task-cluster robustness bounds: - mean absolute innovation-soma correlation at most 0.10 (upper bound 0.12), with mean raw-minus-innovation correlation at least 0.20 (lower bound 0.15); - preceding surrounding-network state predicts amplification at mean balanced accuracy at least 55% (lower bound 52%), while decoder distance has mean residual correlation at least 0.10 (lower bound 0.02); - residual population outcome accuracy averages at least 57% (lower bound 53%) and exceeds soma by at least 3 points on average with nonnegative lower bound; - causal-role sign inversion is positive in at least 25/30 records and in all six task-cluster means; - performance velocity has at least 0.05 mean absolute CV-correlation advantage over instantaneous error, with nonnegative lower bound; and - early residual predicts late-minus-early somatic activity with mean correlation at least 0.30 and lower bound at least 0.10. The executable runner, immutable complete-grid analyzer, and entry point are `experiments/bci_td_confirmation.py`, `experiments/analyze_bci_td_confirmation.py`, and `experiments/bci_td_confirmation.sh`. The analyzer accepts exactly the 30 named records. A complete pass raises the strict reviewer score from 7 to 8 and establishes oral-B **innovation-guided plasticity**. Because `kappa=0`, it explicitly does not revive the failed online-control or desired-velocity claim. Any failed R2 gate is retained, leaves the score at 7, and closes this recovery without threshold repair.