# SDIL — results log (audited sections report n; timan107 GTX-1080 / ep_pascal) ## 2026-07-21 audit and version boundary Git was initialized after inheriting the project. The original code/results are preserved at `215acff`. The following audit findings changed the canonical implementation: 1. `49b86d6`: residual-block local updates now include the branch Jacobian `res_alpha=1/sqrt(depth)`. The omission preserved cosine direction but inflated update norms by `sqrt(depth)`, confounding cross-depth optimization comparisons. 2. `5255166`: the Harnett baseline is now a per-neuron affine soma-dendrite fit `ahat_l = p_l * h_l + b_l`, rather than a dense population map. This follows the paper's per-cell least-squares somatic-event to dendritic-event relation. 3. `c2769e6`: causal perturbation targets and predicted residuals are measured on the same pre-update network state. 4. `0f107d1` / `e0bfdfd`: simultaneous hidden-layer perturbations reduce calibration from quadratic to linear depth work; batching directions makes the reduction visible in GPU wall time. `66e3034` fixes diagnostics so raw-apical ablations report the signal actually used. 5. `2f3f44d`: the raw-apical control can be matched per sample to the innovation norm. This separates subtraction/direction from a trivial update-magnitude explanation. 6. `6093fe5` / `d6a00aa`: FA, PEPITA, Forward-Forward, and EP were reimplemented against their papers and reference code. EP uses persistent free particles; deep PEPITA requires an explicit learning rate; FA now logs the signal actually transported through its fixed feedback path. All older `deepres_*` files predate item 1. Their alignment directions remain informative, but their accuracy and stability should not be pooled with the corrected runs below. ## Method SDIL (Somato-Dendritic Innovation Learning), a local non-backprop rule inspired by Harnett 2026. Per hidden layer l: - forward: `u_l = W_l h_{l-1}`, `h_l = φ(u_l)` - apical feedback `a_l = A_l c` (c = output error e = softmax−onehot, broadcast) - default per-neuron predictor `ĥ_l = p_l ⊙ h_l + b_l`; teaching signal `r_l = a_l − ĥ_l` - three-factor update `ΔW_l = η (r_l ⊙ φ'(u_l)) h_{l-1}^T` - `A_l` LEARNED by amortized node perturbation (q_l ≈ −∇_{h_l}L, forward-only, no weight transport) - `P_l` learned on neutral (c=0) periods only, so it strips the soma-predictable nuisance without eating the teaching signal. The scalable perturbation mode injects independent Rademacher interventions at all hidden layers in the same antithetic forward evaluations. Cross-layer interference adds variance but averages to zero; it trains all `A_l` using `O(depth)` rather than `O(depth^2)` forward work. Distinct from the failed `~/sdrn` (fixed-random feedback + post-hoc residualization) and from Dual Prop / DFA / FA: the feedback pathway is *learned by causal perturbation*, so the residual is aligned by construction rather than being a random projection. ## Critical fix Node-perturbation estimator must divide by σ, not σ² (with δ=σξ, ĝ = δ·ΔL/σ² = ξ·ΔL/σ ⇒ ‖q‖≈‖∇L‖). The σ² bug made the teaching signal ~100× too large ⇒ effective LR ~100× ⇒ SDIL stalled. Cosine is magnitude-invariant so alignment looked fine while optimization blew up. Fixing it turned SDIL from "worse than DFA everywhere" to "matches DFA on MNIST, beats it on FMNIST/CIFAR". (Codex independently flagged the same σ-scale issue.) ## Verified narrow depth scaling (CIFAR-10, width 64, 5 epochs, five seeds) Mean ± sample SD. Alignment is the mean per-sample cosine over the earliest third of hidden layers at the final checkpoint. SDIL uses simultaneous perturbations with 16 batched directions. | depth | method | test acc (%) | early-third alignment | wall time (s) | |------:|:-------|-------------:|----------------------:|--------------:| | 30 | BP | 41.908 ± 0.697 | — | 29.8 ± 1.3 | | 30 | FA | 38.158 ± 0.517 | 0.593 ± 0.014 | 26.9 ± 1.9 | | 30 | DFA | 36.988 ± 0.524 | 0.087 ± 0.007 | 23.4 ± 0.8 | | 30 | SDIL | 41.198 ± 0.523 | 0.366 ± 0.007 | 47.8 ± 1.4 | | 60 | BP | 41.624 ± 0.781 | — | 58.4 ± 1.0 | | 60 | FA | 38.428 ± 0.387 | 0.577 ± 0.011 | 49.8 ± 1.8 | | 60 | DFA | 36.898 ± 0.444 | 0.047 ± 0.005 | 47.6 ± 5.6 | | 60 | SDIL | 41.246 ± 0.323 | 0.336 ± 0.013 | 96.4 ± 12.4 | The paired SDIL−DFA accuracy gap is `4.210 ± 0.426` points at d30 and `4.348 ± 0.554` at d60. SDIL wall time grows 2.02x when depth doubles, so the batched simultaneous calibration removes the former superlinear depth cost. Its accuracy is depth-stable and remains within 0.71 points of BP at d30 and 0.38 points at d60. Flattened CIFAR itself is depth-flat here, so these runs establish preservation rather than positive accuracy scaling. DFA's early-layer direction nearly halves with depth. Residual FA does **not** show that failure: its identity skips transport most of the state gradient exactly, leaving cosine near 0.58–0.59. Nevertheless FA plateaus around 38.3%, roughly three points below SDIL. Thus the supported claim is not “all FA signals lose directional alignment”; direction, signal scale, serial transport, and optimization efficiency must be separated. ## Harnett-specific residualization (MNIST, d3/w256, 15 epochs, five seeds) The matched-raw control preserves the raw apical direction but rescales every sample to the innovation's norm. | nuisance rho | raw apical acc (%) | matched raw acc (%) | residual acc (%) | |-------------:|-------------------:|--------------------:|-----------------:| | 0.00 | 97.450 ± 0.046 | 97.450 ± 0.046 | 97.450 ± 0.046 | | 0.05 | 96.438 ± 0.077 | 96.980 ± 0.148 | 97.382 ± 0.154 | | 0.20 | 95.512 ± 0.150 | 95.974 ± 0.144 | 97.370 ± 0.083 | | 0.50 | 10.380 ± 0.548 | 10.310 ± 0.633 | 97.348 ± 0.133 | At zero nuisance all three signals are identical. As normal soma-dendrite coupling grows, raw feedback degrades and eventually reaches chance; residual learning stays at 97.35–97.45%. Matching the raw signal's norm recovers only part of the low-rho loss and does nothing at rho=0.5. The strong result therefore depends on subtracting the soma-predictable component, not merely clipping or normalizing an oversized update. This is the most direct evidence that the Harnett-specific innovation operation is algorithmically necessary. ## Legacy accuracy (pre-audit implementation) | task | BP | DFA | SDIL | |------------------------------|-----------|-------|---------------| | MNIST d3 / d5 / d7 / d10 | 98.3 | 97.3 / 97.2 / 97.1 / 96.8 | 97.4 / 97.3 / 97.1 / 97.1 | | FashionMNIST d3 / d5 / d7 / d10 | 88–89 | 86.3 / 87.1 / 86.2 / 86.9 | 88.4 / 87.6 / 87.3 / 88.1 | | CIFAR-10 (no-BN residual) d5/10/20 | 44.2 / 44.8 / 45.5 | 40.6 / 39.9 / 37.4 | 42.4 / 42.6 / 42.4 | SDIL matches DFA on easy MNIST and **beats DFA on FashionMNIST (+1–2 pts) and CIFAR (+2 to +5 pts)**. On CIFAR the SDIL−DFA gap GROWS with depth (+1.8 → +2.7 → +5.0 at d5/10/20): DFA degrades with depth, SDIL holds. ## Legacy credit assignment / depth utility Per-layer alignment cos(r_l, −∇_{h_l}L), early-third average, CIFAR no-BN residual: | depth | DFA early-align | SDIL early-align | |-------|-----------------|------------------| | 5 | 0.55 | 0.83 | | 10 | 0.36 | 0.86 | | 20 | 0.25 | 0.88 | | 30 | 0.19 | 0.87 | DFA's teaching signal to early/deep layers collapses toward random (~0.05–0.19) as depth grows; SDIL stays ~0.85 and is depth-invariant. This is the "does credit reach early layers" metric that DFA fails and SDIL passes — the core advantage. ## Paper-faithful non-backprop baselines `PETITE` in the initial request is treated as **PEPITA**. Implementations and hyperparameters were checked against the [FA paper](https://www.nature.com/articles/ncomms13276), [PEPITA paper](https://proceedings.mlr.press/v162/dellaferrera22a.html) and [author code](https://github.com/GiorgiaD/PEPITA), [Forward-Forward paper](https://www.cs.toronto.edu/~hinton/absps/FFXfinal.pdf) and [reference implementation](https://github.com/mpezeshki/pytorch_forward_forward), and the [EP paper](https://www.frontiersin.org/journals/computational-neuroscience/articles/10.3389/fncom.2017.00024/full) and [author code](https://github.com/bscellier/Towards-a-Biologically-Plausible-Backprop). Native protocol results are not an equal-compute comparison: | method | architecture | budget | n | last test acc (%) | wall time (s) | |:-------|:-------------|:-------|--:|------------------:|--------------:| | PEPITA | d3/w256 | 60 epochs | 5 | 91.094 ± 0.214 | 89.2 ± 3.5 | | Forward-Forward | d3/w256 | 15 epochs/layer | 5 | 93.558 ± 0.232 | 68.2 ± 0.8 | | EP | d1/w500 | 25 epochs | 2 | 97.040 ± 0.269 | 839.8 ± 78.4 | PEPITA uses its two forward presentations, He-uniform initialization, shared 10% dropout mask, and an explicit deep-network learning rate of 0.001. Forward-Forward trains greedy local goodness layers and has no classifier readout. EP uses the original hard-sigmoid energy dynamics, random beta sign, 20 free steps, 4 nudged steps, layer rates `[0.1, 0.05]`, and persistent free particles for the fixed first 50k examples. This corrected EP is the strongest baseline and is also by far the most expensive. ### EP exact-architecture comparison All methods below use d1/w500, the same first 50k MNIST examples, and 25 epochs. EP retains its native raw-pixel energy protocol; BP/DFA/SDIL retain the shared z-score/tanh feedforward protocol. Forward parameter count is 397,510 for every row. | method | n | last test acc (%) | best test acc (%) | wall time (s) | |:-------|--:|------------------:|------------------:|--------------:| | BP | 2 | 98.190 ± 0.014 | 98.215 ± 0.007 | 15.1 ± 0.9 | | DFA | 2 | 97.620 ± 0.071 | 97.690 ± 0.099 | 9.9 ± 0.1 | | EP | 2 | 97.040 ± 0.269 | 97.220 ± 0.283 | 839.8 ± 78.4 | | SDIL | 2 | 97.510 ± 0.028 | 97.570 ± 0.028 | 15.4 ± 0.4 | SDIL is 0.47 points above EP on last-epoch mean and about 54.5x faster in wall time. This establishes that SDIL does not lose to the strongest baseline in the exact shallow architecture. It is **not** evidence of a unique scaling advantage: DFA is also strong on this easy d1 MNIST problem, and only two EP seeds have been run. At similar forward-parameter count but d3/w256, SDIL and EP are tied within noise (97.13% vs 97.04% last-epoch mean), again with a ~26x wall-time advantage for SDIL. ## How to run `experiments/run.py --mode {bp,dfa,sdil} --dataset {mnist,fmnist,cifar10} --depth D --residual {0,1} --act {tanh,gelu,silu,relu}` Batteries: `experiments/run_v2.sh "" "" `. Canonical local baselines: `experiments/baseline_sweep.sh`; original d1 EP: `experiments/ep_original_sweep.sh`; EP comparison: `experiments/ep_matched_sweep.sh`. Audited aggregate tables: `experiments/analyze_verified.py`. Smoke/mechanism checks: `experiments/smoke.py` and `experiments/baseline_smoke.py`. ## Open items - Compare simultaneous calibration at fixed loss-evaluation budgets and sweep directions (4/8/16/32). - Validate SDIL on a genuinely depth-necessary compositional task and a CNN; flattened CIFAR does not gain accuracy with depth. - Diagnose why residual FA has high directional cosine but lower accuracy: log update-norm ratios, parameter-space cosine, and serial feedback latency/cost. - Increase EP and exact-architecture comparisons from two to five seeds. If compute permits, run the author's d2 EP schedule (60 epochs, 150 free steps, 6 nudged steps) rather than inventing a cheap deep-EP approximation. - Tune SDIL schedules using a held-out validation split, never the test trajectory, then freeze the protocol before expanding the EP comparison. - Extend PEPITA and Forward-Forward beyond shallow MNIST before making any scaling claim about them.