1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
|
# ICLR-style reviewer scorecard
This is a deliberately adversarial paper-level assessment, not an average of the best runs. It is
updated only when a frozen experiment is complete and audited. Development pilots, partial jobs,
and post-hoc subsets cannot raise the score. Failed gates remain visible and narrow the claim.
## Scoring convention
The overall score uses a 1--10 ICLR-like scale: 1 strong reject, 3 reject, 5 borderline reject,
6 borderline accept, 8 accept, and 10 strong accept. Component scores use 1--4. The overall score
is a reviewer judgment rather than the arithmetic mean of the components. Confidence uses 1--5.
Every formal result report records:
- the previous and new overall score;
- soundness, novelty, significance, empirical support, and reproducibility;
- which frozen evidence changed the assessment;
- the strongest remaining reviewer objections;
- whether the evidence is development, validation, or untouched confirmation.
## Current paper-level assessment (2026-07-22)
| Dimension | Score | Strict reviewer assessment |
|:--|--:|:--|
| Soundness | 3/4 | Theory, local-gradient checks, causal diagnostics, cost accounting, and frozen stop rules are unusually careful. The learned apical vectorizer remains an unresolved failure mode. |
| Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. |
| Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, but the current flattened-CIFAR task does not benefit from depth and the frozen standard-ResNet recipe failed. |
| Empirical support | 2/4 | Five-depth scaling and the residualization ablation are strong. Useful-depth C2, broad endogenous C1, oral-B, and full ResNet A3 gates failed; the untouched A4 panel correctly remained sealed. |
| Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. |
| **Overall** | **5/10** | **Borderline reject: a strong core result without a completed standard-scale or useful-depth demonstration.** |
| Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. |
### Evidence already carrying the paper
- On flattened CIFAR-10, SDIL changes by only `-0.214 +/- 0.349` accuracy points from depth 5 to
depth 60, while DFA's early-layer alignment falls from `0.514` to `0.047`.
- Under strong soma-predictable apical traffic, raw and norm-matched raw learning fall to about
chance (`10.38%` and `10.31%`) while innovation learning retains `97.35%` accuracy.
- The local update has a proved descent condition and an explicit query/MAC/memory audit; direct
node perturbation isolates the learned vectorizer as the useful-depth bottleneck.
### Objections currently preventing acceptance confidence
1. The depth result is preservation on a depth-flat task, not evidence that SDIL uses added depth.
2. The amortized apical vectorizer failed the frozen useful-depth gate; direct causal targets work
but cost `68.4x` ordinary forward-equivalent work.
3. Innovation was not uniformly beneficial for arbitrary endogenous top-down traffic, so the
supported mechanism is narrower than the initial claim.
4. The frozen oral-B screen falsified the desired-velocity and Harnett error-derivative claims.
5. The frozen ResNet-20 A3 run became nonfinite and ended at chance. The failed gate forbids the
A4 test panel; completed native baselines improve fairness but do not supply SDIL scale evidence.
## Score trajectory and prospective gates
| Checkpoint | Overall | What changed | Remaining ceiling |
|:--|--:|:--|:--|
| Current audited package | 5 | Strong depth-preservation and residual-necessity evidence; negative gates retained | Standard useful scale is absent |
| Native baselines complete | 5 | BurstCCN is below its published endpoint; Dual Prop reproduces 92.46% versus 92.41%, with strict provenance and cost semantics | Fairness objection narrows, but SDIL gains no standard-scale evidence |
| Oral-A A1/A2 | 5 | BP reached 91.62%; short channel-gated SDIL reached 41.98% versus tuned DFA at 37.16% | Development screening alone cannot raise the score |
| Oral-A A3 fails | 5 | Full ResNet-20 SDIL became nonfinite at epoch 89 and ended at 10%; DFA ended finite at 33.06% | Standard-scale and oral-A claims are closed; A4 remains untouched |
| Oral-A-v2 causal capture fails | 5 | Structured perturbation improves early/all-layer alignment 6.5x/4.5x but misses the frozen early gate | Mechanism diagnosis sharpens; no full standard-scale endpoint opens |
| Oral-A-v3 causal capture fails | 5 | Direct A/G perturbation lowers synthetic estimator variance but leaves real early alignment at 0.0071 | Coefficient regression is not the dominant early bottleneck; full scale remains closed |
| Fixed hierarchical FA short gate fails | 5 | HFA reaches 43.52%, above DFA 37.16% and failed-v1 SDIL 41.98%, but below its frozen 50% full-run threshold | Spatial hierarchy helps; fixed random hierarchy remains insufficient and supplies no SDIL scale evidence |
| Hierarchical task-scalar V4 fails | 5 | Stable calibration leaves early alignment near zero; higher rates explode to 58.5x norm or become nonfinite after 800 queries | Correct mechanics and the right spatial family are insufficient when one global scalar estimates 267,904 feedback parameters |
| Normalized response mirror capture passes | 5 | Twenty local observations reach 0.446 early alignment and 0.915 feedback/forward cosine at zero task-loss queries | Strongly resolves the engineering bottleneck, but all credit belongs to an inherited weight-estimation baseline until innovation is load-bearing |
| Normalized response mirror short gate fails | 5 | WM reaches 64.04% and 0.939 early alignment at 0.9968x BP MACs, but misses its two accuracy gates by under one point | Strong inherited baseline and useful warning that alignment is insufficient; full run and any SDIL scale claim remain closed |
| Residual response mirror capture passes | 5 | RRM reaches 0.665 early and 0.730 all-layer alignment with 0.954 feedback/forward cosine after 20 zero-query local observations | Opens the frozen short accuracy gate, but remains inherited predictive weight estimation rather than SDIL evidence |
| Residual response mirror short gate passes | 5 | RRM reaches 65.08%, within 9.86 points of matched BP, at 0.9969x BP MACs and nearly exact alignment | Opens one full validation run; the inherited baseline still supplies no Harnett-specific evidence |
| Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available |
These are conditional reviewer forecasts, not promised scores. A failed stage leaves its negative
result in the record and can lower the score if it invalidates a current claim. Oral-B does not get
retroactively reopened by success on standard vision benchmarks.
## Evidence-to-score log
| Date / revision | Evidence status | Overall change | Reviewer interpretation |
|:--|:--|:--:|:--|
| 2026-07-22 / `2304e83` | Audit of all completed frozen branches | baseline → 5 | Strong preservation/residualization core, but no standard useful-scale result |
| 2026-07-22 / `c753f51`, `1b24c87`, `6d19078` | Existing 60-run innovation panel promoted to a strict main figure; conditional-projection and norm-direction identities made executable | 5 → 5 | Closes a presentation/theory objection and makes the narrow novelty legible, but adds no new held-out evidence and therefore earns no score inflation |
| 2026-07-22 / frozen Oral-A A1--A3 | A1 and A2 pass; full A3 SDIL becomes nonfinite and fails four of six checks; A4 untouched | 5 → 5 | Closes the standard-scale question negatively. The narrow mechanism paper survives, while any standard-ResNet or oral claim does not |
| 2026-07-22 / native C4 | BurstCCN and Dual Prop author-code records pass the strict audit; Dual Prop reproduces 92.46% test in 23119.8 s | 5 → 5 | Closes a baseline-fidelity objection and confirms a strong expensive comparator, but does not repair SDIL's failed A3 evidence |
| 2026-07-22 / Oral-A-v2 V2-1 | Six clean frozen-forward records: structured calibration raises early/all-layer alignment to 0.0072/0.0527 but fails two frozen advancement checks | 5 → 5 | Confirms representable-subspace variance was real, while showing that early-layer causal credit remains below the standard-depth gate; V2-2 and confirmation stay closed |
| 2026-07-22 / post-failure representation oracle | Current family has a 0.0240 cross-validated early-alignment ceiling on the fixed probe versus 0.0072 learned; unconstrained basis coefficients reach 0.0549 | 5 → 5 | Localizes both a causal-regression gap and an output-error-only capacity gap, but an oracle audit supplies no task-performance evidence |
| 2026-07-22 / Oral-A-v3 V3-1 | Four clean frozen-forward records: vectorizer-space estimation raises all-layer alignment to 0.0626 but leaves early alignment at 0.0071 and fails three advancement checks | 5 → 5 | Exact lower-variance mechanics do not repair early credit; no full ResNet or confirmation evidence is opened |
| 2026-07-22 / hierarchical oracle | Held-out gated 3x3 maps from exact child error fields reach 0.9998 early alignment; ordinary activation context remains 0.0321 | 5 → 5 | Strongly localizes the missing information to spatial hierarchical error fields, but exact-gradient oracle inputs provide no evidence that the proposed local learner can obtain them |
| 2026-07-22 / fixed HFA S1 | Three clean matched ResNet-20 records; selected HFA reaches 43.52% and 0.0404 early alignment but misses the frozen 50% gate | 5 → 5 | Establishes a stronger zero-query local baseline and confirms hierarchy helps, while closing an uncalibrated full run and leaving standard SDIL evidence absent |
| 2026-07-22 / hierarchical V4-1 | Four frozen-forward records: etaA 0.1 leaves early alignment near zero, etaA 1 explodes one feedback norm to 58.5x, and etaA 10 is nonfinite | 5 → 5 | Closes global-task-scalar calibration of the full hierarchy at the fixed budget; richer local information is required before another accuracy endpoint |
| 2026-07-22 / response-mirror WM-1 | Four clean frozen-forward records; selected etaM 0.1 reaches 0.4461 early and 0.5422 all-layer alignment with zero task-loss queries | 5 → 5 | Opens a strong inherited-baseline accuracy test and isolates information source as the bottleneck, but supplies no Harnett-specific task evidence |
| 2026-07-22 / response-mirror WM-2 | Two clean short ResNet records; selected WM reaches 64.04% with 0.9393 early alignment and 0.9968x BP MACs, missing both accuracy gates narrowly | 5 → 5 | Substantially strengthens the comparator and cost story, but closes its full run and demonstrates that high credit cosine alone is not a scale result |
| 2026-07-22 / residual response-mirror RRM-1 | Four frozen-forward records; selected etaM 0.1 reaches 0.6651 early and 0.7303 all-layer alignment with controlled norms, zero task-loss queries, and 2.4458e9 MACs | 5 → 5 | Residual prediction removes much of the fixed-point estimator noise and opens the short task gate, but the gain belongs to an inherited baseline and does not establish somato-dendritic innovation |
| 2026-07-22 / residual response-mirror RRM-2 | Two clean short ResNet records; selected RRM reaches 65.08%, 9.86 points below BP, with 0.9991 early alignment and 0.9969x BP MACs | 5 → 5 | Narrowly opens the full baseline run and strengthens the efficient comparator; the remaining accuracy gap despite near-exact direction warns that alignment is not trajectory equivalence |
Future rows are appended only after an audited frozen stage. A score staying flat is informative:
engineering, theory exposition, or visualization may make the paper more defensible without
resolving the empirical objection that determines the recommendation.
|