summaryrefslogtreecommitdiff
path: root/REVIEW_SCORECARD.md
blob: eded7194c46f92a586caf7d77a0d6b49e19bdf58 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
# ICLR-style reviewer scorecard

This is a deliberately adversarial paper-level assessment, not an average of the best runs. It is
updated only when a frozen experiment is complete and audited. Development pilots, partial jobs,
and post-hoc subsets cannot raise the score. Failed gates remain visible and narrow the claim.

## Scoring convention

The overall score uses a 1--10 ICLR-like scale: 1 strong reject, 3 reject, 5 borderline reject,
6 borderline accept, 8 accept, and 10 strong accept. Component scores use 1--4. The overall score
is a reviewer judgment rather than the arithmetic mean of the components. Confidence uses 1--5.

Every formal result report records:

- the previous and new overall score;
- soundness, novelty, significance, empirical support, and reproducibility;
- which frozen evidence changed the assessment;
- the strongest remaining reviewer objections;
- whether the evidence is development, validation, or untouched confirmation.

## Current paper-level assessment (2026-07-27)

| Dimension | Score | Strict reviewer assessment |
|:--|--:|:--|
| Soundness | 4/4 | Theory, exact local-update checks, causal lesions, clustered uncertainty, cost accounting, and retained frozen failures make the implemented claims unusually well identified. |
| Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. |
| Significance | 4/4 | Dynamic innovation now tracks BP/KP while gaining accuracy from ResNet-20 to ResNet-56 over five paired seeds; a separate local actor--critic reproduces role-vectorized outcome surprise. Original-data biological validation remains absent. |
| Empirical support | 4/4 | Residual necessity, untouched ResNet-20 confirmation, the complete 60-record standard-depth panel, and a complete 30-record task-clustered BCI confirmation are strong. Useful-depth C2, broad endogenous C1, old oral-B, A3, and the original MT-1 remain disclosed failures. |
| Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. |
| **Overall** | **9/10** | **Internal oral-A milestone: the preregistered 60-record panel establishes standard ResNet depth scaling in addition to load-bearing innovation and role-vectorized TD outcome surprise.** |
| Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. |

The conservative external-review forecast is **8/10**, not 9.  The
standard-depth result is unusually strong, but its clean-task transport
substrate is inherited reciprocal KP and the Harnett-specific gain remains a
conditional mixed-traffic result.  A reviewer can also discount the synthetic
BCI because terminal reward is supplied and the psychometric target range is
calibrated per policy.  The 9/10 value is the repository's predeclared
evidence milestone; 8/10 is the recommendation I would actually submit as a
reviewer today.  The ongoing 81-cell BP/FA/DFA/PEPITA/FF/EP/DualProp/KP/SDIL
crossover targets the strongest remaining baseline-and-cost objection and
cannot change this score until its complete frozen panels are audited.

### Evidence already carrying the paper

- On flattened CIFAR-10, SDIL changes by only `-0.214 +/- 0.349` accuracy points from depth 5 to
  depth 60, while DFA's early-layer alignment falls from `0.514` to `0.047`.
- Under strong soma-predictable apical traffic, raw and norm-matched raw learning fall to about
  chance (`10.38%` and `10.31%`) while innovation learning retains `97.35%` accuracy.
- On untouched CIFAR-10 test endpoints, dynamic innovation reaches `91.584%` versus clean KP's
  `91.388%` over five paired ResNet-20 seeds; its paired deficit upper bound is only `0.131`
  points and every frozen mechanism/cost invariant passes.
- In the complete 60-record standard-depth panel, dynamic innovation rises
  from `91.584%` at ResNet-20 to `92.254%` at ResNet-32 and `92.760%` at
  ResNet-56.  All five paired seeds improve from depth 20 to 56; mean gain is
  `1.176` points, while early-third alignment remains
  `0.99969/0.99961/0.99942` across depths 20/32/56.
- In the untouched six-task by five-model BCI panel, final task success is
  `100%`, terminal residual outcome decoding is `99.83%`, the acute
  outcome-lesion separation drop is `0.400`, critic expectedness is `0.319`,
  and all 30 causal signs are positive after label-free calibration.
- The local update has a proved descent condition and an explicit query/MAC/memory audit; direct
  node perturbation isolates the learned vectorizer as the useful-depth bottleneck.

### Strongest remaining objections

1. The matched standard-depth panel contains BP, tuned DFA, clean KP, and
   dynamic innovation, but not yet matched PEPITA, Forward-Forward, EP, Dual
   Propagation, and ordinary FA at every depth.  The frozen 81-cell crossover
   is active but incomplete.
2. The clean scaling result primarily establishes the inherited reciprocal-KP
   substrate; SDIL-specific necessity still comes from controlled predictable
   traffic and cannot be inferred from clean KP/SDIL equality.
3. The amortized perturbation-trained apical vectorizer failed the frozen
   useful-depth gate; the successful standard-depth path uses reciprocal local
   plasticity and a paired neutral microphase.
4. Innovation was not uniformly beneficial for arbitrary endogenous top-down traffic, so the
   supported mechanism is narrower than the initial claim.
5. The passed BCI is synthetic: reward is directly supplied, causal roles are
   experimenter-defined for diagnostics, and target quantiles are calibrated
   on a separate cursor split. Longitudinal prediction remains failed, and no
   original Francioni/Harnett event-level data are tested.

## Score trajectory and prospective gates

| Checkpoint | Overall | What changed | Remaining ceiling |
|:--|--:|:--|:--|
| Current audited package | 9 | D4 confirms load-bearing ResNet-20 innovation; calibrated oral-B-v2 confirms role-vectorized TD outcome surprise; oral-A-v2 confirms five-seed ResNet-20/32/56 scaling over 60 records | Complete matched strong-baseline crossover and original-data biological validation remain absent |
| Native baselines complete | 5 | BurstCCN is below its published endpoint; Dual Prop reproduces 92.46% versus 92.41%, with strict provenance and cost semantics | Fairness objection narrows, but SDIL gains no standard-scale evidence |
| Oral-A A1/A2 | 5 | BP reached 91.62%; short channel-gated SDIL reached 41.98% versus tuned DFA at 37.16% | Development screening alone cannot raise the score |
| Oral-A A3 fails | 5 | Full ResNet-20 SDIL became nonfinite at epoch 89 and ended at 10%; DFA ended finite at 33.06% | Standard-scale and oral-A claims are closed; A4 remains untouched |
| Oral-A-v2 causal capture fails | 5 | Structured perturbation improves early/all-layer alignment 6.5x/4.5x but misses the frozen early gate | Mechanism diagnosis sharpens; no full standard-scale endpoint opens |
| Oral-A-v3 causal capture fails | 5 | Direct A/G perturbation lowers synthetic estimator variance but leaves real early alignment at 0.0071 | Coefficient regression is not the dominant early bottleneck; full scale remains closed |
| Fixed hierarchical FA short gate fails | 5 | HFA reaches 43.52%, above DFA 37.16% and failed-v1 SDIL 41.98%, but below its frozen 50% full-run threshold | Spatial hierarchy helps; fixed random hierarchy remains insufficient and supplies no SDIL scale evidence |
| Hierarchical task-scalar V4 fails | 5 | Stable calibration leaves early alignment near zero; higher rates explode to 58.5x norm or become nonfinite after 800 queries | Correct mechanics and the right spatial family are insufficient when one global scalar estimates 267,904 feedback parameters |
| Normalized response mirror capture passes | 5 | Twenty local observations reach 0.446 early alignment and 0.915 feedback/forward cosine at zero task-loss queries | Strongly resolves the engineering bottleneck, but all credit belongs to an inherited weight-estimation baseline until innovation is load-bearing |
| Normalized response mirror short gate fails | 5 | WM reaches 64.04% and 0.939 early alignment at 0.9968x BP MACs, but misses its two accuracy gates by under one point | Strong inherited baseline and useful warning that alignment is insufficient; full run and any SDIL scale claim remain closed |
| Residual response mirror capture passes | 5 | RRM reaches 0.665 early and 0.730 all-layer alignment with 0.954 feedback/forward cosine after 20 zero-query local observations | Opens the frozen short accuracy gate, but remains inherited predictive weight estimation rather than SDIL evidence |
| Residual response mirror short gate passes | 5 | RRM reaches 65.08%, within 9.86 points of matched BP, at 0.9969x BP MACs and nearly exact alignment | Opens one full validation run; the inherited baseline still supplies no Harnett-specific evidence |
| Residual response mirror full gate fails | 5 | RRM ends at 10% with NaN validation loss although endpoint Q/W cosine is 0.999998 | Closes intermittent mirroring and proves endpoint alignment is an inadequate trajectory certificate; KP remains the separate strong substrate |
| Modified KP short gate passes | 5 | KP reaches 82.66% versus BP's epoch-20 81.02%, with 0.886 early alignment and 1.326x BP MACs | Strongly solves the substrate at useful scale, but all credit belongs to inherited reciprocal plasticity until innovation is load-bearing |
| Modified KP full gate passes | 5 | KP reaches 91.26% versus matched BP's 91.62%, with 0.9994 early alignment, 0.9997 late feedback cosine, zero queries, and 1.326x BP MACs | Establishes a stable near-BP ResNet substrate and opens MT-1, but inherited KP evidence cannot raise the SDIL score |
| Mixed-traffic MT-1 fails | 5 | Raw, norm-matched raw, and innovation all become nonfinite in epoch 1 and end at 10%; calibration and predictor warmup still pass | Closes the controlled standard-ResNet innovation path; MT-2/MT-3 remain untouched and no weaker traffic rescue is allowed |
| Mixed-traffic MT-2 | not opened | The prerequisite MT-1 gate failed | No full seed-0 validation claim is available |
| Mixed-traffic MT-3 | not opened | MT-2 was never opened; all five confirmation seeds remain untouched | No controlled ResNet innovation confirmation claim is available |
| Dynamic projection D1 | 5 | All 352 training-only steps remain finite while the fast neutral fit holds residual coupling near zero | Mechanics only; no held-out endpoint |
| Dynamic projection D2 | 5 | One frozen 20-epoch record reaches 83.58%, above clean KP's 82.66%, with all stability and cost gates passing | Short single-seed validation cannot establish the full scaling claim |
| Dynamic projection D3 | 6 | All 19 frozen checks pass at 91.18%, within 0.44/0.08 points of BP/clean KP, with 0.9994 early alignment and 1.326x BP MACs | One validation seed cannot establish robustness |
| Dynamic projection D4 | 7 | All ten untouched records pass: dynamic 91.584% versus clean KP 91.388%, paired upper deficit bound 0.131 points, early alignment 0.999687, and no invariant failures | Establishes ResNet-20 robustness/noninferiority, not positive depth utility |
| Oral-B recovery R1/R2 | failed at R2 | R1 selects eta 0.1 with 98.05% worst-task success; untouched R2 retains 99.53% mean success and 30/30 positive signs but fails outcome-vectorization and longitudinal gates | Score remains 7; the joint oral-B claim is not established |
| Oral-B-v2 initial grid | failed at development | All 24 records preserve role learning and residual identification but fail from a one-quarter dense-signal cold start | Failure retained; no confirmation touched |
| Oral-B-v2 fixed-target recovery | failed at development | Two seeds pass 18/18; the third passes 17/18 but its fixed target ladder has 98.96% success | Mechanism works, absolute assay scale does not generalize |
| Oral-B-v2 calibrated R1/R2 | 8 | Three fresh development seeds pass, then all 30 untouched records and every task-cluster bound pass under independent label-free calibration/evaluation splits | Establishes synthetic outcome surprise; does not establish cortex or added-depth utility |
| Oral-A dynamic depth recovery | closed | A 60-cell ResNet-20/32/56 BP/DFA/clean-KP/dynamic panel was frozen before any new endpoint | Its oral-B R2 prerequisite failed, so none of the 50 new cells may run |
| Oral-A-v2 standard-depth scaling | 9 | After the calibrated BCI-v2 prerequisite passed, all 60 records and every preregistered accuracy, depth-gain, alignment, mechanism, cost, query, and memory check pass | Strong matched-baseline breadth and original neural-data validation remain absent |
| Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available |

These are conditional reviewer forecasts, not promised scores. A failed stage leaves its negative
result in the record and can lower the score if it invalidates a current claim. The original
oral-B branch and both v2 development failures remain failed. The calibrated
v2 pass does not reopen the old oral-A panel; it permits only a new
independently frozen oral-A-v2 protocol.

## Evidence-to-score log

| Date / revision | Evidence status | Overall change | Reviewer interpretation |
|:--|:--|:--:|:--|
| 2026-07-22 / `2304e83` | Audit of all completed frozen branches | baseline → 5 | Strong preservation/residualization core, but no standard useful-scale result |
| 2026-07-22 / `c753f51`, `1b24c87`, `6d19078` | Existing 60-run innovation panel promoted to a strict main figure; conditional-projection and norm-direction identities made executable | 5 → 5 | Closes a presentation/theory objection and makes the narrow novelty legible, but adds no new held-out evidence and therefore earns no score inflation |
| 2026-07-22 / frozen Oral-A A1--A3 | A1 and A2 pass; full A3 SDIL becomes nonfinite and fails four of six checks; A4 untouched | 5 → 5 | Closes the standard-scale question negatively. The narrow mechanism paper survives, while any standard-ResNet or oral claim does not |
| 2026-07-22 / native C4 | BurstCCN and Dual Prop author-code records pass the strict audit; Dual Prop reproduces 92.46% test in 23119.8 s | 5 → 5 | Closes a baseline-fidelity objection and confirms a strong expensive comparator, but does not repair SDIL's failed A3 evidence |
| 2026-07-22 / Oral-A-v2 V2-1 | Six clean frozen-forward records: structured calibration raises early/all-layer alignment to 0.0072/0.0527 but fails two frozen advancement checks | 5 → 5 | Confirms representable-subspace variance was real, while showing that early-layer causal credit remains below the standard-depth gate; V2-2 and confirmation stay closed |
| 2026-07-22 / post-failure representation oracle | Current family has a 0.0240 cross-validated early-alignment ceiling on the fixed probe versus 0.0072 learned; unconstrained basis coefficients reach 0.0549 | 5 → 5 | Localizes both a causal-regression gap and an output-error-only capacity gap, but an oracle audit supplies no task-performance evidence |
| 2026-07-22 / Oral-A-v3 V3-1 | Four clean frozen-forward records: vectorizer-space estimation raises all-layer alignment to 0.0626 but leaves early alignment at 0.0071 and fails three advancement checks | 5 → 5 | Exact lower-variance mechanics do not repair early credit; no full ResNet or confirmation evidence is opened |
| 2026-07-22 / hierarchical oracle | Held-out gated 3x3 maps from exact child error fields reach 0.9998 early alignment; ordinary activation context remains 0.0321 | 5 → 5 | Strongly localizes the missing information to spatial hierarchical error fields, but exact-gradient oracle inputs provide no evidence that the proposed local learner can obtain them |
| 2026-07-22 / fixed HFA S1 | Three clean matched ResNet-20 records; selected HFA reaches 43.52% and 0.0404 early alignment but misses the frozen 50% gate | 5 → 5 | Establishes a stronger zero-query local baseline and confirms hierarchy helps, while closing an uncalibrated full run and leaving standard SDIL evidence absent |
| 2026-07-22 / hierarchical V4-1 | Four frozen-forward records: etaA 0.1 leaves early alignment near zero, etaA 1 explodes one feedback norm to 58.5x, and etaA 10 is nonfinite | 5 → 5 | Closes global-task-scalar calibration of the full hierarchy at the fixed budget; richer local information is required before another accuracy endpoint |
| 2026-07-22 / response-mirror WM-1 | Four clean frozen-forward records; selected etaM 0.1 reaches 0.4461 early and 0.5422 all-layer alignment with zero task-loss queries | 5 → 5 | Opens a strong inherited-baseline accuracy test and isolates information source as the bottleneck, but supplies no Harnett-specific task evidence |
| 2026-07-22 / response-mirror WM-2 | Two clean short ResNet records; selected WM reaches 64.04% with 0.9393 early alignment and 0.9968x BP MACs, missing both accuracy gates narrowly | 5 → 5 | Substantially strengthens the comparator and cost story, but closes its full run and demonstrates that high credit cosine alone is not a scale result |
| 2026-07-22 / residual response-mirror RRM-1 | Four frozen-forward records; selected etaM 0.1 reaches 0.6651 early and 0.7303 all-layer alignment with controlled norms, zero task-loss queries, and 2.4458e9 MACs | 5 → 5 | Residual prediction removes much of the fixed-point estimator noise and opens the short task gate, but the gain belongs to an inherited baseline and does not establish somato-dendritic innovation |
| 2026-07-22 / residual response-mirror RRM-2 | Two clean short ResNet records; selected RRM reaches 65.08%, 9.86 points below BP, with 0.9991 early alignment and 0.9969x BP MACs | 5 → 5 | Narrowly opens the full baseline run and strengthens the efficient comparator; the remaining accuracy gap despite near-exact direction warns that alignment is not trajectory equivalence |
| 2026-07-22 / residual response-mirror RRM-3 | The sole frozen full record reaches only 10%, has NaN validation loss, and peaks at `1.5808e16` training loss despite 0.999998 endpoint Q/W cosine | 5 → 5 | Closes intermittent residual mirroring and turns alignment latency into a measured negative; it neither weakens the surviving SDIL mechanism claim nor supplies positive scale evidence |
| 2026-07-22 / modified KP-1 | One clean constant-LR ResNet record reaches 82.66%, above BP's epoch-20 81.02%, with 0.8856 early alignment, 0.8546 train-period feedback cosine, zero queries, and 1.326x BP MACs | 5 → 5 | Establishes a stable strong substrate under a gate that observes the training trajectory, but reciprocal Kolen--Pollack plasticity is prior art and adds no Harnett-specific evidence |
| 2026-07-22 / `2ef7f94` modified KP-2 | One clean full ResNet record reaches 91.26%, 0.36 points below matched BP, with 0.9994 early alignment, 0.9997 late feedback cosine, zero queries, and 1.326x BP MACs | 5 → 5 | Removes substrate stability as the immediate blocker and opens the load-bearing innovation experiment, but prior-art reciprocal plasticity earns no novelty credit |
| 2026-07-22 / mixed-traffic MT-0 | Exact raw/matched/innovation mechanics pass; synthetic ResNet-20 matched batch peaks at 0.857 GB allocated on GTX 1080 without task data | 5 → 5 | Removes implementation, graph, and memory objections before endpoint access; supplies no evidence yet that innovation is useful |
| 2026-07-22 / `f310ad5` mixed-traffic MT-1 | All three frozen conditions become nonfinite in epoch 1 and end at 10%; ratio calibration, predictor warmup, zero-query, and cost checks pass | 5 → 5 | Fails to make innovation load-bearing on a standard ResNet and closes MT-2/MT-3; the prior mechanism-only evidence survives but empirical support does not improve |
| 2026-07-22 / `6531aed`, `d818fc2` MT-1 diagnosis | Training-only localization finds the first active failure at stem steps 23/69/74 for raw/matched/innovation; residual coupling yields the verified multiplicative operator `D W C` | 5 → 5 | Sharpens the failure into an operator-stability requirement and improves soundness of the negative analysis, but adds no successful held-out SDIL endpoint |
| 2026-07-22 / `8d28fd9` stability S0 | A frozen training-only four-margin grid has no eligible candidate; sign-certified margins either overflow BN state or produce enormous transient loss and parameter growth | 5 → 5 | Closes fixed one-sided predictor bias and motivates a two-sided dynamic gain condition, but no validation endpoint or empirical support is gained |
| 2026-07-22 / `6b85249` dynamic projection D1 | The sole frozen training-prefix record passes 14/14 checks; a 0.0058 neutral residual is reduced to 3.0e-8 while weights, momentum, BN, and loss stay bounded | 5 → 5 | Establishes a nontrivial operator-stability repair, but training-only mechanics cannot change the recommendation |
| 2026-07-22 / `9bc58d6` dynamic projection D2 | The sole frozen short validation record reaches 83.58% versus clean KP's 82.66% and failed controls' 10%, with 0.884 early alignment, zero queries, and 1.326x BP MACs | 5 → 5 | First positive standard-ResNet load-bearing innovation endpoint, but one short seed is insufficient; full D3 and independent confirmation remain decisive |
| 2026-07-22 / `15c60d0` dynamic projection D3 | The sole frozen full validation record passes 19/19 checks at 91.18% versus BP's 91.62% and clean KP's 91.26%, with 0.9994 early alignment, zero queries, and 1.326x BP MACs | 5 → 6 | Resolves the primary full-scale objection and opens the already frozen D4 panel; one validation seed and narrow biological scope prevent a stronger recommendation |
| 2026-07-22 / `32122d0` oral-B recovery freeze | Before any recovery task endpoint, the untouched 30-record R2 protocol, task-cluster uncertainty, original B1/B2 signatures, plasticity lesion, digest binding, and immutable analyzer are executable | 6 → 6 | Removes a preregistration gap but supplies no empirical evidence; R1 is still sealed behind D4 and confirmation seeds remain untouched |
| 2026-07-22 / `1853620` oral-A recovery freeze | Before any new standard-depth endpoint, a D4-reusing 60-cell ResNet-20/32/56 panel and exact runner/analyzer contract are executable, with BP/DFA/KP controls, positive depth gain, alignment, projection, query, MAC, and memory gates | 6 → 6 | Precommits the final oral claim without bypassing the required sequence; no experiment opens unless D4 reaches 7 and oral-B R2 reaches 8 |
| 2026-07-23 / `2fed62a` dynamic projection D4 | All ten untouched paired test records pass the frozen gate: dynamic 91.584% versus clean KP 91.388%, clean-minus-dynamic upper bound 0.131 points, 0.999687 early alignment, and zero invariant failures | 6 → 7 | Establishes the strict accept bar with independent ResNet-20 robustness/noninferiority; opens oral-B R1 but does not support added-depth or desired-velocity claims |
| 2026-07-23 / `1ff7cb7` oral-B recovery R1 | The complete two-rate development grid selects eta 0.1; all 30 seed-level checks pass, with 98.05% worst-task final success, positive sign inversion, and 0.945 minimum role cosine | 7 → 7 | Strong development support for the repaired causal-role/velocity factorization, but R1 cannot change the score and only opens untouched R2 |
| 2026-07-23 / `03c94a1` oral-B recovery R2 | Thirty untouched records retain 99.53% mean success, 90.45-point gain, 30/30 positive signs, and strong decorrelation, but fail seven population-vectorization/longitudinal checks | 7 → 7 | The recovery fixes learning, causal role, and sign but not the broader Harnett-like signature; oral-B and oral-A close without threshold repair |
| 2026-07-23 / `2a6f72e` audited D4 main figure | The strict renderer independently rechecks the ten D4 records and visualizes paired test accuracy, layerwise raw-versus-innovation direction, all 200 tracking epochs, and explicit neutral/MAC/memory/wall costs | 7 → 7 | Makes the accept evidence reviewable without adding or selecting data; presentation improves, but visualization alone cannot repair oral-B or justify score inflation |
| 2026-07-23 / `87cfb93` evidence-bound manuscript | A 3,238-word working draft binds 34 central numbers and all four figures to source manifests, retains the passed D4 gate and all seven failed R2 checks, and is re-audited by the accept finalizer | 7 → 7 | Substantially improves submission readiness and guards against claim drift; it adds no empirical evidence, so soundness and recommendation do not inflate |
| 2026-07-23 / `eb021a6` oral-B-v2 initial R1 | The complete 24-record grid learns causal roles and identifies innovations but reaches at most 0.78% evaluation success because terminal reward remains unreachable | 7 → 7 | Localizes a cold-start created by scaling the only pre-reward drive to one quarter; confirmation remains untouched |
| 2026-07-23 / `378e68d` fixed-target recovery R1 | Unit dense velocity restores 100% task learning and 17--18 biological checks per seed, but one fresh seed has 98.96% challenge success | 7 → 7 | Confirms the algorithmic repair while falsifying an absolute target ladder as a model-independent assay |
| 2026-07-23 / `70e180c`, `9a8c057` calibrated oral-B-v2 R1/R2 | Three new development seeds pass all 18 gates; all 30 untouched confirmation records then pass every clustered learning, innovation, decoder, lesion, and expectedness bound | 7 → 8 | Establishes role-vectorized TD outcome surprise in the synthetic paradigm and raises the formal milestone; ecological validity and added depth remain the external-review ceiling |
| 2026-07-23 / `75a6488` audited oral-B-v2 figure | The strict renderer visualizes all 30 records, task-cluster learning, residualization, independent psychometrics, and acute lesions | 8 → 8 | Improves reviewability without adding evidence or inflating the score |
| 2026-07-26 / `cdc7d5e` oral-A-v2 scaling | All 60 preregistered ResNet-20/32/56 records pass; dynamic accuracy rises `91.584% → 92.254% → 92.760%`, all five depth pairs improve, and mean early alignment remains above `0.9994` | 8 → 9 | Establishes the internal standard-depth oral-A milestone. The conservative external recommendation is 8 because reciprocal KP is inherited and the full matched strong-baseline crossover is not yet complete |

Future rows are appended only after an audited frozen stage. A score staying flat is informative:
engineering, theory exposition, or visualization may make the paper more defensible without
resolving the empirical objection that determines the recommendation.