summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
-rw-r--r--REVIEW_SCORECARD.md52
1 files changed, 34 insertions, 18 deletions
diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md
index 36da020..eded719 100644
--- a/REVIEW_SCORECARD.md
+++ b/REVIEW_SCORECARD.md
@@ -18,23 +18,28 @@ Every formal result report records:
- the strongest remaining reviewer objections;
- whether the evidence is development, validation, or untouched confirmation.
-## Current paper-level assessment (2026-07-23)
+## Current paper-level assessment (2026-07-27)
| Dimension | Score | Strict reviewer assessment |
|:--|--:|:--|
| Soundness | 4/4 | Theory, exact local-update checks, causal lesions, clustered uncertainty, cost accounting, and retained frozen failures make the implemented claims unusually well identified. |
| Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. |
-| Significance | 3/4 | Dynamic innovation reaches near-BP ResNet-20 accuracy and a separate local actor--critic reproduces role-vectorized outcome surprise. Added-depth utility and original-data biological validation remain absent. |
-| Empirical support | 4/4 | Five-depth preservation, residual necessity, untouched ResNet-20 confirmation, and a complete 30-record task-clustered BCI confirmation are strong. Useful-depth C2, broad endogenous C1, old oral-B, A3, and the original MT-1 remain disclosed failures. |
+| Significance | 4/4 | Dynamic innovation now tracks BP/KP while gaining accuracy from ResNet-20 to ResNet-56 over five paired seeds; a separate local actor--critic reproduces role-vectorized outcome surprise. Original-data biological validation remains absent. |
+| Empirical support | 4/4 | Residual necessity, untouched ResNet-20 confirmation, the complete 60-record standard-depth panel, and a complete 30-record task-clustered BCI confirmation are strong. Useful-depth C2, broad endogenous C1, old oral-B, A3, and the original MT-1 remain disclosed failures. |
| Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. |
-| **Overall** | **8/10** | **Internal accept milestone: untouched confirmations establish both load-bearing ResNet-20 innovation and role-vectorized TD outcome surprise under the declared synthetic BCI paradigm.** |
+| **Overall** | **9/10** | **Internal oral-A milestone: the preregistered 60-record panel establishes standard ResNet depth scaling in addition to load-bearing innovation and role-vectorized TD outcome surprise.** |
| Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. |
-The conservative external-review forecast is **7/10**, not 8: a reviewer can
-reasonably discount the synthetic BCI because terminal reward is supplied and
-the psychometric target range is calibrated per policy. The 8/10 value is the
-repository's predeclared evidence milestone; the external forecast is the
-recommendation I would actually submit as a reviewer today.
+The conservative external-review forecast is **8/10**, not 9. The
+standard-depth result is unusually strong, but its clean-task transport
+substrate is inherited reciprocal KP and the Harnett-specific gain remains a
+conditional mixed-traffic result. A reviewer can also discount the synthetic
+BCI because terminal reward is supplied and the psychometric target range is
+calibrated per policy. The 9/10 value is the repository's predeclared
+evidence milestone; 8/10 is the recommendation I would actually submit as a
+reviewer today. The ongoing 81-cell BP/FA/DFA/PEPITA/FF/EP/DualProp/KP/SDIL
+crossover targets the strongest remaining baseline-and-cost objection and
+cannot change this score until its complete frozen panels are audited.
### Evidence already carrying the paper
@@ -45,6 +50,11 @@ recommendation I would actually submit as a reviewer today.
- On untouched CIFAR-10 test endpoints, dynamic innovation reaches `91.584%` versus clean KP's
`91.388%` over five paired ResNet-20 seeds; its paired deficit upper bound is only `0.131`
points and every frozen mechanism/cost invariant passes.
+- In the complete 60-record standard-depth panel, dynamic innovation rises
+ from `91.584%` at ResNet-20 to `92.254%` at ResNet-32 and `92.760%` at
+ ResNet-56. All five paired seeds improve from depth 20 to 56; mean gain is
+ `1.176` points, while early-third alignment remains
+ `0.99969/0.99961/0.99942` across depths 20/32/56.
- In the untouched six-task by five-model BCI panel, final task success is
`100%`, terminal residual outcome decoding is `99.83%`, the acute
outcome-lesion separation drop is `0.400`, critic expectedness is `0.319`,
@@ -54,24 +64,28 @@ recommendation I would actually submit as a reviewer today.
### Strongest remaining objections
-1. The depth result is preservation on a depth-flat task, not evidence that SDIL uses added depth.
-2. The amortized apical vectorizer failed the frozen useful-depth gate; direct causal targets work
- but cost `68.4x` ordinary forward-equivalent work.
-3. Innovation was not uniformly beneficial for arbitrary endogenous top-down traffic, so the
+1. The matched standard-depth panel contains BP, tuned DFA, clean KP, and
+ dynamic innovation, but not yet matched PEPITA, Forward-Forward, EP, Dual
+ Propagation, and ordinary FA at every depth. The frozen 81-cell crossover
+ is active but incomplete.
+2. The clean scaling result primarily establishes the inherited reciprocal-KP
+ substrate; SDIL-specific necessity still comes from controlled predictable
+ traffic and cannot be inferred from clean KP/SDIL equality.
+3. The amortized perturbation-trained apical vectorizer failed the frozen
+ useful-depth gate; the successful standard-depth path uses reciprocal local
+ plasticity and a paired neutral microphase.
+4. Innovation was not uniformly beneficial for arbitrary endogenous top-down traffic, so the
supported mechanism is narrower than the initial claim.
-4. The passed BCI is synthetic: reward is directly supplied, causal roles are
+5. The passed BCI is synthetic: reward is directly supplied, causal roles are
experimenter-defined for diagnostics, and target quantiles are calibrated
on a separate cursor split. Longitudinal prediction remains failed, and no
original Francioni/Harnett event-level data are tested.
-5. The standard-network result inherits reciprocal KP and pays for a paired neutral microphase.
- D4 establishes the innovation operation, not a new credit-transport mechanism or a literal
- cortical implementation.
## Score trajectory and prospective gates
| Checkpoint | Overall | What changed | Remaining ceiling |
|:--|--:|:--|:--|
-| Current audited package | 8 | D4 confirms load-bearing ResNet-20 innovation; calibrated oral-B-v2 confirms role-vectorized TD outcome surprise over 30 untouched records with all cluster bounds passing | Added-depth utility and original-data biological validation remain absent |
+| Current audited package | 9 | D4 confirms load-bearing ResNet-20 innovation; calibrated oral-B-v2 confirms role-vectorized TD outcome surprise; oral-A-v2 confirms five-seed ResNet-20/32/56 scaling over 60 records | Complete matched strong-baseline crossover and original-data biological validation remain absent |
| Native baselines complete | 5 | BurstCCN is below its published endpoint; Dual Prop reproduces 92.46% versus 92.41%, with strict provenance and cost semantics | Fairness objection narrows, but SDIL gains no standard-scale evidence |
| Oral-A A1/A2 | 5 | BP reached 91.62%; short channel-gated SDIL reached 41.98% versus tuned DFA at 37.16% | Development screening alone cannot raise the score |
| Oral-A A3 fails | 5 | Full ResNet-20 SDIL became nonfinite at epoch 89 and ended at 10%; DFA ended finite at 33.06% | Standard-scale and oral-A claims are closed; A4 remains untouched |
@@ -98,6 +112,7 @@ recommendation I would actually submit as a reviewer today.
| Oral-B-v2 fixed-target recovery | failed at development | Two seeds pass 18/18; the third passes 17/18 but its fixed target ladder has 98.96% success | Mechanism works, absolute assay scale does not generalize |
| Oral-B-v2 calibrated R1/R2 | 8 | Three fresh development seeds pass, then all 30 untouched records and every task-cluster bound pass under independent label-free calibration/evaluation splits | Establishes synthetic outcome surprise; does not establish cortex or added-depth utility |
| Oral-A dynamic depth recovery | closed | A 60-cell ResNet-20/32/56 BP/DFA/clean-KP/dynamic panel was frozen before any new endpoint | Its oral-B R2 prerequisite failed, so none of the 50 new cells may run |
+| Oral-A-v2 standard-depth scaling | 9 | After the calibrated BCI-v2 prerequisite passed, all 60 records and every preregistered accuracy, depth-gain, alignment, mechanism, cost, query, and memory check pass | Strong matched-baseline breadth and original neural-data validation remain absent |
| Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available |
These are conditional reviewer forecasts, not promised scores. A failed stage leaves its negative
@@ -145,6 +160,7 @@ independently frozen oral-A-v2 protocol.
| 2026-07-23 / `378e68d` fixed-target recovery R1 | Unit dense velocity restores 100% task learning and 17--18 biological checks per seed, but one fresh seed has 98.96% challenge success | 7 → 7 | Confirms the algorithmic repair while falsifying an absolute target ladder as a model-independent assay |
| 2026-07-23 / `70e180c`, `9a8c057` calibrated oral-B-v2 R1/R2 | Three new development seeds pass all 18 gates; all 30 untouched confirmation records then pass every clustered learning, innovation, decoder, lesion, and expectedness bound | 7 → 8 | Establishes role-vectorized TD outcome surprise in the synthetic paradigm and raises the formal milestone; ecological validity and added depth remain the external-review ceiling |
| 2026-07-23 / `75a6488` audited oral-B-v2 figure | The strict renderer visualizes all 30 records, task-cluster learning, residualization, independent psychometrics, and acute lesions | 8 → 8 | Improves reviewability without adding evidence or inflating the score |
+| 2026-07-26 / `cdc7d5e` oral-A-v2 scaling | All 60 preregistered ResNet-20/32/56 records pass; dynamic accuracy rises `91.584% → 92.254% → 92.760%`, all five depth pairs improve, and mean early alignment remains above `0.9994` | 8 → 9 | Establishes the internal standard-depth oral-A milestone. The conservative external recommendation is 8 because reciprocal KP is inherited and the full matched strong-baseline crossover is not yet complete |
Future rows are appended only after an audited frozen stage. A score staying flat is informative:
engineering, theory exposition, or visualization may make the paper more defensible without