diff options
| -rw-r--r-- | REVIEW_SCORECARD.md | 52 |
1 files changed, 34 insertions, 18 deletions
diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md index 36da020..eded719 100644 --- a/REVIEW_SCORECARD.md +++ b/REVIEW_SCORECARD.md @@ -18,23 +18,28 @@ Every formal result report records: - the strongest remaining reviewer objections; - whether the evidence is development, validation, or untouched confirmation. -## Current paper-level assessment (2026-07-23) +## Current paper-level assessment (2026-07-27) | Dimension | Score | Strict reviewer assessment | |:--|--:|:--| | Soundness | 4/4 | Theory, exact local-update checks, causal lesions, clustered uncertainty, cost accounting, and retained frozen failures make the implemented claims unusually well identified. | | Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. | -| Significance | 3/4 | Dynamic innovation reaches near-BP ResNet-20 accuracy and a separate local actor--critic reproduces role-vectorized outcome surprise. Added-depth utility and original-data biological validation remain absent. | -| Empirical support | 4/4 | Five-depth preservation, residual necessity, untouched ResNet-20 confirmation, and a complete 30-record task-clustered BCI confirmation are strong. Useful-depth C2, broad endogenous C1, old oral-B, A3, and the original MT-1 remain disclosed failures. | +| Significance | 4/4 | Dynamic innovation now tracks BP/KP while gaining accuracy from ResNet-20 to ResNet-56 over five paired seeds; a separate local actor--critic reproduces role-vectorized outcome surprise. Original-data biological validation remains absent. | +| Empirical support | 4/4 | Residual necessity, untouched ResNet-20 confirmation, the complete 60-record standard-depth panel, and a complete 30-record task-clustered BCI confirmation are strong. Useful-depth C2, broad endogenous C1, old oral-B, A3, and the original MT-1 remain disclosed failures. | | Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. | -| **Overall** | **8/10** | **Internal accept milestone: untouched confirmations establish both load-bearing ResNet-20 innovation and role-vectorized TD outcome surprise under the declared synthetic BCI paradigm.** | +| **Overall** | **9/10** | **Internal oral-A milestone: the preregistered 60-record panel establishes standard ResNet depth scaling in addition to load-bearing innovation and role-vectorized TD outcome surprise.** | | Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. | -The conservative external-review forecast is **7/10**, not 8: a reviewer can -reasonably discount the synthetic BCI because terminal reward is supplied and -the psychometric target range is calibrated per policy. The 8/10 value is the -repository's predeclared evidence milestone; the external forecast is the -recommendation I would actually submit as a reviewer today. +The conservative external-review forecast is **8/10**, not 9. The +standard-depth result is unusually strong, but its clean-task transport +substrate is inherited reciprocal KP and the Harnett-specific gain remains a +conditional mixed-traffic result. A reviewer can also discount the synthetic +BCI because terminal reward is supplied and the psychometric target range is +calibrated per policy. The 9/10 value is the repository's predeclared +evidence milestone; 8/10 is the recommendation I would actually submit as a +reviewer today. The ongoing 81-cell BP/FA/DFA/PEPITA/FF/EP/DualProp/KP/SDIL +crossover targets the strongest remaining baseline-and-cost objection and +cannot change this score until its complete frozen panels are audited. ### Evidence already carrying the paper @@ -45,6 +50,11 @@ recommendation I would actually submit as a reviewer today. - On untouched CIFAR-10 test endpoints, dynamic innovation reaches `91.584%` versus clean KP's `91.388%` over five paired ResNet-20 seeds; its paired deficit upper bound is only `0.131` points and every frozen mechanism/cost invariant passes. +- In the complete 60-record standard-depth panel, dynamic innovation rises + from `91.584%` at ResNet-20 to `92.254%` at ResNet-32 and `92.760%` at + ResNet-56. All five paired seeds improve from depth 20 to 56; mean gain is + `1.176` points, while early-third alignment remains + `0.99969/0.99961/0.99942` across depths 20/32/56. - In the untouched six-task by five-model BCI panel, final task success is `100%`, terminal residual outcome decoding is `99.83%`, the acute outcome-lesion separation drop is `0.400`, critic expectedness is `0.319`, @@ -54,24 +64,28 @@ recommendation I would actually submit as a reviewer today. ### Strongest remaining objections -1. The depth result is preservation on a depth-flat task, not evidence that SDIL uses added depth. -2. The amortized apical vectorizer failed the frozen useful-depth gate; direct causal targets work - but cost `68.4x` ordinary forward-equivalent work. -3. Innovation was not uniformly beneficial for arbitrary endogenous top-down traffic, so the +1. The matched standard-depth panel contains BP, tuned DFA, clean KP, and + dynamic innovation, but not yet matched PEPITA, Forward-Forward, EP, Dual + Propagation, and ordinary FA at every depth. The frozen 81-cell crossover + is active but incomplete. +2. The clean scaling result primarily establishes the inherited reciprocal-KP + substrate; SDIL-specific necessity still comes from controlled predictable + traffic and cannot be inferred from clean KP/SDIL equality. +3. The amortized perturbation-trained apical vectorizer failed the frozen + useful-depth gate; the successful standard-depth path uses reciprocal local + plasticity and a paired neutral microphase. +4. Innovation was not uniformly beneficial for arbitrary endogenous top-down traffic, so the supported mechanism is narrower than the initial claim. -4. The passed BCI is synthetic: reward is directly supplied, causal roles are +5. The passed BCI is synthetic: reward is directly supplied, causal roles are experimenter-defined for diagnostics, and target quantiles are calibrated on a separate cursor split. Longitudinal prediction remains failed, and no original Francioni/Harnett event-level data are tested. -5. The standard-network result inherits reciprocal KP and pays for a paired neutral microphase. - D4 establishes the innovation operation, not a new credit-transport mechanism or a literal - cortical implementation. ## Score trajectory and prospective gates | Checkpoint | Overall | What changed | Remaining ceiling | |:--|--:|:--|:--| -| Current audited package | 8 | D4 confirms load-bearing ResNet-20 innovation; calibrated oral-B-v2 confirms role-vectorized TD outcome surprise over 30 untouched records with all cluster bounds passing | Added-depth utility and original-data biological validation remain absent | +| Current audited package | 9 | D4 confirms load-bearing ResNet-20 innovation; calibrated oral-B-v2 confirms role-vectorized TD outcome surprise; oral-A-v2 confirms five-seed ResNet-20/32/56 scaling over 60 records | Complete matched strong-baseline crossover and original-data biological validation remain absent | | Native baselines complete | 5 | BurstCCN is below its published endpoint; Dual Prop reproduces 92.46% versus 92.41%, with strict provenance and cost semantics | Fairness objection narrows, but SDIL gains no standard-scale evidence | | Oral-A A1/A2 | 5 | BP reached 91.62%; short channel-gated SDIL reached 41.98% versus tuned DFA at 37.16% | Development screening alone cannot raise the score | | Oral-A A3 fails | 5 | Full ResNet-20 SDIL became nonfinite at epoch 89 and ended at 10%; DFA ended finite at 33.06% | Standard-scale and oral-A claims are closed; A4 remains untouched | @@ -98,6 +112,7 @@ recommendation I would actually submit as a reviewer today. | Oral-B-v2 fixed-target recovery | failed at development | Two seeds pass 18/18; the third passes 17/18 but its fixed target ladder has 98.96% success | Mechanism works, absolute assay scale does not generalize | | Oral-B-v2 calibrated R1/R2 | 8 | Three fresh development seeds pass, then all 30 untouched records and every task-cluster bound pass under independent label-free calibration/evaluation splits | Establishes synthetic outcome surprise; does not establish cortex or added-depth utility | | Oral-A dynamic depth recovery | closed | A 60-cell ResNet-20/32/56 BP/DFA/clean-KP/dynamic panel was frozen before any new endpoint | Its oral-B R2 prerequisite failed, so none of the 50 new cells may run | +| Oral-A-v2 standard-depth scaling | 9 | After the calibrated BCI-v2 prerequisite passed, all 60 records and every preregistered accuracy, depth-gain, alignment, mechanism, cost, query, and memory check pass | Strong matched-baseline breadth and original neural-data validation remain absent | | Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available | These are conditional reviewer forecasts, not promised scores. A failed stage leaves its negative @@ -145,6 +160,7 @@ independently frozen oral-A-v2 protocol. | 2026-07-23 / `378e68d` fixed-target recovery R1 | Unit dense velocity restores 100% task learning and 17--18 biological checks per seed, but one fresh seed has 98.96% challenge success | 7 → 7 | Confirms the algorithmic repair while falsifying an absolute target ladder as a model-independent assay | | 2026-07-23 / `70e180c`, `9a8c057` calibrated oral-B-v2 R1/R2 | Three new development seeds pass all 18 gates; all 30 untouched confirmation records then pass every clustered learning, innovation, decoder, lesion, and expectedness bound | 7 → 8 | Establishes role-vectorized TD outcome surprise in the synthetic paradigm and raises the formal milestone; ecological validity and added depth remain the external-review ceiling | | 2026-07-23 / `75a6488` audited oral-B-v2 figure | The strict renderer visualizes all 30 records, task-cluster learning, residualization, independent psychometrics, and acute lesions | 8 → 8 | Improves reviewability without adding evidence or inflating the score | +| 2026-07-26 / `cdc7d5e` oral-A-v2 scaling | All 60 preregistered ResNet-20/32/56 records pass; dynamic accuracy rises `91.584% → 92.254% → 92.760%`, all five depth pairs improve, and mean early alignment remains above `0.9994` | 8 → 9 | Establishes the internal standard-depth oral-A milestone. The conservative external recommendation is 8 because reciprocal KP is inherited and the full matched strong-baseline crossover is not yet complete | Future rows are appended only after an audited frozen stage. A score staying flat is informative: engineering, theory exposition, or visualization may make the paper more defensible without |
