summaryrefslogtreecommitdiff
path: root/REVIEW_SCORECARD.md
diff options
context:
space:
mode:
authorYurenHao0426 <Blackhao0426@gmail.com>2026-07-23 06:55:15 -0500
committerYurenHao0426 <Blackhao0426@gmail.com>2026-07-23 06:55:15 -0500
commit2b8b36b249e5d05386460c58aab268ca4d02eb0e (patch)
treecf92a8c627f760095dcf95fe994ad2c4aeaabaa2 /REVIEW_SCORECARD.md
parent2fed62a2486962e0054fbb9d38661ed76a2820f4 (diff)
docs: record accept-gate confirmation pass
Diffstat (limited to 'REVIEW_SCORECARD.md')
-rw-r--r--REVIEW_SCORECARD.md23
1 files changed, 14 insertions, 9 deletions
diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md
index 63f8d45..4efdf72 100644
--- a/REVIEW_SCORECARD.md
+++ b/REVIEW_SCORECARD.md
@@ -18,16 +18,16 @@ Every formal result report records:
- the strongest remaining reviewer objections;
- whether the evidence is development, validation, or untouched confirmation.
-## Current paper-level assessment (2026-07-22)
+## Current paper-level assessment (2026-07-23)
| Dimension | Score | Strict reviewer assessment |
|:--|--:|:--|
| Soundness | 3/4 | Theory, local-gradient checks, causal diagnostics, cost accounting, and frozen stop rules are unusually careful. The learned apical vectorizer remains an unresolved failure mode. |
| Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. |
| Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, and dynamic innovation now reaches near-BP accuracy on a standard ResNet-20. Added-depth utility and broader biological generality remain absent. |
-| Empirical support | 3/4 | Five-depth scaling, residual necessity, and a fully frozen 91.18% ResNet-20 validation endpoint are strong. D4 test confirmation is still running/open, while useful-depth C2, broad endogenous C1, oral-B, A3, and the original MT-1 failed. |
+| Empirical support | 4/4 | Five-depth scaling, residual necessity, the frozen 91.18% ResNet-20 validation endpoint, and an untouched five-seed paired test confirmation at 91.584% are strong. Useful-depth C2, broad endogenous C1, oral-B, A3, and the original MT-1 remain disclosed failures. |
| Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. |
-| **Overall** | **6/10** | **Borderline accept: the frozen full standard-ResNet endpoint resolves the primary scale objection; independent confirmation and broader novelty are still missing.** |
+| **Overall** | **7/10** | **Weak accept: untouched five-seed test confirmation establishes that dynamic innovation is robust and noninferior to strong clean KP on ResNet-20. Positive added-depth utility and the biological signature remain oral-level gaps.** |
| Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. |
### Evidence already carrying the paper
@@ -36,10 +36,13 @@ Every formal result report records:
depth 60, while DFA's early-layer alignment falls from `0.514` to `0.047`.
- Under strong soma-predictable apical traffic, raw and norm-matched raw learning fall to about
chance (`10.38%` and `10.31%`) while innovation learning retains `97.35%` accuracy.
+- On untouched CIFAR-10 test endpoints, dynamic innovation reaches `91.584%` versus clean KP's
+ `91.388%` over five paired ResNet-20 seeds; its paired deficit upper bound is only `0.131`
+ points and every frozen mechanism/cost invariant passes.
- The local update has a proved descent condition and an explicit query/MAC/memory audit; direct
node perturbation isolates the learned vectorizer as the useful-depth bottleneck.
-### Objections currently preventing acceptance confidence
+### Strongest remaining objections
1. The depth result is preservation on a depth-flat task, not evidence that SDIL uses added depth.
2. The amortized apical vectorizer failed the frozen useful-depth gate; direct causal targets work
@@ -48,14 +51,15 @@ Every formal result report records:
supported mechanism is narrower than the initial claim.
4. The frozen oral-B screen falsified the desired-velocity and Harnett error-derivative claims.
A temporal-difference factorization repairs the sign algebra only; it has no task endpoint.
-5. The new full ResNet-20 evidence is one validation seed. The paired five-seed D4 test panel is
- open but incomplete, so robustness and noninferiority to clean KP are not yet established.
+5. The standard-network result inherits reciprocal KP and pays for a paired neutral microphase.
+ D4 establishes the innovation operation, not a new credit-transport mechanism or a literal
+ cortical implementation.
## Score trajectory and prospective gates
| Checkpoint | Overall | What changed | Remaining ceiling |
|:--|--:|:--|:--|
-| Current audited package | 6 | Strong depth-preservation/residual-necessity evidence plus a frozen 91.18% ResNet-20 validation endpoint; negative gates retained | Independent test confirmation and oral evidence remain absent |
+| Current audited package | 7 | The untouched D4 panel reaches 91.584% dynamic versus 91.388% clean KP over five paired test seeds, with every mechanism and cost invariant passing | Added-depth utility and oral-B biology remain absent |
| Native baselines complete | 5 | BurstCCN is below its published endpoint; Dual Prop reproduces 92.46% versus 92.41%, with strict provenance and cost semantics | Fairness objection narrows, but SDIL gains no standard-scale evidence |
| Oral-A A1/A2 | 5 | BP reached 91.62%; short channel-gated SDIL reached 41.98% versus tuned DFA at 37.16% | Development screening alone cannot raise the score |
| Oral-A A3 fails | 5 | Full ResNet-20 SDIL became nonfinite at epoch 89 and ended at 10%; DFA ended finite at 33.06% | Standard-scale and oral-A claims are closed; A4 remains untouched |
@@ -76,8 +80,8 @@ Every formal result report records:
| Dynamic projection D1 | 5 | All 352 training-only steps remain finite while the fast neutral fit holds residual coupling near zero | Mechanics only; no held-out endpoint |
| Dynamic projection D2 | 5 | One frozen 20-epoch record reaches 83.58%, above clean KP's 82.66%, with all stability and cost gates passing | Short single-seed validation cannot establish the full scaling claim |
| Dynamic projection D3 | 6 | All 19 frozen checks pass at 91.18%, within 0.44/0.08 points of BP/clean KP, with 0.9994 early alignment and 1.326x BP MACs | One validation seed cannot establish robustness |
-| Dynamic projection D4 | open | Before observing D3, a paired clean-KP/dynamic five-seed test protocol and executable gate were frozen | A complete D4 pass, with no seed replacement, is required for 7/10 |
-| Oral-B recovery R1/R2 | not opened | Role/velocity factorization passes mechanics; the two-rate R1 and untouched 6-task-by-5-model R2 gates are frozen before any recovery task endpoint | R1 remains hard-gated on D4; only a complete R2 pass can establish innovation-guided plasticity and move 7 to 8 |
+| Dynamic projection D4 | 7 | All ten untouched records pass: dynamic 91.584% versus clean KP 91.388%, paired upper deficit bound 0.131 points, early alignment 0.999687, and no invariant failures | Establishes ResNet-20 robustness/noninferiority, not positive depth utility |
+| Oral-B recovery R1/R2 | R1 open | Role/velocity factorization passes mechanics; the two-rate R1 and untouched 6-task-by-5-model R2 gates were frozen before any recovery task endpoint | Only a complete R2 pass can establish innovation-guided plasticity and move 7 to 8 |
| Oral-A dynamic depth recovery | not opened | A 60-cell ResNet-20/32/56 BP/DFA/clean-KP/dynamic panel, positive-depth-benefit gate, mechanism invariants, and fair cost bounds are frozen before any new endpoint | It is hard-gated on D4 and oral-B R2; only a complete pass can establish standard-depth scaling and move 8 to 9 |
| Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available |
@@ -116,6 +120,7 @@ plasticity-only recovery must pass its own separately frozen R1 and R2 gates.
| 2026-07-22 / `15c60d0` dynamic projection D3 | The sole frozen full validation record passes 19/19 checks at 91.18% versus BP's 91.62% and clean KP's 91.26%, with 0.9994 early alignment, zero queries, and 1.326x BP MACs | 5 → 6 | Resolves the primary full-scale objection and opens the already frozen D4 panel; one validation seed and narrow biological scope prevent a stronger recommendation |
| 2026-07-22 / `32122d0` oral-B recovery freeze | Before any recovery task endpoint, the untouched 30-record R2 protocol, task-cluster uncertainty, original B1/B2 signatures, plasticity lesion, digest binding, and immutable analyzer are executable | 6 → 6 | Removes a preregistration gap but supplies no empirical evidence; R1 is still sealed behind D4 and confirmation seeds remain untouched |
| 2026-07-22 / `1853620` oral-A recovery freeze | Before any new standard-depth endpoint, a D4-reusing 60-cell ResNet-20/32/56 panel and exact runner/analyzer contract are executable, with BP/DFA/KP controls, positive depth gain, alignment, projection, query, MAC, and memory gates | 6 → 6 | Precommits the final oral claim without bypassing the required sequence; no experiment opens unless D4 reaches 7 and oral-B R2 reaches 8 |
+| 2026-07-23 / `2fed62a` dynamic projection D4 | All ten untouched paired test records pass the frozen gate: dynamic 91.584% versus clean KP 91.388%, clean-minus-dynamic upper bound 0.131 points, 0.999687 early alignment, and zero invariant failures | 6 → 7 | Establishes the strict accept bar with independent ResNet-20 robustness/noninferiority; opens oral-B R1 but does not support added-depth or desired-velocity claims |
Future rows are appended only after an audited frozen stage. A score staying flat is informative:
engineering, theory exposition, or visualization may make the paper more defensible without