summaryrefslogtreecommitdiff
path: root/REVIEW_SCORECARD.md
diff options
context:
space:
mode:
authorYurenHao0426 <Blackhao0426@gmail.com>2026-07-22 16:31:15 -0500
committerYurenHao0426 <Blackhao0426@gmail.com>2026-07-22 16:31:15 -0500
commit524c0704029c4b831d9f19413995e8cff1508b8b (patch)
tree4df653547478723fa61df58b37af423bc3709583 /REVIEW_SCORECARD.md
parentf310ad55b76ab96de488cdb4f2d05861ed4616f2 (diff)
docs: close mixed-traffic path after short failure
Diffstat (limited to 'REVIEW_SCORECARD.md')
-rw-r--r--REVIEW_SCORECARD.md14
1 files changed, 8 insertions, 6 deletions
diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md
index 73fb424..bbc5bd3 100644
--- a/REVIEW_SCORECARD.md
+++ b/REVIEW_SCORECARD.md
@@ -25,7 +25,7 @@ Every formal result report records:
| Soundness | 3/4 | Theory, local-gradient checks, causal diagnostics, cost accounting, and frozen stop rules are unusually careful. The learned apical vectorizer remains an unresolved failure mode. |
| Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. |
| Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, but the current flattened-CIFAR task does not benefit from depth and the frozen standard-ResNet recipe failed. |
-| Empirical support | 2/4 | Five-depth scaling and the residualization ablation are strong. Useful-depth C2, broad endogenous C1, oral-B, and full ResNet A3 gates failed; the untouched A4 panel correctly remained sealed. |
+| Empirical support | 2/4 | Five-depth scaling and the residualization ablation are strong. Useful-depth C2, broad endogenous C1, oral-B, full ResNet A3, and KP mixed-traffic MT-1 gates failed; untouched confirmation panels correctly remained sealed. |
| Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. |
| **Overall** | **5/10** | **Borderline reject: a strong core result without a completed standard-scale or useful-depth demonstration.** |
| Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. |
@@ -47,8 +47,9 @@ Every formal result report records:
3. Innovation was not uniformly beneficial for arbitrary endogenous top-down traffic, so the
supported mechanism is narrower than the initial claim.
4. The frozen oral-B screen falsified the desired-velocity and Harnett error-derivative claims.
-5. The frozen ResNet-20 A3 run became nonfinite and ended at chance. The failed gate forbids the
- A4 test panel; completed native baselines improve fairness but do not supply SDIL scale evidence.
+5. The frozen ResNet-20 A3 run became nonfinite and ended at chance; the later KP mixed-traffic
+ MT-1 panel also became nonfinite under all three signals. A4, MT-2, and MT-3 remain sealed, so
+ completed native baselines improve fairness but do not supply SDIL scale evidence.
## Score trajectory and prospective gates
@@ -69,9 +70,9 @@ Every formal result report records:
| Residual response mirror full gate fails | 5 | RRM ends at 10% with NaN validation loss although endpoint Q/W cosine is 0.999998 | Closes intermittent mirroring and proves endpoint alignment is an inadequate trajectory certificate; KP remains the separate strong substrate |
| Modified KP short gate passes | 5 | KP reaches 82.66% versus BP's epoch-20 81.02%, with 0.886 early alignment and 1.326x BP MACs | Strongly solves the substrate at useful scale, but all credit belongs to inherited reciprocal plasticity until innovation is load-bearing |
| Modified KP full gate passes | 5 | KP reaches 91.26% versus matched BP's 91.62%, with 0.9994 early alignment, 0.9997 late feedback cosine, zero queries, and 1.326x BP MACs | Establishes a stable near-BP ResNet substrate and opens MT-1, but inherited KP evidence cannot raise the SDIL score |
-| Mixed-traffic MT-1 | 5 if passed | A frozen short raw/matched/innovation screen must preserve clean-KP utility and show accuracy plus same-state directional gains | One seed and 20 epochs can only open full validation |
-| Mixed-traffic MT-2 | 5 if passed | A full seed-0 panel must reach 88%, stay within 3 points of clean KP, and beat both controls under audited cost | A single development seed still cannot establish acceptance |
-| Mixed-traffic MT-3 | 6 if passed | Five untouched network/traffic draws must retain 88% mean test accuracy and positive confidence-bounded paired gains | Establishes the controlled ResNet innovation claim; natural traffic and broad depth/architecture scaling remain open |
+| Mixed-traffic MT-1 fails | 5 | Raw, norm-matched raw, and innovation all become nonfinite in epoch 1 and end at 10%; calibration and predictor warmup still pass | Closes the controlled standard-ResNet innovation path; MT-2/MT-3 remain untouched and no weaker traffic rescue is allowed |
+| Mixed-traffic MT-2 | not opened | The prerequisite MT-1 gate failed | No full seed-0 validation claim is available |
+| Mixed-traffic MT-3 | not opened | MT-2 was never opened; all five confirmation seeds remain untouched | No controlled ResNet innovation confirmation claim is available |
| Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available |
These are conditional reviewer forecasts, not promised scores. A failed stage leaves its negative
@@ -100,6 +101,7 @@ retroactively reopened by success on standard vision benchmarks.
| 2026-07-22 / modified KP-1 | One clean constant-LR ResNet record reaches 82.66%, above BP's epoch-20 81.02%, with 0.8856 early alignment, 0.8546 train-period feedback cosine, zero queries, and 1.326x BP MACs | 5 → 5 | Establishes a stable strong substrate under a gate that observes the training trajectory, but reciprocal Kolen--Pollack plasticity is prior art and adds no Harnett-specific evidence |
| 2026-07-22 / `2ef7f94` modified KP-2 | One clean full ResNet record reaches 91.26%, 0.36 points below matched BP, with 0.9994 early alignment, 0.9997 late feedback cosine, zero queries, and 1.326x BP MACs | 5 → 5 | Removes substrate stability as the immediate blocker and opens the load-bearing innovation experiment, but prior-art reciprocal plasticity earns no novelty credit |
| 2026-07-22 / mixed-traffic MT-0 | Exact raw/matched/innovation mechanics pass; synthetic ResNet-20 matched batch peaks at 0.857 GB allocated on GTX 1080 without task data | 5 → 5 | Removes implementation, graph, and memory objections before endpoint access; supplies no evidence yet that innovation is useful |
+| 2026-07-22 / `f310ad5` mixed-traffic MT-1 | All three frozen conditions become nonfinite in epoch 1 and end at 10%; ratio calibration, predictor warmup, zero-query, and cost checks pass | 5 → 5 | Fails to make innovation load-bearing on a standard ResNet and closes MT-2/MT-3; the prior mechanism-only evidence survives but empirical support does not improve |
Future rows are appended only after an audited frozen stage. A score staying flat is informative:
engineering, theory exposition, or visualization may make the paper more defensible without