summaryrefslogtreecommitdiff
path: root/REVIEW_SCORECARD.md
diff options
context:
space:
mode:
Diffstat (limited to 'REVIEW_SCORECARD.md')
-rw-r--r--REVIEW_SCORECARD.md41
1 files changed, 33 insertions, 8 deletions
diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md
index eded719..b2660bc 100644
--- a/REVIEW_SCORECARD.md
+++ b/REVIEW_SCORECARD.md
@@ -1,14 +1,39 @@
# ICLR-style reviewer scorecard
-This is a deliberately adversarial paper-level assessment, not an average of the best runs. It is
-updated only when a frozen experiment is complete and audited. Development pilots, partial jobs,
-and post-hoc subsets cannot raise the score. Failed gates remain visible and narrow the claim.
+This file starts with the current hardware-imperfection paper. The older
+KP/BCI clean-scaling assessment is retained below as project history. Scores
+use frozen, completed evidence; failed gates remain visible.
-## Scoring convention
+## Current three-part evidence assessment (2026-08-29)
-The overall score uses a 1--10 ICLR-like scale: 1 strong reject, 3 reject, 5 borderline reject,
-6 borderline accept, 8 accept, and 10 strong accept. Component scores use 1--4. The overall score
-is a reviewer judgment rather than the arithmetic mean of the components. Confidence uses 1--5.
+This score evaluates the evidence package described in
+`THREE_PART_EVIDENCE.md`. The existing `paper/MANUSCRIPT.md` still describes
+the earlier KP/BCI project and requires a full rewrite before submission.
+
+| Dimension | Score (1–5) | Confidence (1–5) | Evidence basis | Deduction / score-change condition |
+|:--|--:|--:|:--|:--|
+| Novelty | 3 | 4 | Harnett-inspired conditional innovation; transfer across Dual Propagation, EP, CLLN, and overclamping; distinction from AIMC update-asymmetry correction | The operation is simple and adjacent to autozero, residual-array, and dynamic-calibration methods. A sharper theorem or fabricated-hardware result would raise this dimension. |
+| Soundness | 4 | 4 | Conditional-projection identity, local quadratic bias result, exact locality audit, matched-noise and calibration controls | The largest-scale ladder uses a linearized digital CLLN; a second nonlinear large-scale substrate would close this gap. |
+| Evidence | 4 | 5 | Four digital learner forms; 3,600 core scaling trajectories; six-size overclamp confirmation; 120-pair nonlinear hardware simulation; retained EP failure | Fabricated-hardware SDIL training is absent, and the released ring tasks remain small classification problems. |
+| Significance | 4 | 4 | State-dependent local bias appears in released physical traces; SDIL prevents imperfection growth through 2,048 edges | Broader significance depends on showing the same problem and correction in a second physical workload or device family. |
+| Clarity | 3 | 4 | `THREE_PART_EVIDENCE.md` gives a compact three-part story and explicit costs | The current manuscript is obsolete. Rewriting it around teaching-signal imperfection is required before review. |
+| Reproducibility | 5 | 5 | Frozen protocols, exact seeds, source JSON/CSV, confirmation analyzers, failure records, and embedded-font figures are committed | Preserve source-to-claim links when moving into LaTeX. |
+| Ethics / limitations | 4 | 4 | Hardware simulation, descriptive real traces, cost tradeoffs, and the EP failure are labeled directly | Keep fabricated-hardware and energy claims outside the supported scope. |
+
+**Overall: 7/10 — weak accept for the evidence package. Scholarly confidence: 3/5.**
+
+The decisive positive is the combination of a simple local mechanism, broad
+additive transfer, and two independently confirmed six-size scaling results.
+The decisive limit is physical realism: real traces support the problem, while
+the SDIL intervention itself remains simulated. A closed-loop fabricated CLLN
+experiment, or a second nonlinear physical substrate with task-level scaling,
+would move the package toward 8/10. Failure of same-state neutral sampling at
+larger nonlinear scale would move it to 6/10.
+
+## Historical scoring convention
+
+The archived assessment below used a 1--10 overall scale and 1--4 component
+scores. These anchors apply only to the historical program.
Every formal result report records:
@@ -18,7 +43,7 @@ Every formal result report records:
- the strongest remaining reviewer objections;
- whether the evidence is development, validation, or untouched confirmation.
-## Current paper-level assessment (2026-07-27)
+## Historical clean-scaling assessment (2026-07-27)
| Dimension | Score | Strict reviewer assessment |
|:--|--:|:--|