diff options
| author | YurenHao0426 <Blackhao0426@gmail.com> | 2026-08-29 21:18:28 -0500 |
|---|---|---|
| committer | YurenHao0426 <Blackhao0426@gmail.com> | 2026-08-29 21:18:28 -0500 |
| commit | f4842c5fb1f71377572c887e0563ddd631cf92ad (patch) | |
| tree | 136d59b990c062707ed74f0295b1380aaf7cabad /REVIEW_SCORECARD.md | |
| parent | a1e684f277b85eacea18ab1dc6a59e4f528e14d2 (diff) | |
docs: audit the three-part evidence and reset reviewer score
Diffstat (limited to 'REVIEW_SCORECARD.md')
| -rw-r--r-- | REVIEW_SCORECARD.md | 41 |
1 files changed, 33 insertions, 8 deletions
diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md index eded719..b2660bc 100644 --- a/REVIEW_SCORECARD.md +++ b/REVIEW_SCORECARD.md @@ -1,14 +1,39 @@ # ICLR-style reviewer scorecard -This is a deliberately adversarial paper-level assessment, not an average of the best runs. It is -updated only when a frozen experiment is complete and audited. Development pilots, partial jobs, -and post-hoc subsets cannot raise the score. Failed gates remain visible and narrow the claim. +This file starts with the current hardware-imperfection paper. The older +KP/BCI clean-scaling assessment is retained below as project history. Scores +use frozen, completed evidence; failed gates remain visible. -## Scoring convention +## Current three-part evidence assessment (2026-08-29) -The overall score uses a 1--10 ICLR-like scale: 1 strong reject, 3 reject, 5 borderline reject, -6 borderline accept, 8 accept, and 10 strong accept. Component scores use 1--4. The overall score -is a reviewer judgment rather than the arithmetic mean of the components. Confidence uses 1--5. +This score evaluates the evidence package described in +`THREE_PART_EVIDENCE.md`. The existing `paper/MANUSCRIPT.md` still describes +the earlier KP/BCI project and requires a full rewrite before submission. + +| Dimension | Score (1–5) | Confidence (1–5) | Evidence basis | Deduction / score-change condition | +|:--|--:|--:|:--|:--| +| Novelty | 3 | 4 | Harnett-inspired conditional innovation; transfer across Dual Propagation, EP, CLLN, and overclamping; distinction from AIMC update-asymmetry correction | The operation is simple and adjacent to autozero, residual-array, and dynamic-calibration methods. A sharper theorem or fabricated-hardware result would raise this dimension. | +| Soundness | 4 | 4 | Conditional-projection identity, local quadratic bias result, exact locality audit, matched-noise and calibration controls | The largest-scale ladder uses a linearized digital CLLN; a second nonlinear large-scale substrate would close this gap. | +| Evidence | 4 | 5 | Four digital learner forms; 3,600 core scaling trajectories; six-size overclamp confirmation; 120-pair nonlinear hardware simulation; retained EP failure | Fabricated-hardware SDIL training is absent, and the released ring tasks remain small classification problems. | +| Significance | 4 | 4 | State-dependent local bias appears in released physical traces; SDIL prevents imperfection growth through 2,048 edges | Broader significance depends on showing the same problem and correction in a second physical workload or device family. | +| Clarity | 3 | 4 | `THREE_PART_EVIDENCE.md` gives a compact three-part story and explicit costs | The current manuscript is obsolete. Rewriting it around teaching-signal imperfection is required before review. | +| Reproducibility | 5 | 5 | Frozen protocols, exact seeds, source JSON/CSV, confirmation analyzers, failure records, and embedded-font figures are committed | Preserve source-to-claim links when moving into LaTeX. | +| Ethics / limitations | 4 | 4 | Hardware simulation, descriptive real traces, cost tradeoffs, and the EP failure are labeled directly | Keep fabricated-hardware and energy claims outside the supported scope. | + +**Overall: 7/10 — weak accept for the evidence package. Scholarly confidence: 3/5.** + +The decisive positive is the combination of a simple local mechanism, broad +additive transfer, and two independently confirmed six-size scaling results. +The decisive limit is physical realism: real traces support the problem, while +the SDIL intervention itself remains simulated. A closed-loop fabricated CLLN +experiment, or a second nonlinear physical substrate with task-level scaling, +would move the package toward 8/10. Failure of same-state neutral sampling at +larger nonlinear scale would move it to 6/10. + +## Historical scoring convention + +The archived assessment below used a 1--10 overall scale and 1--4 component +scores. These anchors apply only to the historical program. Every formal result report records: @@ -18,7 +43,7 @@ Every formal result report records: - the strongest remaining reviewer objections; - whether the evidence is development, validation, or untouched confirmation. -## Current paper-level assessment (2026-07-27) +## Historical clean-scaling assessment (2026-07-27) | Dimension | Score | Strict reviewer assessment | |:--|--:|:--| |
