From f4842c5fb1f71377572c887e0563ddd631cf92ad Mon Sep 17 00:00:00 2001 From: YurenHao0426 Date: Sat, 29 Aug 2026 21:18:28 -0500 Subject: docs: audit the three-part evidence and reset reviewer score --- README.md | 10 ++++------ REVIEW_SCORECARD.md | 41 +++++++++++++++++++++++++++++++++-------- THREE_PART_EVIDENCE.md | 9 ++++++--- 3 files changed, 43 insertions(+), 17 deletions(-) diff --git a/README.md b/README.md index 4da092e..bb1a4c4 100644 --- a/README.md +++ b/README.md @@ -362,9 +362,7 @@ accounting, and final finiteness. Development, validation, and untouched confirmation results are never pooled. Failed gates close their branch instead of triggering seed deletion or post-hoc threshold changes. -The current formal milestone is `9/10` after the untouched D4, oral-B-v2, and -standard-depth confirmations. A conservative external-review forecast is -`8/10`: the full matched nine-method crossover, original-data biological -validation, and a credit pathway novel beyond inherited perturbation/KP -mechanisms remain open. Scores change only after an audited frozen stage, not -after a pilot or presentation improvement. +That previous program reached an internal `9/10` milestone under its own +frozen gates. It is retained as an audit trail and does not score the current +three-part hardware-imperfection paper. The current project-level assessment +is maintained in [`REVIEW_SCORECARD.md`](REVIEW_SCORECARD.md). diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md index eded719..b2660bc 100644 --- a/REVIEW_SCORECARD.md +++ b/REVIEW_SCORECARD.md @@ -1,14 +1,39 @@ # ICLR-style reviewer scorecard -This is a deliberately adversarial paper-level assessment, not an average of the best runs. It is -updated only when a frozen experiment is complete and audited. Development pilots, partial jobs, -and post-hoc subsets cannot raise the score. Failed gates remain visible and narrow the claim. +This file starts with the current hardware-imperfection paper. The older +KP/BCI clean-scaling assessment is retained below as project history. Scores +use frozen, completed evidence; failed gates remain visible. -## Scoring convention +## Current three-part evidence assessment (2026-08-29) -The overall score uses a 1--10 ICLR-like scale: 1 strong reject, 3 reject, 5 borderline reject, -6 borderline accept, 8 accept, and 10 strong accept. Component scores use 1--4. The overall score -is a reviewer judgment rather than the arithmetic mean of the components. Confidence uses 1--5. +This score evaluates the evidence package described in +`THREE_PART_EVIDENCE.md`. The existing `paper/MANUSCRIPT.md` still describes +the earlier KP/BCI project and requires a full rewrite before submission. + +| Dimension | Score (1–5) | Confidence (1–5) | Evidence basis | Deduction / score-change condition | +|:--|--:|--:|:--|:--| +| Novelty | 3 | 4 | Harnett-inspired conditional innovation; transfer across Dual Propagation, EP, CLLN, and overclamping; distinction from AIMC update-asymmetry correction | The operation is simple and adjacent to autozero, residual-array, and dynamic-calibration methods. A sharper theorem or fabricated-hardware result would raise this dimension. | +| Soundness | 4 | 4 | Conditional-projection identity, local quadratic bias result, exact locality audit, matched-noise and calibration controls | The largest-scale ladder uses a linearized digital CLLN; a second nonlinear large-scale substrate would close this gap. | +| Evidence | 4 | 5 | Four digital learner forms; 3,600 core scaling trajectories; six-size overclamp confirmation; 120-pair nonlinear hardware simulation; retained EP failure | Fabricated-hardware SDIL training is absent, and the released ring tasks remain small classification problems. | +| Significance | 4 | 4 | State-dependent local bias appears in released physical traces; SDIL prevents imperfection growth through 2,048 edges | Broader significance depends on showing the same problem and correction in a second physical workload or device family. | +| Clarity | 3 | 4 | `THREE_PART_EVIDENCE.md` gives a compact three-part story and explicit costs | The current manuscript is obsolete. Rewriting it around teaching-signal imperfection is required before review. | +| Reproducibility | 5 | 5 | Frozen protocols, exact seeds, source JSON/CSV, confirmation analyzers, failure records, and embedded-font figures are committed | Preserve source-to-claim links when moving into LaTeX. | +| Ethics / limitations | 4 | 4 | Hardware simulation, descriptive real traces, cost tradeoffs, and the EP failure are labeled directly | Keep fabricated-hardware and energy claims outside the supported scope. | + +**Overall: 7/10 — weak accept for the evidence package. Scholarly confidence: 3/5.** + +The decisive positive is the combination of a simple local mechanism, broad +additive transfer, and two independently confirmed six-size scaling results. +The decisive limit is physical realism: real traces support the problem, while +the SDIL intervention itself remains simulated. A closed-loop fabricated CLLN +experiment, or a second nonlinear physical substrate with task-level scaling, +would move the package toward 8/10. Failure of same-state neutral sampling at +larger nonlinear scale would move it to 6/10. + +## Historical scoring convention + +The archived assessment below used a 1--10 overall scale and 1--4 component +scores. These anchors apply only to the historical program. Every formal result report records: @@ -18,7 +43,7 @@ Every formal result report records: - the strongest remaining reviewer objections; - whether the evidence is development, validation, or untouched confirmation. -## Current paper-level assessment (2026-07-27) +## Historical clean-scaling assessment (2026-07-27) | Dimension | Score | Strict reviewer assessment | |:--|--:|:--| diff --git a/THREE_PART_EVIDENCE.md b/THREE_PART_EVIDENCE.md index b8785e4..e0a52c4 100644 --- a/THREE_PART_EVIDENCE.md +++ b/THREE_PART_EVIDENCE.md @@ -187,12 +187,15 @@ teaching rule itself is contrastive and local. 1. Method and transfer across Dual Propagation, EP, standard CLLN, and overclamped CLLN. -2. Digital CLLN scaling: final error, stable success, and learning curves. +2. Digital CLLN scaling: final error, error AUC, and stable failure. 3. Hardware-realistic CLLN: raw, calibration, overclamping, SDIL, and nonideal-CDS robustness. 4. Mechanism boundary: matched noise, state dependence, sampling mismatch, and refresh interval. -## Remaining gates +## Audit status -1. Audit every publication-facing number against its source JSON. +The publication-facing values above have been checked against their source +JSON and CSV files. All four main PDFs use embedded TrueType fonts. The EP +strict-gate failure remains visible in Part 1, and the real hardware traces in +Part 3 remain descriptive evidence of state dependence. -- cgit v1.2.3