diff options
| author | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-23 07:10:38 -0500 |
|---|---|---|
| committer | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-23 07:10:38 -0500 |
| commit | 81f32238ae1faf420c96bf48bc0507fe0c1f8fd9 (patch) | |
| tree | 07ae3386035acc452a6803ec4ac89923402582bc | |
| parent | 2a6f72e22a814fb4093174a2fb1ca2af56c7e63c (diff) | |
docs: log audited D4 figure without score inflation
| -rw-r--r-- | REVIEW_SCORECARD.md | 1 |
1 files changed, 1 insertions, 0 deletions
diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md index 483114c..d449a44 100644 --- a/REVIEW_SCORECARD.md +++ b/REVIEW_SCORECARD.md @@ -125,6 +125,7 @@ remains closed. | 2026-07-23 / `2fed62a` dynamic projection D4 | All ten untouched paired test records pass the frozen gate: dynamic 91.584% versus clean KP 91.388%, clean-minus-dynamic upper bound 0.131 points, 0.999687 early alignment, and zero invariant failures | 6 → 7 | Establishes the strict accept bar with independent ResNet-20 robustness/noninferiority; opens oral-B R1 but does not support added-depth or desired-velocity claims | | 2026-07-23 / `1ff7cb7` oral-B recovery R1 | The complete two-rate development grid selects eta 0.1; all 30 seed-level checks pass, with 98.05% worst-task final success, positive sign inversion, and 0.945 minimum role cosine | 7 → 7 | Strong development support for the repaired causal-role/velocity factorization, but R1 cannot change the score and only opens untouched R2 | | 2026-07-23 / `03c94a1` oral-B recovery R2 | Thirty untouched records retain 99.53% mean success, 90.45-point gain, 30/30 positive signs, and strong decorrelation, but fail seven population-vectorization/longitudinal checks | 7 → 7 | The recovery fixes learning, causal role, and sign but not the broader Harnett-like signature; oral-B and oral-A close without threshold repair | +| 2026-07-23 / `2a6f72e` audited D4 main figure | The strict renderer independently rechecks the ten D4 records and visualizes paired test accuracy, layerwise raw-versus-innovation direction, all 200 tracking epochs, and explicit neutral/MAC/memory/wall costs | 7 → 7 | Makes the accept evidence reviewable without adding or selecting data; presentation improves, but visualization alone cannot repair oral-B or justify score inflation | Future rows are appended only after an audited frozen stage. A score staying flat is informative: engineering, theory exposition, or visualization may make the paper more defensible without |
