summaryrefslogtreecommitdiff
path: root/REVIEW_SCORECARD.md
diff options
context:
space:
mode:
authorYurenHao0426 <Blackhao0426@gmail.com>2026-07-22 12:10:39 -0500
committerYurenHao0426 <Blackhao0426@gmail.com>2026-07-22 12:10:39 -0500
commit04de9326e60bb9e3592c0ddb5474265d228cbce0 (patch)
tree2066af6e8b164ea825934a416581f73dbaa13c58 /REVIEW_SCORECARD.md
parent09e8ccdf2404ba0691bffeab00cf223fcbd13222 (diff)
results: record failed Oral-A full ResNet gate
Diffstat (limited to 'REVIEW_SCORECARD.md')
-rw-r--r--REVIEW_SCORECARD.md13
1 files changed, 8 insertions, 5 deletions
diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md
index 1978688..25d1247 100644
--- a/REVIEW_SCORECARD.md
+++ b/REVIEW_SCORECARD.md
@@ -24,8 +24,8 @@ Every formal result report records:
|:--|--:|:--|
| Soundness | 3/4 | Theory, local-gradient checks, causal diagnostics, cost accounting, and frozen stop rules are unusually careful. The learned apical vectorizer remains an unresolved failure mode. |
| Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. |
-| Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, but the current flattened-CIFAR task does not benefit from depth. |
-| Empirical support | 2/4 | Five-depth scaling and the residualization ablation are strong. Useful-depth C2, broad endogenous C1, and oral-B gates failed; standard ResNet-scale confirmation is still pending. |
+| Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, but the current flattened-CIFAR task does not benefit from depth and the frozen standard-ResNet recipe failed. |
+| Empirical support | 2/4 | Five-depth scaling and the residualization ablation are strong. Useful-depth C2, broad endogenous C1, oral-B, and full ResNet A3 gates failed; the untouched A4 panel correctly remained sealed. |
| Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. |
| **Overall** | **5/10** | **Borderline reject: a strong core result without a completed standard-scale or useful-depth demonstration.** |
| Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. |
@@ -47,7 +47,8 @@ Every formal result report records:
3. Innovation was not uniformly beneficial for arbitrary endogenous top-down traffic, so the
supported mechanism is narrower than the initial claim.
4. The frozen oral-B screen falsified the desired-velocity and Harnett error-derivative claims.
-5. Native BurstCCN/Dual Prop endpoints and the standard CIFAR ResNet panel are not yet complete.
+5. The frozen ResNet-20 A3 run became nonfinite and ended at chance; native Dual Prop is the only
+ remaining incomplete baseline endpoint. The failed gate forbids the A4 test panel.
## Score trajectory and prospective gates
@@ -55,8 +56,9 @@ Every formal result report records:
|:--|--:|:--|:--|
| Current audited package | 5 | Strong depth-preservation and residual-necessity evidence; negative gates retained | Standard useful scale is absent |
| Native baselines complete | pending | Can close fairness/completeness objections, but cannot by itself establish the main claim | Usually no automatic score increase |
-| Oral-A A3 passes | pending | A full-data ResNet20 SDIL run is within 5 points of BP, beats DFA by 2, stays aligned, finite, and locally cost-audited | Plausible 6 if the evidence is clean |
-| Oral-A A4 passes | pending | Five untouched seeds across ResNet20/32/56 pass the frozen accuracy, scaling, alignment, and Pareto gates | Plausible 7--8; oral discussion becomes credible |
+| Oral-A A1/A2 | 5 | BP reached 91.62%; short channel-gated SDIL reached 41.98% versus tuned DFA at 37.16% | Development screening alone cannot raise the score |
+| Oral-A A3 fails | 5 | Full ResNet-20 SDIL became nonfinite at epoch 90 and ended at 10%; DFA ended finite at 33.06% | Standard-scale and oral-A claims are closed; A4 remains untouched |
+| Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available |
These are conditional reviewer forecasts, not promised scores. A failed stage leaves its negative
result in the record and can lower the score if it invalidates a current claim. Oral-B does not get
@@ -68,6 +70,7 @@ retroactively reopened by success on standard vision benchmarks.
|:--|:--|:--:|:--|
| 2026-07-22 / `2304e83` | Audit of all completed frozen branches | baseline → 5 | Strong preservation/residualization core, but no standard useful-scale result |
| 2026-07-22 / `c753f51`, `1b24c87`, `6d19078` | Existing 60-run innovation panel promoted to a strict main figure; conditional-projection and norm-direction identities made executable | 5 → 5 | Closes a presentation/theory objection and makes the narrow novelty legible, but adds no new held-out evidence and therefore earns no score inflation |
+| 2026-07-22 / frozen Oral-A A1--A3 | A1 and A2 pass; full A3 SDIL becomes nonfinite and fails four of six checks; A4 untouched | 5 → 5 | Closes the standard-scale question negatively. The narrow mechanism paper survives, while any standard-ResNet or oral claim does not |
Future rows are appended only after an audited frozen stage. A score staying flat is informative:
engineering, theory exposition, or visualization may make the paper more defensible without