summaryrefslogtreecommitdiff
path: root/REVIEW_SCORECARD.md
diff options
context:
space:
mode:
authorYurenHao0426 <Blackhao0426@gmail.com>2026-07-22 19:45:46 -0500
committerYurenHao0426 <Blackhao0426@gmail.com>2026-07-22 19:45:46 -0500
commit0008f2cd96c06bea98f43560867d93967c597711 (patch)
tree9699171527b91b5468a15828ee501eeca4bf72cd /REVIEW_SCORECARD.md
parent15c60d05a7b6e173f6962d644ab23905d3441501 (diff)
docs: promote full dynamic ResNet result
Diffstat (limited to 'REVIEW_SCORECARD.md')
-rw-r--r--REVIEW_SCORECARD.md19
1 files changed, 9 insertions, 10 deletions
diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md
index 96272ab..0d03486 100644
--- a/REVIEW_SCORECARD.md
+++ b/REVIEW_SCORECARD.md
@@ -24,10 +24,10 @@ Every formal result report records:
|:--|--:|:--|
| Soundness | 3/4 | Theory, local-gradient checks, causal diagnostics, cost accounting, and frozen stop rules are unusually careful. The learned apical vectorizer remains an unresolved failure mode. |
| Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. |
-| Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, but the current flattened-CIFAR task does not benefit from depth and the frozen standard-ResNet recipe failed. |
-| Empirical support | 2/4 | Five-depth scaling and the residualization ablation are strong. Dynamic neutral projection now passes one standard-ResNet short validation gate, but useful-depth C2, broad endogenous C1, oral-B, full ResNet A3, and the original KP mixed-traffic MT-1 gates failed; full D3 and confirmation remain open. |
+| Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, and dynamic innovation now reaches near-BP accuracy on a standard ResNet-20. Added-depth utility and broader biological generality remain absent. |
+| Empirical support | 3/4 | Five-depth scaling, residual necessity, and a fully frozen 91.18% ResNet-20 validation endpoint are strong. D4 test confirmation is still running/open, while useful-depth C2, broad endogenous C1, oral-B, A3, and the original MT-1 failed. |
| Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. |
-| **Overall** | **5/10** | **Borderline reject: a strong core plus a promising short standard-scale result, but no completed full or independently confirmed endpoint.** |
+| **Overall** | **6/10** | **Borderline accept: the frozen full standard-ResNet endpoint resolves the primary scale objection; independent confirmation and broader novelty are still missing.** |
| Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. |
### Evidence already carrying the paper
@@ -47,16 +47,14 @@ Every formal result report records:
3. Innovation was not uniformly beneficial for arbitrary endogenous top-down traffic, so the
supported mechanism is narrower than the initial claim.
4. The frozen oral-B screen falsified the desired-velocity and Harnett error-derivative claims.
-5. The frozen ResNet-20 A3 run became nonfinite and ended at chance; the later KP mixed-traffic
- MT-1 panel also became nonfinite under all three signals. Dynamic neutral projection repairs a
- 20-epoch seed-0 validation run, but its full D3 endpoint and independent confirmation are not
- complete. A4, MT-2, and MT-3 remain sealed.
+5. The new full ResNet-20 evidence is one validation seed. The paired five-seed D4 test panel is
+ open but incomplete, so robustness and noninferiority to clean KP are not yet established.
## Score trajectory and prospective gates
| Checkpoint | Overall | What changed | Remaining ceiling |
|:--|--:|:--|:--|
-| Current audited package | 5 | Strong depth-preservation and residual-necessity evidence; negative gates retained | Standard useful scale is absent |
+| Current audited package | 6 | Strong depth-preservation/residual-necessity evidence plus a frozen 91.18% ResNet-20 validation endpoint; negative gates retained | Independent test confirmation and oral evidence remain absent |
| Native baselines complete | 5 | BurstCCN is below its published endpoint; Dual Prop reproduces 92.46% versus 92.41%, with strict provenance and cost semantics | Fairness objection narrows, but SDIL gains no standard-scale evidence |
| Oral-A A1/A2 | 5 | BP reached 91.62%; short channel-gated SDIL reached 41.98% versus tuned DFA at 37.16% | Development screening alone cannot raise the score |
| Oral-A A3 fails | 5 | Full ResNet-20 SDIL became nonfinite at epoch 89 and ended at 10%; DFA ended finite at 33.06% | Standard-scale and oral-A claims are closed; A4 remains untouched |
@@ -76,8 +74,8 @@ Every formal result report records:
| Mixed-traffic MT-3 | not opened | MT-2 was never opened; all five confirmation seeds remain untouched | No controlled ResNet innovation confirmation claim is available |
| Dynamic projection D1 | 5 | All 352 training-only steps remain finite while the fast neutral fit holds residual coupling near zero | Mechanics only; no held-out endpoint |
| Dynamic projection D2 | 5 | One frozen 20-epoch record reaches 83.58%, above clean KP's 82.66%, with all stability and cost gates passing | Short single-seed validation cannot establish the full scaling claim |
-| Dynamic projection D3 | running | A predeclared 200-epoch near-BP validation gate is open | Only a complete pass can make 6/10 eligible and open confirmation |
-| Dynamic projection D4 | sealed | Before observing D3, a paired clean-KP/dynamic five-seed test protocol and executable gate were frozen | D3 must pass first; a complete D4 pass, with no seed replacement, is required for 7/10 |
+| Dynamic projection D3 | 6 | All 19 frozen checks pass at 91.18%, within 0.44/0.08 points of BP/clean KP, with 0.9994 early alignment and 1.326x BP MACs | One validation seed cannot establish robustness |
+| Dynamic projection D4 | open | Before observing D3, a paired clean-KP/dynamic five-seed test protocol and executable gate were frozen | A complete D4 pass, with no seed replacement, is required for 7/10 |
| Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available |
These are conditional reviewer forecasts, not promised scores. A failed stage leaves its negative
@@ -111,6 +109,7 @@ retroactively reopened by success on standard vision benchmarks.
| 2026-07-22 / `8d28fd9` stability S0 | A frozen training-only four-margin grid has no eligible candidate; sign-certified margins either overflow BN state or produce enormous transient loss and parameter growth | 5 → 5 | Closes fixed one-sided predictor bias and motivates a two-sided dynamic gain condition, but no validation endpoint or empirical support is gained |
| 2026-07-22 / `6b85249` dynamic projection D1 | The sole frozen training-prefix record passes 14/14 checks; a 0.0058 neutral residual is reduced to 3.0e-8 while weights, momentum, BN, and loss stay bounded | 5 → 5 | Establishes a nontrivial operator-stability repair, but training-only mechanics cannot change the recommendation |
| 2026-07-22 / `9bc58d6` dynamic projection D2 | The sole frozen short validation record reaches 83.58% versus clean KP's 82.66% and failed controls' 10%, with 0.884 early alignment, zero queries, and 1.326x BP MACs | 5 → 5 | First positive standard-ResNet load-bearing innovation endpoint, but one short seed is insufficient; full D3 and independent confirmation remain decisive |
+| 2026-07-22 / `15c60d0` dynamic projection D3 | The sole frozen full validation record passes 19/19 checks at 91.18% versus BP's 91.62% and clean KP's 91.26%, with 0.9994 early alignment, zero queries, and 1.326x BP MACs | 5 → 6 | Resolves the primary full-scale objection and opens the already frozen D4 panel; one validation seed and narrow biological scope prevent a stronger recommendation |
Future rows are appended only after an audited frozen stage. A score staying flat is informative:
engineering, theory exposition, or visualization may make the paper more defensible without