diff options
| author | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-22 19:45:46 -0500 |
|---|---|---|
| committer | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-22 19:45:46 -0500 |
| commit | 0008f2cd96c06bea98f43560867d93967c597711 (patch) | |
| tree | 9699171527b91b5468a15828ee501eeca4bf72cd | |
| parent | 15c60d05a7b6e173f6962d644ab23905d3441501 (diff) | |
docs: promote full dynamic ResNet result
| -rw-r--r-- | DYNAMIC_INNOVATION.md | 15 | ||||
| -rw-r--r-- | PAPER_PLAN.md | 13 | ||||
| -rw-r--r-- | README.md | 8 | ||||
| -rw-r--r-- | RESULTS.md | 20 | ||||
| -rw-r--r-- | REVIEW_SCORECARD.md | 19 | ||||
| -rw-r--r-- | ROADMAP.md | 8 | ||||
| -rwxr-xr-x | experiments/finalize_accept.sh | 2 |
7 files changed, 66 insertions, 19 deletions
diff --git a/DYNAMIC_INNOVATION.md b/DYNAMIC_INNOVATION.md index d12dda3..2a22ea9 100644 --- a/DYNAMIC_INNOVATION.md +++ b/DYNAMIC_INNOVATION.md @@ -167,6 +167,21 @@ would then open a separately frozen independent multi-seed test confirmation; no confirmation seed or test endpoint may be touched before that protocol is committed. A D3 failure leaves the score at 5 and closes this branch. +## Audited D3 outcome + +D3 passes all 19 frozen checks at source revision `d945c42`. The sole +endpoint reaches `91.18%` validation accuracy, only 0.44 points below BP +(`91.62%`) and 0.08 points below clean KP (`91.26%`). Early-third alignment +is `0.999353`, final feedback/forward cosine is `0.999663`, and the epoch +151--200 mean is also `0.999663`. The maximum unprojected neutral/traffic +ratio is `0.008474`; the projected ratio and soma slope remain below +`3.04e-8` and `2.13e-8`, with zero task-instruction observations in the fast +fit. The run uses zero task-loss queries, `1.3261x` BP MACs plus separately +reported elementwise work, 2.165 GB peak allocated memory, one validation +evaluation, and zero test evaluations. Under the predeclared score rule this +raises the strict reviewer score from 5 to 6 and opens, but does not complete, +D4. + ## D4: conditionally frozen paired test confirmation The D4 protocol and executable gate are committed before any D3 endpoint is diff --git a/PAPER_PLAN.md b/PAPER_PLAN.md index 45ff8e5..e6c2465 100644 --- a/PAPER_PLAN.md +++ b/PAPER_PLAN.md @@ -14,8 +14,9 @@ direct feedback in the audited regime. This sentence deliberately says neither that cortex implements backpropagation nor that the current model explains Harnett's temporal-error -signature. “Scales on standard residual networks” may be added only after the -untouched Oral-A confirmation passes. +signature. The frozen D3 validation endpoint supports a near-BP standard +ResNet-20 result, but “scales on standard residual networks” may be added only +after the untouched multi-seed/depth confirmations pass. ## Working title @@ -74,10 +75,10 @@ not imply that no-traffic scaling establishes the necessity of residualization. vectorization, the Harnett desired-velocity signature, and full standard ResNet stability fail their frozen gates, defining the method's actual boundary. -5. **Standard scale.** Do not claim it yet: A3 failed and the dynamic - projection branch currently has only one short validation endpoint. A - predeclared full D3 result plus the already frozen paired five-seed D4 - confirmation is required. +5. **Standard scale.** D3 passes all frozen checks at `91.18%` validation on + ResNet-20 versus BP's `91.62%`, but it is one seed. Do not promote this to + a robustness or depth-scaling claim until the already frozen paired + five-seed D4 confirmation and a separate depth panel pass. ## Figure order in the manuscript @@ -71,9 +71,11 @@ and scaling behavior. See `NOVELTY.md` for the exact prior-art boundary. 352-step trajectory is finite, and D2 reaches `83.58%` after 20 epochs versus clean KP's `82.66%` and the failed mixed-traffic controls' `10%`, at `1.3261x` BP MACs and zero task-loss queries. This is still one short validation seed; - the predeclared 200-epoch D3 run is in progress and test remains untouched. - A paired five-seed clean-KP/dynamic D4 protocol is already frozen but remains - hard-gated on D3; no D4 test endpoint has been accessed. + the predeclared 200-epoch D3 run subsequently passes all 19 gates at `91.18%` + versus BP's `91.62%` and clean KP's `91.26%`, with `0.9994` early alignment, + zero task-loss queries, and `1.3261x` BP MACs. This raises the strict score + to 6/10. A paired five-seed clean-KP/dynamic D4 protocol is frozen and now + open; no D4 test endpoint had been accessed at the D3 audit. - Native author-code fidelity is complete. BurstCCN reaches `80.10%` at its validation-selected epoch versus published `82.97 +/- 0.21%`; Dual Prop reaches `92.46%` versus published `92.41 +/- 0.07%`. Their audited walls are @@ -995,6 +995,26 @@ innovation result on the standard ResNet substrate, but it is still a one-seed short validation record. Per the frozen scoring rule it leaves the reviewer score at 5/10 and opens only the predeclared D3 full validation run. +The predeclared D3 full validation record then passes all 19 frozen gates: + +| method | validation accuracy | early alignment | epoch-151--200 feedback cosine | MAC ratio to BP | +|:--|--:|--:|--:|--:| +| dynamic neutral-projection innovation | **91.18%** | 0.999353 | 0.999663 | 1.3261x | +| clean KP | 91.26% | 0.9994 | 0.9997 | 1.326x | +| matched BP | 91.62% | exact | exact | 1.000x | + +D3 remains finite for 200 epochs and ends with validation loss `0.425251`. +The maximum unprojected neutral/traffic ratio is `0.008474`, whereas the +projected remainder and absolute soma slope stay below `3.04e-8` and +`2.13e-8`; used-signal/instruction RMS differs by at most `6.73e-10`. It uses +zero task-loss queries, 1.452e15 MACs, 4.070e13 separately reported elementwise +operations, one neutral observation per ordinary example, 2.165 GB peak +allocated memory, one final validation evaluation, and zero test evaluations. +The source revision is `d945c42`, and the gate was frozen before the endpoint. +This is the first full near-BP standard-ResNet SDIL validation result and raises +the strict reviewer score from 5 to 6. Independent five-seed paired test +confirmation remains open rather than assumed. + ## How to run `experiments/run.py --mode {bp,fa,dfa,sdil} --dataset {mnist,fmnist,cifar10} --depth D --residual {0,1} --act {tanh,gelu,silu,relu}` Batteries: `experiments/run_v2.sh <ds> "<depths>" <res> <act> "<seeds>" <ep> <pfx>`. diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md index 96272ab..0d03486 100644 --- a/REVIEW_SCORECARD.md +++ b/REVIEW_SCORECARD.md @@ -24,10 +24,10 @@ Every formal result report records: |:--|--:|:--| | Soundness | 3/4 | Theory, local-gradient checks, causal diagnostics, cost accounting, and frozen stop rules are unusually careful. The learned apical vectorizer remains an unresolved failure mode. | | Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. | -| Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, but the current flattened-CIFAR task does not benefit from depth and the frozen standard-ResNet recipe failed. | -| Empirical support | 2/4 | Five-depth scaling and the residualization ablation are strong. Dynamic neutral projection now passes one standard-ResNet short validation gate, but useful-depth C2, broad endogenous C1, oral-B, full ResNet A3, and the original KP mixed-traffic MT-1 gates failed; full D3 and confirmation remain open. | +| Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, and dynamic innovation now reaches near-BP accuracy on a standard ResNet-20. Added-depth utility and broader biological generality remain absent. | +| Empirical support | 3/4 | Five-depth scaling, residual necessity, and a fully frozen 91.18% ResNet-20 validation endpoint are strong. D4 test confirmation is still running/open, while useful-depth C2, broad endogenous C1, oral-B, A3, and the original MT-1 failed. | | Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. | -| **Overall** | **5/10** | **Borderline reject: a strong core plus a promising short standard-scale result, but no completed full or independently confirmed endpoint.** | +| **Overall** | **6/10** | **Borderline accept: the frozen full standard-ResNet endpoint resolves the primary scale objection; independent confirmation and broader novelty are still missing.** | | Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. | ### Evidence already carrying the paper @@ -47,16 +47,14 @@ Every formal result report records: 3. Innovation was not uniformly beneficial for arbitrary endogenous top-down traffic, so the supported mechanism is narrower than the initial claim. 4. The frozen oral-B screen falsified the desired-velocity and Harnett error-derivative claims. -5. The frozen ResNet-20 A3 run became nonfinite and ended at chance; the later KP mixed-traffic - MT-1 panel also became nonfinite under all three signals. Dynamic neutral projection repairs a - 20-epoch seed-0 validation run, but its full D3 endpoint and independent confirmation are not - complete. A4, MT-2, and MT-3 remain sealed. +5. The new full ResNet-20 evidence is one validation seed. The paired five-seed D4 test panel is + open but incomplete, so robustness and noninferiority to clean KP are not yet established. ## Score trajectory and prospective gates | Checkpoint | Overall | What changed | Remaining ceiling | |:--|--:|:--|:--| -| Current audited package | 5 | Strong depth-preservation and residual-necessity evidence; negative gates retained | Standard useful scale is absent | +| Current audited package | 6 | Strong depth-preservation/residual-necessity evidence plus a frozen 91.18% ResNet-20 validation endpoint; negative gates retained | Independent test confirmation and oral evidence remain absent | | Native baselines complete | 5 | BurstCCN is below its published endpoint; Dual Prop reproduces 92.46% versus 92.41%, with strict provenance and cost semantics | Fairness objection narrows, but SDIL gains no standard-scale evidence | | Oral-A A1/A2 | 5 | BP reached 91.62%; short channel-gated SDIL reached 41.98% versus tuned DFA at 37.16% | Development screening alone cannot raise the score | | Oral-A A3 fails | 5 | Full ResNet-20 SDIL became nonfinite at epoch 89 and ended at 10%; DFA ended finite at 33.06% | Standard-scale and oral-A claims are closed; A4 remains untouched | @@ -76,8 +74,8 @@ Every formal result report records: | Mixed-traffic MT-3 | not opened | MT-2 was never opened; all five confirmation seeds remain untouched | No controlled ResNet innovation confirmation claim is available | | Dynamic projection D1 | 5 | All 352 training-only steps remain finite while the fast neutral fit holds residual coupling near zero | Mechanics only; no held-out endpoint | | Dynamic projection D2 | 5 | One frozen 20-epoch record reaches 83.58%, above clean KP's 82.66%, with all stability and cost gates passing | Short single-seed validation cannot establish the full scaling claim | -| Dynamic projection D3 | running | A predeclared 200-epoch near-BP validation gate is open | Only a complete pass can make 6/10 eligible and open confirmation | -| Dynamic projection D4 | sealed | Before observing D3, a paired clean-KP/dynamic five-seed test protocol and executable gate were frozen | D3 must pass first; a complete D4 pass, with no seed replacement, is required for 7/10 | +| Dynamic projection D3 | 6 | All 19 frozen checks pass at 91.18%, within 0.44/0.08 points of BP/clean KP, with 0.9994 early alignment and 1.326x BP MACs | One validation seed cannot establish robustness | +| Dynamic projection D4 | open | Before observing D3, a paired clean-KP/dynamic five-seed test protocol and executable gate were frozen | A complete D4 pass, with no seed replacement, is required for 7/10 | | Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available | These are conditional reviewer forecasts, not promised scores. A failed stage leaves its negative @@ -111,6 +109,7 @@ retroactively reopened by success on standard vision benchmarks. | 2026-07-22 / `8d28fd9` stability S0 | A frozen training-only four-margin grid has no eligible candidate; sign-certified margins either overflow BN state or produce enormous transient loss and parameter growth | 5 → 5 | Closes fixed one-sided predictor bias and motivates a two-sided dynamic gain condition, but no validation endpoint or empirical support is gained | | 2026-07-22 / `6b85249` dynamic projection D1 | The sole frozen training-prefix record passes 14/14 checks; a 0.0058 neutral residual is reduced to 3.0e-8 while weights, momentum, BN, and loss stay bounded | 5 → 5 | Establishes a nontrivial operator-stability repair, but training-only mechanics cannot change the recommendation | | 2026-07-22 / `9bc58d6` dynamic projection D2 | The sole frozen short validation record reaches 83.58% versus clean KP's 82.66% and failed controls' 10%, with 0.884 early alignment, zero queries, and 1.326x BP MACs | 5 → 5 | First positive standard-ResNet load-bearing innovation endpoint, but one short seed is insufficient; full D3 and independent confirmation remain decisive | +| 2026-07-22 / `15c60d0` dynamic projection D3 | The sole frozen full validation record passes 19/19 checks at 91.18% versus BP's 91.62% and clean KP's 91.26%, with 0.9994 early alignment, zero queries, and 1.326x BP MACs | 5 → 6 | Resolves the primary full-scale objection and opens the already frozen D4 panel; one validation seed and narrow biological scope prevent a stronger recommendation | Future rows are appended only after an audited frozen stage. A score staying flat is informative: engineering, theory exposition, or visualization may make the paper more defensible without @@ -538,6 +538,14 @@ hard-gated on a complete D3 pass, evaluates test exactly once per run, and can raise 6-to-7 only through its predeclared accuracy, paired noninferiority, alignment, leakage, query, and cost checks. +**Dynamic neutral projection D3: passed.** The sole frozen 200-epoch record +passes all 19 gates at source revision `d945c42`: `91.18%` validation versus +BP's `91.62%` and clean KP's `91.26%`, `0.999353` early alignment, `0.999663` +late feedback cosine, zero task-loss queries, and `1.3261x` BP MACs. Exactly +one validation and zero test evaluations occur. Per the frozen rule, the +strict reviewer score rises from 5 to 6 and the already committed D4 paired +five-seed test panel is now open. + Prepare convolutional local-update primitives and ResNet-20/32/56 protocols early. Queue frozen runs opportunistically on authorized idle GPUs. Because BurstCCN already reports CIFAR-10 and ImageNet scaling, dataset scale alone is not novel. The oral-level target is a memorable joint diff --git a/experiments/finalize_accept.sh b/experiments/finalize_accept.sh index 96f9ae9..44ad848 100755 --- a/experiments/finalize_accept.sh +++ b/experiments/finalize_accept.sh @@ -14,6 +14,8 @@ experiments/finalize_claims.sh experiments/analyze_kp_dynamic_projection.py >/dev/null /home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \ experiments/analyze_kp_dynamic_projection_short.py >/dev/null +/home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \ + experiments/analyze_kp_dynamic_projection_full.py >/dev/null /home/yurenh2/miniconda3/bin/python \ experiments/plot_tracking_failure.py >/dev/null /home/yurenh2/miniconda3/bin/python \ |
