diff options
| author | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-27 14:59:34 -0500 |
|---|---|---|
| committer | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-27 14:59:34 -0500 |
| commit | 94b77d8b76bead67fff9e91f24c7c981cf083b0c (patch) | |
| tree | 9ac8c667f96a57120864648d356dd4ea00e0a102 | |
| parent | 8e5473036e9db8d4a863dc707cb73a811bb162b1 (diff) | |
docs: align project status with depth pass
| -rw-r--r-- | PAPER_PLAN.md | 17 | ||||
| -rw-r--r-- | README.md | 60 | ||||
| -rw-r--r-- | RESULTS.md | 31 | ||||
| -rw-r--r-- | ROADMAP.md | 13 |
4 files changed, 92 insertions, 29 deletions
diff --git a/PAPER_PLAN.md b/PAPER_PLAN.md index ecb3bc8..ec52c74 100644 --- a/PAPER_PLAN.md +++ b/PAPER_PLAN.md @@ -200,8 +200,12 @@ should be reassessed from the frozen confirmation only. The original A3/A4 branch is permanently failed, so this paragraph can now apply only to the separately frozen dynamic-depth recovery. Its full 60-cell -pass is required before a standard-depth abstract/title claim or a 9/10 score; -D4 alone remains a single-architecture robustness result. +pass was required before a standard-depth abstract/title claim or a 9/10 +score; D4 alone remained a single-architecture robustness result. The +calibrated-BCI-gated oral-A-v2 panel subsequently passed all 60 record, +accuracy, paired depth-gain, alignment, mechanism, query, MAC, and memory +checks. The manuscript may therefore make the standard-depth claim, while +attributing clean-task transport substantially to inherited reciprocal KP. ### If A3 or A4 fails @@ -211,9 +215,12 @@ causal-feedback backbone preserves depth on the controlled task, but standard useful scaling remains unresolved. The correct response is a narrower title, claim, and score—not another post-hoc ResNet tuning branch. -This is the realized branch: A3 SDIL became nonfinite at epoch 89 and ended at -chance. A4 remained untouched. The working title therefore stays -mechanism-focused, and the standard-ResNet result is a disclosed limitation. +This is the realized outcome for the original A3/A4 branch: A3 SDIL became +nonfinite at epoch 89 and ended at chance, and A4 remained untouched. It is +retained as a failed direct-vectorizer path rather than treated as the final +standard-depth outcome. The independently authorized dynamic-innovation v2 +recovery later passed ResNet-20/32/56, so the remaining limitation is matched +baseline and architecture breadth rather than positive depth utility itself. ## Post-A3 accept recovery: innovation on a strong inherited substrate @@ -81,8 +81,8 @@ and scaling behavior. See `NOVELTY.md` for the exact prior-art boundary. Every dynamic seed reaches at least `91.51%`, mean early alignment is `0.999687`, and all trajectory, projection, leakage, query, hardware, MAC, memory, split, and test-isolation checks pass. This untouched confirmation - establishes the strict 7/10 accept bar; it does not establish that added - standard-network depth is useful. + establishes the strict 7/10 accept bar; D4 alone does not establish that + added standard-network depth is useful. - Native author-code fidelity is complete. BurstCCN reaches `80.10%` at its validation-selected epoch versus published `82.97 +/- 0.21%`; Dual Prop reaches `92.46%` versus published `92.41 +/- 0.07%`. Their audited walls are @@ -105,8 +105,8 @@ nevertheless fails: residual outcome decoding is only 47.33%, residuals trail soma by 4.32 points, and the longitudinal correlation is `-0.013`. Thus the recovery sharpens the boundary--causal-role temporal-difference plasticity works in this synthetic task, but the broader Harnett-like population -vectorization signature is not established. The strict score remains 7/10 -and the oral-A depth panel stays sealed. +vectorization signature is not established. At that stage the strict score +remained 7/10 and the old oral-A depth panel stayed sealed. An independent oral-B-v2 then adds an explicit terminal reward/timeout phase, a local linear TD critic, and temporal eligibility traces. Its first frozen @@ -122,6 +122,24 @@ drop in role separation, and 30/30 positive causal signs. This raises the formal milestone to 8/10 and permits only a separately frozen oral-A-v2 protocol; the old depth panel remains closed. +That separately frozen oral-A-v2 protocol is now complete. Its audited 60 +records cross BP, tuned DFA, clean reciprocal KP, and dynamic SDIL over +ResNet-20/32/56 with five paired seeds. Dynamic SDIL rises +`91.584% -> 92.254% -> 92.760%`; all five depth-20-to-56 pairs improve, with a +mean gain of `1.176` points. ResNet-56 SDIL reaches `92.760%` versus `92.632%` +BP, `92.670%` clean KP, and `30.850%` DFA while retaining `0.999423` +early-third alignment at `1.331x` the matched BP MAC estimate. This establishes +the internal 9/10 standard-depth milestone. It does not erase the failed old +branch or show that residualization, rather than the inherited reciprocal +backbone, causes clean-task scaling. + +The next matched crossover is mechanically registered as nine distinct +architecture/size points times nine methods: miniCNN/VGGlike/VGG16, +ResNet-20/32/56, and decoder-Transformer-4/8/12, each with BP, ordinary FA, +DFA, PEPITA, Forward--Forward, EP, Dual Propagation, clean KP, and SDIL. All 81 +cells are mandatory; adding a width, context, or depth point adds all nine +methods rather than an SDIL-only extension. + ## Publication-facing artifacts - `RESULTS.md`: audited positive and negative results; @@ -140,20 +158,26 @@ protocol; the old depth panel remains closed. - `ORAL_A.md`: frozen standard CIFAR ResNet funnel; - `ORAL_A_V2.md`: frozen post-failure representable-subspace funnel; - `ORAL_A_V3.md`: frozen vectorizer-space causal-calibration funnel; +- `ORAL_A_RECOVERY_V2.md`: passed calibrated-BCI-gated + ResNet-20/32/56 scaling protocol; +- `CROSS_ARCHITECTURE_CROSSOVER.md`, `RESNET_CROSSOVER.md`, and + `TRANSFORMER_CROSSOVER.md`: frozen 81-cell matched-crossover contracts; - `REVIEW_SCORECARD.md`: adversarial ICLR-style score trajectory; - `paper/MANUSCRIPT.md`: evidence-bound ICLR working draft; - `paper/CLAIM_LEDGER.json` and `paper/manuscript_audit.json`: direct bindings from manuscript numbers, gate statuses, figures, and claim boundaries to their audited source files; - `results/figs/`: deterministic PDF/PNG main figures, captions, and a source - hash manifest, including the untouched D4 ResNet-20 and oral-B-v2 - confirmations, plus audited RRM failure and dynamic-stability supplements. + hash manifest, including the untouched D4 ResNet-20, oral-B-v2, and + 60-record standard-depth confirmations, plus audited RRM failure and + dynamic-stability supplements. The current main figures show the local-method Pareto frontier, credit assignment versus depth, the load-bearing innovation ablation, and the -untouched standard-ResNet confirmation. The first three require exactly seeds -0--4; the ResNet panel requires paired test seeds 10--14. Both strict renderers -refuse missing, dirty, protocol-mixed, or gate-inconsistent cells. +untouched standard-ResNet confirmations, including positive ResNet-20-to-56 +scaling. The first three require exactly seeds 0--4; the ResNet panels require +paired test seeds 10--14. Strict renderers refuse missing, dirty, +protocol-mixed, source-drifted, or gate-inconsistent cells. ## Verification @@ -165,9 +189,9 @@ experiments/finalize_accept.sh ``` This regenerates all publication-facing figures and manifests, enforces the -frozen Pareto/scaling and D4 confirmation gates, verifies baseline and protocol -mechanics, checks the local-rule smoke tests, evaluates the theory identities, -re-audits completed oral-B outcomes, and rebuilds the audited tables. +frozen Pareto/scaling, D4, oral-B-v2, and 60-record standard-depth gates, +verifies baseline and protocol mechanics, checks the local-rule smoke tests, +evaluates the theory identities, and rebuilds the audited tables. The convolutional infrastructure can be checked independently: @@ -239,9 +263,9 @@ until a new frozen gate is completed. `experiments/finalize_accept.sh` additionally requires strict imports of the frozen BurstCCN and Dual Propagation author-code runs. Both records now pass; -the accept-bar mechanical audit is green. The standard ResNet branch stopped -at its failed A3 validation gate, and the untouched A4 test panel remains -sealed. +the accept-bar mechanical audit is green. The original standard-ResNet branch +stopped at its failed A3 validation gate and its A4 panel remains sealed; the +independently authorized v2 recovery is the separate passed 60-record panel. ## Result discipline @@ -251,9 +275,9 @@ accounting, and final finiteness. Development, validation, and untouched confirmation results are never pooled. Failed gates close their branch instead of triggering seed deletion or post-hoc threshold changes. -The current formal milestone is `8/10` (accept, confidence `4/5`) after the -untouched D4 and oral-B-v2 confirmations. A conservative external-review -forecast is `7/10`: positive added-depth utility, original-data biological +The current formal milestone is `9/10` after the untouched D4, oral-B-v2, and +standard-depth confirmations. A conservative external-review forecast is +`8/10`: the full matched nine-method crossover, original-data biological validation, and a credit pathway novel beyond inherited perturbation/KP mechanisms remain open. Scores change only after an audited frozen stage, not after a pilot or presentation improvement. @@ -1091,6 +1091,28 @@ trajectory, and resource audit directly from these ten records. Its strict source-hash manifest refuses an incomplete seed set, failed D4 check, dirty provenance, protocol drift, or disagreement between the gate and records. +The separately authorized oral-A-v2 confirmation subsequently reuses those ten +D4 records verbatim and adds exactly 50 preregistered endpoints. Its complete +60-record, five-seed panel passes every accuracy, paired depth-gain, alignment, +mechanism, query, MAC, memory, provenance, and test-isolation check: + +| depth | BP test acc. | DFA test acc. | clean KP test acc. | dynamic SDIL test acc. | SDIL early alignment | SDIL/BP MACs | +|--:|--:|--:|--:|--:|--:|--:| +| 20 | 91.624% | 31.878% | 91.388% | 91.584% | 0.999687 | 1.326x | +| 32 | 92.302% | 32.684% | 92.332% | 92.254% | 0.999613 | 1.329x | +| 56 | 92.632% | 30.850% | 92.670% | **92.760%** | 0.999423 | 1.331x | + +Dynamic SDIL gains `1.176` accuracy points from ResNet-20 to ResNet-56 on +average, and all five paired seeds improve. Its ResNet-56 advantage over DFA is +`61.910` points, while its mean endpoint differs from BP and clean KP by only +`+0.128` and `+0.090` points. The passed gate establishes positive +standard-depth scaling and raises the internal milestone from 8 to 9. It does +not assign the clean-task gain uniquely to residualization: the inherited clean +KP substrate scales equally well, so SDIL's method-specific evidence remains +its survival under the four-times-RMS mixed-traffic intervention. +`results/figs/figure6_standard_depth_scaling.{pdf,png}` is regenerated from all +60 records and verifies their hashes plus the historical training source. + ## How to run `experiments/run.py --mode {bp,fa,dfa,sdil} --dataset {mnist,fmnist,cifar10} --depth D --residual {0,1} --act {tanh,gelu,silu,relu}` Batteries: `experiments/run_v2.sh <ds> "<depths>" <res> <act> "<seeds>" <ep> <pfx>`. @@ -1104,7 +1126,7 @@ C2 direct causal diagnosis: `experiments/c2_nodepert_validation.sh` followed by Theory checks: `python experiments/verify_theory.py`. Audited aggregate tables: `experiments/analyze_verified.py`. The complete accept-bar finalizer is `experiments/finalize_accept.sh`; it rejects -missing/dirty or gate-inconsistent inputs, regenerates the four main figures, +missing/dirty or gate-inconsistent inputs, regenerates the six main figures, their captions and source-hash manifests, rebuilds `results/audited_tables.md`, and runs all smoke and completed-gate audits. @@ -1115,10 +1137,9 @@ their captions and source-hash manifests, rebuilds SDIL. - Compare simultaneous calibration at fixed loss-evaluation budgets and sweep directions (4/8/16/32). -- A genuinely depth-necessary standard-network comparison remains absent. The - frozen ResNet-20/32/56 route is sealed by the failed oral-B prerequisite; any - replacement would require a new, independently preregistered mechanism - branch rather than reusing the untouched panel. +- Complete the registered 81-cell crossover: all nine methods at each of + miniCNN/VGGlike/VGG16, ResNet-20/32/56, and Transformer-4/8/12. No depth, + width, or context extension counts unless all nine method cells are added. - Diagnose why residual FA has high directional cosine but lower accuracy: log update-norm ratios, parameter-space cosine, and serial feedback latency/cost. - Reproduce the d2 EP basin sensitivity in the original Theano revision and/or a strong modern EP @@ -393,7 +393,8 @@ ended finite at `33.06%`; SDIL became nonfinite at epoch 89 and ended at `10.00%`, failing alignment, accuracy, and finiteness gates. The MAC gate alone passed. Per the stop rule, the five-seed ResNet-20/32/56 test panel was never run. The exact BatchNorm/local-gradient mechanics remain verified, but the -current training recipe does not establish standard-network scaling. +original channel-gated training recipe does not establish standard-network +scaling. A separate dynamic-innovation recovery is now frozen in `ORAL_A_RECOVERY.md` before any new endpoint. It does not reopen failed A3/A4: @@ -403,6 +404,16 @@ KP/dynamic cells verbatim; exactly 50 new cells are permitted. A full pass is the sole 8-to-9 score gate and must show a positive paired depth benefit, not merely survival on another depth-flat task. +**Independent oral-A-v2 status: passed.** After the calibrated oral-B-v2 +confirmation separately satisfied the prerequisite, `ORAL_A_RECOVERY_V2.md` +froze a new protocol without reopening the old panel. All 60 +BP/DFA/clean-KP/dynamic ResNet-20/32/56 records pass. Dynamic innovation rises +`91.584% -> 92.254% -> 92.760%`; every depth-20-to-56 seed pair improves, with +a `1.176`-point mean gain, and mean early-third alignment remains +`0.999687/0.999613/0.999423`. The internal milestone is therefore 9/10. The +next evidence gap is a full matched strong-baseline and architecture crossover, +not positive standard-depth utility. + Post-failure diagnosis is recorded separately in `results/oral_a_failure_diagnosis.json`: prediction--target cosine remains below `8.51e-5` in magnitude, calibration MSE is indistinguishable from target |
