diff options
| author | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-23 06:55:15 -0500 |
|---|---|---|
| committer | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-23 06:55:15 -0500 |
| commit | 2b8b36b249e5d05386460c58aab268ca4d02eb0e (patch) | |
| tree | cf92a8c627f760095dcf95fe994ad2c4aeaabaa2 | |
| parent | 2fed62a2486962e0054fbb9d38661ed76a2820f4 (diff) | |
docs: record accept-gate confirmation pass
| -rw-r--r-- | DYNAMIC_INNOVATION.md | 19 | ||||
| -rw-r--r-- | PAPER_PLAN.md | 21 | ||||
| -rw-r--r-- | README.md | 11 | ||||
| -rw-r--r-- | RESULTS.md | 31 | ||||
| -rw-r--r-- | REVIEW_SCORECARD.md | 23 | ||||
| -rw-r--r-- | ROADMAP.md | 12 |
6 files changed, 99 insertions, 18 deletions
diff --git a/DYNAMIC_INNOVATION.md b/DYNAMIC_INNOVATION.md index ae841ea..00ace0e 100644 --- a/DYNAMIC_INNOVATION.md +++ b/DYNAMIC_INNOVATION.md @@ -221,3 +221,22 @@ arbitrary clean revision is insufficient. Every record must also report an NVIDIA GTX 1080 with physical visibility restricted to the authorized timan107 GPU 5 or 7; device identity is an audit condition rather than an inferred property of the launch script. + +### D4 realized result + +D4 passes on 2026-07-23 without seed replacement, threshold changes, or +intermediate endpoint access. Dynamic innovation reaches `91.584%` mean test +accuracy across seeds 10--14 versus clean KP's `91.388%`; the paired +clean-minus-dynamic differences are +`[+0.06, -0.03, -0.78, -0.01, -0.22]` percentage points, with mean `-0.196` +and one-sided 95% upper bound `+0.131` points. Every dynamic seed reaches at +least `91.51%`, and mean early-third alignment is `0.999687`. + +All ten trajectories and audited metrics are finite. The five dynamic records +also pass the frozen feedback, projection, instruction-leakage, query, MAC, +elementwise-work, peak-memory, hardware, source, data-split, and evaluation +checks with no invariant failure. The immutable gate is +`results/kp_dynamic_projection_confirmation_gate.json`, and the ten raw +records are in `results/kp_dynamic_projection_confirmation/`. This realizes +the predeclared 6-to-7 score change and opens oral-B R1; it does not establish +positive utility from added standard-network depth. diff --git a/PAPER_PLAN.md b/PAPER_PLAN.md index cc556d6..04f1d2f 100644 --- a/PAPER_PLAN.md +++ b/PAPER_PLAN.md @@ -263,8 +263,19 @@ at `10%`; all projection, query, MAC, memory, and evaluation-boundary checks pass. This is the first standard-ResNet evidence for load-bearing innovation, but it is one short validation run and therefore leaves the score at 5/10. -The predeclared D3 run is the next paper-changing gate. Only a complete -near-BP 200-epoch pass can move the simulated reviewer score to 6 and authorize -an independently frozen multi-seed test confirmation. If D3 fails, preserve -the dynamic short result as a stability diagnosis and do not replace the -schedule, thresholds, or traffic intervention. +The predeclared D3 run passes all 19 gates at `91.18%` validation accuracy and +opens the independently frozen D4 confirmation. D4 then passes without seed +replacement or threshold changes: dynamic innovation reaches `91.584%` mean +test accuracy across seeds 10--14 versus clean KP's `91.388%`, with a +clean-minus-dynamic one-sided 95% upper bound of `0.131` points. All mechanism, +query, cost, memory, provenance, split, and endpoint-isolation checks pass. +The strict reviewer score is therefore 7/10 and the narrow accept claim is now +active: somato-dendritic innovation is load-bearing and robust on a standard +ResNet-20 with an inherited zero-query reciprocal credit path. + +The paper must still attribute reciprocal KP plasticity to prior work, count +the paired neutral microphase and elementwise work, and retain the failed +arbitrary-top-down, desired-velocity, online-control, and earlier unstable +branches. D4 alone cannot support positive added-depth utility. The next +paper-changing gate is oral-B R1/R2; only its separately frozen confirmation +can raise 7 to 8 and unlock the standard ResNet-20/32/56 oral-A panel. @@ -74,8 +74,15 @@ and scaling behavior. See `NOVELTY.md` for the exact prior-art boundary. the predeclared 200-epoch D3 run subsequently passes all 19 gates at `91.18%` versus BP's `91.62%` and clean KP's `91.26%`, with `0.9994` early alignment, zero task-loss queries, and `1.3261x` BP MACs. This raises the strict score - to 6/10. A paired five-seed clean-KP/dynamic D4 protocol is frozen and now - open; no D4 test endpoint had been accessed at the D3 audit. + to 6/10. The independently frozen paired five-seed D4 test panel then + passes every gate: dynamic innovation reaches `91.584%` mean test accuracy + versus clean KP's `91.388%`, wins the mean pairing by `0.196` points, and + has a clean-minus-dynamic one-sided 95% upper bound of only `0.131` points. + Every dynamic seed reaches at least `91.51%`, mean early alignment is + `0.999687`, and all trajectory, projection, leakage, query, hardware, MAC, + memory, split, and test-isolation checks pass. This untouched confirmation + establishes the strict 7/10 accept bar; it does not establish that added + standard-network depth is useful or repair the failed biological signatures. - Native author-code fidelity is complete. BurstCCN reaches `80.10%` at its validation-selected epoch versus published `82.97 +/- 0.21%`; Dual Prop reaches `92.46%` versus published `92.41 +/- 0.07%`. Their audited walls are @@ -1012,8 +1012,35 @@ operations, one neutral observation per ordinary example, 2.165 GB peak allocated memory, one final validation evaluation, and zero test evaluations. The source revision is `d945c42`, and the gate was frozen before the endpoint. This is the first full near-BP standard-ResNet SDIL validation result and raises -the strict reviewer score from 5 to 6. Independent five-seed paired test -confirmation remains open rather than assumed. +the strict reviewer score from 5 to 6. At the D3 audit, independent five-seed +paired test confirmation remained open rather than assumed. + +The independently frozen D4 paired test confirmation subsequently passes every +gate at source revision `0008f2c`. Each of seeds 10--14 trains clean KP and +dynamic innovation from scratch on all 50,000 CIFAR-10 training examples for +200 epochs, evaluates no validation example, and evaluates the 10,000-example +test set exactly once at the endpoint: + +| seed | clean-KP test accuracy | dynamic test accuracy | clean minus dynamic | +|--:|--:|--:|--:| +| 10 | 91.58% | 91.52% | +0.06 points | +| 11 | 91.55% | 91.58% | -0.03 points | +| 12 | 90.87% | 91.65% | -0.78 points | +| 13 | 91.50% | 91.51% | -0.01 points | +| 14 | 91.44% | 91.66% | -0.22 points | +| **mean** | **91.388%** | **91.584%** | **-0.196 points** | + +The one-sided 95% upper confidence bound on the paired clean-minus-dynamic +deficit is `0.1306` points, far inside the frozen 2.5-point bound. All five +dynamic seeds exceed 87%, all five are within two points of paired clean KP, +and mean early-third alignment is `0.999687`. All ten trajectories are +finite; the dynamic runs also pass feedback tracking, neutral projection, +zero instruction leakage, zero task-loss query, `<=1.34x` BP MAC, `<=2.5 GiB` +peak-memory, authorized-hardware, source-provenance, split, and evaluation +checks. D4 therefore raises the strict reviewer score from 6 to 7 and +establishes the accept bar. It is a ResNet-20 robustness/noninferiority result, +not evidence of positive utility from ResNet-20 to ResNet-56 and not evidence +for the previously failed desired-velocity or online-control interpretation. ## How to run `experiments/run.py --mode {bp,fa,dfa,sdil} --dataset {mnist,fmnist,cifar10} --depth D --residual {0,1} --act {tanh,gelu,silu,relu}` diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md index 63f8d45..4efdf72 100644 --- a/REVIEW_SCORECARD.md +++ b/REVIEW_SCORECARD.md @@ -18,16 +18,16 @@ Every formal result report records: - the strongest remaining reviewer objections; - whether the evidence is development, validation, or untouched confirmation. -## Current paper-level assessment (2026-07-22) +## Current paper-level assessment (2026-07-23) | Dimension | Score | Strict reviewer assessment | |:--|--:|:--| | Soundness | 3/4 | Theory, local-gradient checks, causal diagnostics, cost accounting, and frozen stop rules are unusually careful. The learned apical vectorizer remains an unresolved failure mode. | | Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. | | Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, and dynamic innovation now reaches near-BP accuracy on a standard ResNet-20. Added-depth utility and broader biological generality remain absent. | -| Empirical support | 3/4 | Five-depth scaling, residual necessity, and a fully frozen 91.18% ResNet-20 validation endpoint are strong. D4 test confirmation is still running/open, while useful-depth C2, broad endogenous C1, oral-B, A3, and the original MT-1 failed. | +| Empirical support | 4/4 | Five-depth scaling, residual necessity, the frozen 91.18% ResNet-20 validation endpoint, and an untouched five-seed paired test confirmation at 91.584% are strong. Useful-depth C2, broad endogenous C1, oral-B, A3, and the original MT-1 remain disclosed failures. | | Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. | -| **Overall** | **6/10** | **Borderline accept: the frozen full standard-ResNet endpoint resolves the primary scale objection; independent confirmation and broader novelty are still missing.** | +| **Overall** | **7/10** | **Weak accept: untouched five-seed test confirmation establishes that dynamic innovation is robust and noninferior to strong clean KP on ResNet-20. Positive added-depth utility and the biological signature remain oral-level gaps.** | | Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. | ### Evidence already carrying the paper @@ -36,10 +36,13 @@ Every formal result report records: depth 60, while DFA's early-layer alignment falls from `0.514` to `0.047`. - Under strong soma-predictable apical traffic, raw and norm-matched raw learning fall to about chance (`10.38%` and `10.31%`) while innovation learning retains `97.35%` accuracy. +- On untouched CIFAR-10 test endpoints, dynamic innovation reaches `91.584%` versus clean KP's + `91.388%` over five paired ResNet-20 seeds; its paired deficit upper bound is only `0.131` + points and every frozen mechanism/cost invariant passes. - The local update has a proved descent condition and an explicit query/MAC/memory audit; direct node perturbation isolates the learned vectorizer as the useful-depth bottleneck. -### Objections currently preventing acceptance confidence +### Strongest remaining objections 1. The depth result is preservation on a depth-flat task, not evidence that SDIL uses added depth. 2. The amortized apical vectorizer failed the frozen useful-depth gate; direct causal targets work @@ -48,14 +51,15 @@ Every formal result report records: supported mechanism is narrower than the initial claim. 4. The frozen oral-B screen falsified the desired-velocity and Harnett error-derivative claims. A temporal-difference factorization repairs the sign algebra only; it has no task endpoint. -5. The new full ResNet-20 evidence is one validation seed. The paired five-seed D4 test panel is - open but incomplete, so robustness and noninferiority to clean KP are not yet established. +5. The standard-network result inherits reciprocal KP and pays for a paired neutral microphase. + D4 establishes the innovation operation, not a new credit-transport mechanism or a literal + cortical implementation. ## Score trajectory and prospective gates | Checkpoint | Overall | What changed | Remaining ceiling | |:--|--:|:--|:--| -| Current audited package | 6 | Strong depth-preservation/residual-necessity evidence plus a frozen 91.18% ResNet-20 validation endpoint; negative gates retained | Independent test confirmation and oral evidence remain absent | +| Current audited package | 7 | The untouched D4 panel reaches 91.584% dynamic versus 91.388% clean KP over five paired test seeds, with every mechanism and cost invariant passing | Added-depth utility and oral-B biology remain absent | | Native baselines complete | 5 | BurstCCN is below its published endpoint; Dual Prop reproduces 92.46% versus 92.41%, with strict provenance and cost semantics | Fairness objection narrows, but SDIL gains no standard-scale evidence | | Oral-A A1/A2 | 5 | BP reached 91.62%; short channel-gated SDIL reached 41.98% versus tuned DFA at 37.16% | Development screening alone cannot raise the score | | Oral-A A3 fails | 5 | Full ResNet-20 SDIL became nonfinite at epoch 89 and ended at 10%; DFA ended finite at 33.06% | Standard-scale and oral-A claims are closed; A4 remains untouched | @@ -76,8 +80,8 @@ Every formal result report records: | Dynamic projection D1 | 5 | All 352 training-only steps remain finite while the fast neutral fit holds residual coupling near zero | Mechanics only; no held-out endpoint | | Dynamic projection D2 | 5 | One frozen 20-epoch record reaches 83.58%, above clean KP's 82.66%, with all stability and cost gates passing | Short single-seed validation cannot establish the full scaling claim | | Dynamic projection D3 | 6 | All 19 frozen checks pass at 91.18%, within 0.44/0.08 points of BP/clean KP, with 0.9994 early alignment and 1.326x BP MACs | One validation seed cannot establish robustness | -| Dynamic projection D4 | open | Before observing D3, a paired clean-KP/dynamic five-seed test protocol and executable gate were frozen | A complete D4 pass, with no seed replacement, is required for 7/10 | -| Oral-B recovery R1/R2 | not opened | Role/velocity factorization passes mechanics; the two-rate R1 and untouched 6-task-by-5-model R2 gates are frozen before any recovery task endpoint | R1 remains hard-gated on D4; only a complete R2 pass can establish innovation-guided plasticity and move 7 to 8 | +| Dynamic projection D4 | 7 | All ten untouched records pass: dynamic 91.584% versus clean KP 91.388%, paired upper deficit bound 0.131 points, early alignment 0.999687, and no invariant failures | Establishes ResNet-20 robustness/noninferiority, not positive depth utility | +| Oral-B recovery R1/R2 | R1 open | Role/velocity factorization passes mechanics; the two-rate R1 and untouched 6-task-by-5-model R2 gates were frozen before any recovery task endpoint | Only a complete R2 pass can establish innovation-guided plasticity and move 7 to 8 | | Oral-A dynamic depth recovery | not opened | A 60-cell ResNet-20/32/56 BP/DFA/clean-KP/dynamic panel, positive-depth-benefit gate, mechanism invariants, and fair cost bounds are frozen before any new endpoint | It is hard-gated on D4 and oral-B R2; only a complete pass can establish standard-depth scaling and move 8 to 9 | | Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available | @@ -116,6 +120,7 @@ plasticity-only recovery must pass its own separately frozen R1 and R2 gates. | 2026-07-22 / `15c60d0` dynamic projection D3 | The sole frozen full validation record passes 19/19 checks at 91.18% versus BP's 91.62% and clean KP's 91.26%, with 0.9994 early alignment, zero queries, and 1.326x BP MACs | 5 → 6 | Resolves the primary full-scale objection and opens the already frozen D4 panel; one validation seed and narrow biological scope prevent a stronger recommendation | | 2026-07-22 / `32122d0` oral-B recovery freeze | Before any recovery task endpoint, the untouched 30-record R2 protocol, task-cluster uncertainty, original B1/B2 signatures, plasticity lesion, digest binding, and immutable analyzer are executable | 6 → 6 | Removes a preregistration gap but supplies no empirical evidence; R1 is still sealed behind D4 and confirmation seeds remain untouched | | 2026-07-22 / `1853620` oral-A recovery freeze | Before any new standard-depth endpoint, a D4-reusing 60-cell ResNet-20/32/56 panel and exact runner/analyzer contract are executable, with BP/DFA/KP controls, positive depth gain, alignment, projection, query, MAC, and memory gates | 6 → 6 | Precommits the final oral claim without bypassing the required sequence; no experiment opens unless D4 reaches 7 and oral-B R2 reaches 8 | +| 2026-07-23 / `2fed62a` dynamic projection D4 | All ten untouched paired test records pass the frozen gate: dynamic 91.584% versus clean KP 91.388%, clean-minus-dynamic upper bound 0.131 points, 0.999687 early alignment, and zero invariant failures | 6 → 7 | Establishes the strict accept bar with independent ResNet-20 robustness/noninferiority; opens oral-B R1 but does not support added-depth or desired-velocity claims | Future rows are appended only after an audited frozen stage. A score staying flat is informative: engineering, theory exposition, or visualization may make the paper more defensible without @@ -573,6 +573,18 @@ one validation and zero test evaluations occur. Per the frozen rule, the strict reviewer score rises from 5 to 6 and the already committed D4 paired five-seed test panel is now open. +**Dynamic neutral projection D4: passed.** All ten untouched test records at +source revision `0008f2c` pass the frozen audit. Across seeds 10--14, dynamic +innovation reaches `91.584%` mean test accuracy versus clean KP's `91.388%`; +the mean paired clean-minus-dynamic deficit is `-0.196` points and its +one-sided 95% upper bound is `0.131` points. Every dynamic seed reaches at +least `91.51%`, mean early alignment is `0.999687`, and all finite-trajectory, +feedback-tracking, projection, leakage, query, cost, memory, hardware, +provenance, split, and exactly-once test checks pass. Per the frozen rule the +strict reviewer score rises from 6 to 7, establishing the accept bar and +opening oral-B recovery R1. Oral-A remains sealed until the separately frozen +oral-B R2 confirmation passes. + Prepare convolutional local-update primitives and ResNet-20/32/56 protocols early. Queue frozen runs opportunistically on authorized idle GPUs. Because BurstCCN already reports CIFAR-10 and ImageNet scaling, dataset scale alone is not novel. The oral-level target is a memorable joint |
