summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
authorYurenHao0426 <Blackhao0426@gmail.com>2026-07-23 06:55:15 -0500
committerYurenHao0426 <Blackhao0426@gmail.com>2026-07-23 06:55:15 -0500
commit2b8b36b249e5d05386460c58aab268ca4d02eb0e (patch)
treecf92a8c627f760095dcf95fe994ad2c4aeaabaa2
parent2fed62a2486962e0054fbb9d38661ed76a2820f4 (diff)
docs: record accept-gate confirmation pass
-rw-r--r--DYNAMIC_INNOVATION.md19
-rw-r--r--PAPER_PLAN.md21
-rw-r--r--README.md11
-rw-r--r--RESULTS.md31
-rw-r--r--REVIEW_SCORECARD.md23
-rw-r--r--ROADMAP.md12
6 files changed, 99 insertions, 18 deletions
diff --git a/DYNAMIC_INNOVATION.md b/DYNAMIC_INNOVATION.md
index ae841ea..00ace0e 100644
--- a/DYNAMIC_INNOVATION.md
+++ b/DYNAMIC_INNOVATION.md
@@ -221,3 +221,22 @@ arbitrary clean revision is insufficient.
Every record must also report an NVIDIA GTX 1080 with physical visibility
restricted to the authorized timan107 GPU 5 or 7; device identity is an audit
condition rather than an inferred property of the launch script.
+
+### D4 realized result
+
+D4 passes on 2026-07-23 without seed replacement, threshold changes, or
+intermediate endpoint access. Dynamic innovation reaches `91.584%` mean test
+accuracy across seeds 10--14 versus clean KP's `91.388%`; the paired
+clean-minus-dynamic differences are
+`[+0.06, -0.03, -0.78, -0.01, -0.22]` percentage points, with mean `-0.196`
+and one-sided 95% upper bound `+0.131` points. Every dynamic seed reaches at
+least `91.51%`, and mean early-third alignment is `0.999687`.
+
+All ten trajectories and audited metrics are finite. The five dynamic records
+also pass the frozen feedback, projection, instruction-leakage, query, MAC,
+elementwise-work, peak-memory, hardware, source, data-split, and evaluation
+checks with no invariant failure. The immutable gate is
+`results/kp_dynamic_projection_confirmation_gate.json`, and the ten raw
+records are in `results/kp_dynamic_projection_confirmation/`. This realizes
+the predeclared 6-to-7 score change and opens oral-B R1; it does not establish
+positive utility from added standard-network depth.
diff --git a/PAPER_PLAN.md b/PAPER_PLAN.md
index cc556d6..04f1d2f 100644
--- a/PAPER_PLAN.md
+++ b/PAPER_PLAN.md
@@ -263,8 +263,19 @@ at `10%`; all projection, query, MAC, memory, and evaluation-boundary checks
pass. This is the first standard-ResNet evidence for load-bearing innovation,
but it is one short validation run and therefore leaves the score at 5/10.
-The predeclared D3 run is the next paper-changing gate. Only a complete
-near-BP 200-epoch pass can move the simulated reviewer score to 6 and authorize
-an independently frozen multi-seed test confirmation. If D3 fails, preserve
-the dynamic short result as a stability diagnosis and do not replace the
-schedule, thresholds, or traffic intervention.
+The predeclared D3 run passes all 19 gates at `91.18%` validation accuracy and
+opens the independently frozen D4 confirmation. D4 then passes without seed
+replacement or threshold changes: dynamic innovation reaches `91.584%` mean
+test accuracy across seeds 10--14 versus clean KP's `91.388%`, with a
+clean-minus-dynamic one-sided 95% upper bound of `0.131` points. All mechanism,
+query, cost, memory, provenance, split, and endpoint-isolation checks pass.
+The strict reviewer score is therefore 7/10 and the narrow accept claim is now
+active: somato-dendritic innovation is load-bearing and robust on a standard
+ResNet-20 with an inherited zero-query reciprocal credit path.
+
+The paper must still attribute reciprocal KP plasticity to prior work, count
+the paired neutral microphase and elementwise work, and retain the failed
+arbitrary-top-down, desired-velocity, online-control, and earlier unstable
+branches. D4 alone cannot support positive added-depth utility. The next
+paper-changing gate is oral-B R1/R2; only its separately frozen confirmation
+can raise 7 to 8 and unlock the standard ResNet-20/32/56 oral-A panel.
diff --git a/README.md b/README.md
index eded470..0b712c4 100644
--- a/README.md
+++ b/README.md
@@ -74,8 +74,15 @@ and scaling behavior. See `NOVELTY.md` for the exact prior-art boundary.
the predeclared 200-epoch D3 run subsequently passes all 19 gates at `91.18%`
versus BP's `91.62%` and clean KP's `91.26%`, with `0.9994` early alignment,
zero task-loss queries, and `1.3261x` BP MACs. This raises the strict score
- to 6/10. A paired five-seed clean-KP/dynamic D4 protocol is frozen and now
- open; no D4 test endpoint had been accessed at the D3 audit.
+ to 6/10. The independently frozen paired five-seed D4 test panel then
+ passes every gate: dynamic innovation reaches `91.584%` mean test accuracy
+ versus clean KP's `91.388%`, wins the mean pairing by `0.196` points, and
+ has a clean-minus-dynamic one-sided 95% upper bound of only `0.131` points.
+ Every dynamic seed reaches at least `91.51%`, mean early alignment is
+ `0.999687`, and all trajectory, projection, leakage, query, hardware, MAC,
+ memory, split, and test-isolation checks pass. This untouched confirmation
+ establishes the strict 7/10 accept bar; it does not establish that added
+ standard-network depth is useful or repair the failed biological signatures.
- Native author-code fidelity is complete. BurstCCN reaches `80.10%` at its
validation-selected epoch versus published `82.97 +/- 0.21%`; Dual Prop
reaches `92.46%` versus published `92.41 +/- 0.07%`. Their audited walls are
diff --git a/RESULTS.md b/RESULTS.md
index d36edad..436740f 100644
--- a/RESULTS.md
+++ b/RESULTS.md
@@ -1012,8 +1012,35 @@ operations, one neutral observation per ordinary example, 2.165 GB peak
allocated memory, one final validation evaluation, and zero test evaluations.
The source revision is `d945c42`, and the gate was frozen before the endpoint.
This is the first full near-BP standard-ResNet SDIL validation result and raises
-the strict reviewer score from 5 to 6. Independent five-seed paired test
-confirmation remains open rather than assumed.
+the strict reviewer score from 5 to 6. At the D3 audit, independent five-seed
+paired test confirmation remained open rather than assumed.
+
+The independently frozen D4 paired test confirmation subsequently passes every
+gate at source revision `0008f2c`. Each of seeds 10--14 trains clean KP and
+dynamic innovation from scratch on all 50,000 CIFAR-10 training examples for
+200 epochs, evaluates no validation example, and evaluates the 10,000-example
+test set exactly once at the endpoint:
+
+| seed | clean-KP test accuracy | dynamic test accuracy | clean minus dynamic |
+|--:|--:|--:|--:|
+| 10 | 91.58% | 91.52% | +0.06 points |
+| 11 | 91.55% | 91.58% | -0.03 points |
+| 12 | 90.87% | 91.65% | -0.78 points |
+| 13 | 91.50% | 91.51% | -0.01 points |
+| 14 | 91.44% | 91.66% | -0.22 points |
+| **mean** | **91.388%** | **91.584%** | **-0.196 points** |
+
+The one-sided 95% upper confidence bound on the paired clean-minus-dynamic
+deficit is `0.1306` points, far inside the frozen 2.5-point bound. All five
+dynamic seeds exceed 87%, all five are within two points of paired clean KP,
+and mean early-third alignment is `0.999687`. All ten trajectories are
+finite; the dynamic runs also pass feedback tracking, neutral projection,
+zero instruction leakage, zero task-loss query, `<=1.34x` BP MAC, `<=2.5 GiB`
+peak-memory, authorized-hardware, source-provenance, split, and evaluation
+checks. D4 therefore raises the strict reviewer score from 6 to 7 and
+establishes the accept bar. It is a ResNet-20 robustness/noninferiority result,
+not evidence of positive utility from ResNet-20 to ResNet-56 and not evidence
+for the previously failed desired-velocity or online-control interpretation.
## How to run
`experiments/run.py --mode {bp,fa,dfa,sdil} --dataset {mnist,fmnist,cifar10} --depth D --residual {0,1} --act {tanh,gelu,silu,relu}`
diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md
index 63f8d45..4efdf72 100644
--- a/REVIEW_SCORECARD.md
+++ b/REVIEW_SCORECARD.md
@@ -18,16 +18,16 @@ Every formal result report records:
- the strongest remaining reviewer objections;
- whether the evidence is development, validation, or untouched confirmation.
-## Current paper-level assessment (2026-07-22)
+## Current paper-level assessment (2026-07-23)
| Dimension | Score | Strict reviewer assessment |
|:--|--:|:--|
| Soundness | 3/4 | Theory, local-gradient checks, causal diagnostics, cost accounting, and frozen stop rules are unusually careful. The learned apical vectorizer remains an unresolved failure mode. |
| Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. |
| Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, and dynamic innovation now reaches near-BP accuracy on a standard ResNet-20. Added-depth utility and broader biological generality remain absent. |
-| Empirical support | 3/4 | Five-depth scaling, residual necessity, and a fully frozen 91.18% ResNet-20 validation endpoint are strong. D4 test confirmation is still running/open, while useful-depth C2, broad endogenous C1, oral-B, A3, and the original MT-1 failed. |
+| Empirical support | 4/4 | Five-depth scaling, residual necessity, the frozen 91.18% ResNet-20 validation endpoint, and an untouched five-seed paired test confirmation at 91.584% are strong. Useful-depth C2, broad endogenous C1, oral-B, A3, and the original MT-1 remain disclosed failures. |
| Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. |
-| **Overall** | **6/10** | **Borderline accept: the frozen full standard-ResNet endpoint resolves the primary scale objection; independent confirmation and broader novelty are still missing.** |
+| **Overall** | **7/10** | **Weak accept: untouched five-seed test confirmation establishes that dynamic innovation is robust and noninferior to strong clean KP on ResNet-20. Positive added-depth utility and the biological signature remain oral-level gaps.** |
| Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. |
### Evidence already carrying the paper
@@ -36,10 +36,13 @@ Every formal result report records:
depth 60, while DFA's early-layer alignment falls from `0.514` to `0.047`.
- Under strong soma-predictable apical traffic, raw and norm-matched raw learning fall to about
chance (`10.38%` and `10.31%`) while innovation learning retains `97.35%` accuracy.
+- On untouched CIFAR-10 test endpoints, dynamic innovation reaches `91.584%` versus clean KP's
+ `91.388%` over five paired ResNet-20 seeds; its paired deficit upper bound is only `0.131`
+ points and every frozen mechanism/cost invariant passes.
- The local update has a proved descent condition and an explicit query/MAC/memory audit; direct
node perturbation isolates the learned vectorizer as the useful-depth bottleneck.
-### Objections currently preventing acceptance confidence
+### Strongest remaining objections
1. The depth result is preservation on a depth-flat task, not evidence that SDIL uses added depth.
2. The amortized apical vectorizer failed the frozen useful-depth gate; direct causal targets work
@@ -48,14 +51,15 @@ Every formal result report records:
supported mechanism is narrower than the initial claim.
4. The frozen oral-B screen falsified the desired-velocity and Harnett error-derivative claims.
A temporal-difference factorization repairs the sign algebra only; it has no task endpoint.
-5. The new full ResNet-20 evidence is one validation seed. The paired five-seed D4 test panel is
- open but incomplete, so robustness and noninferiority to clean KP are not yet established.
+5. The standard-network result inherits reciprocal KP and pays for a paired neutral microphase.
+ D4 establishes the innovation operation, not a new credit-transport mechanism or a literal
+ cortical implementation.
## Score trajectory and prospective gates
| Checkpoint | Overall | What changed | Remaining ceiling |
|:--|--:|:--|:--|
-| Current audited package | 6 | Strong depth-preservation/residual-necessity evidence plus a frozen 91.18% ResNet-20 validation endpoint; negative gates retained | Independent test confirmation and oral evidence remain absent |
+| Current audited package | 7 | The untouched D4 panel reaches 91.584% dynamic versus 91.388% clean KP over five paired test seeds, with every mechanism and cost invariant passing | Added-depth utility and oral-B biology remain absent |
| Native baselines complete | 5 | BurstCCN is below its published endpoint; Dual Prop reproduces 92.46% versus 92.41%, with strict provenance and cost semantics | Fairness objection narrows, but SDIL gains no standard-scale evidence |
| Oral-A A1/A2 | 5 | BP reached 91.62%; short channel-gated SDIL reached 41.98% versus tuned DFA at 37.16% | Development screening alone cannot raise the score |
| Oral-A A3 fails | 5 | Full ResNet-20 SDIL became nonfinite at epoch 89 and ended at 10%; DFA ended finite at 33.06% | Standard-scale and oral-A claims are closed; A4 remains untouched |
@@ -76,8 +80,8 @@ Every formal result report records:
| Dynamic projection D1 | 5 | All 352 training-only steps remain finite while the fast neutral fit holds residual coupling near zero | Mechanics only; no held-out endpoint |
| Dynamic projection D2 | 5 | One frozen 20-epoch record reaches 83.58%, above clean KP's 82.66%, with all stability and cost gates passing | Short single-seed validation cannot establish the full scaling claim |
| Dynamic projection D3 | 6 | All 19 frozen checks pass at 91.18%, within 0.44/0.08 points of BP/clean KP, with 0.9994 early alignment and 1.326x BP MACs | One validation seed cannot establish robustness |
-| Dynamic projection D4 | open | Before observing D3, a paired clean-KP/dynamic five-seed test protocol and executable gate were frozen | A complete D4 pass, with no seed replacement, is required for 7/10 |
-| Oral-B recovery R1/R2 | not opened | Role/velocity factorization passes mechanics; the two-rate R1 and untouched 6-task-by-5-model R2 gates are frozen before any recovery task endpoint | R1 remains hard-gated on D4; only a complete R2 pass can establish innovation-guided plasticity and move 7 to 8 |
+| Dynamic projection D4 | 7 | All ten untouched records pass: dynamic 91.584% versus clean KP 91.388%, paired upper deficit bound 0.131 points, early alignment 0.999687, and no invariant failures | Establishes ResNet-20 robustness/noninferiority, not positive depth utility |
+| Oral-B recovery R1/R2 | R1 open | Role/velocity factorization passes mechanics; the two-rate R1 and untouched 6-task-by-5-model R2 gates were frozen before any recovery task endpoint | Only a complete R2 pass can establish innovation-guided plasticity and move 7 to 8 |
| Oral-A dynamic depth recovery | not opened | A 60-cell ResNet-20/32/56 BP/DFA/clean-KP/dynamic panel, positive-depth-benefit gate, mechanism invariants, and fair cost bounds are frozen before any new endpoint | It is hard-gated on D4 and oral-B R2; only a complete pass can establish standard-depth scaling and move 8 to 9 |
| Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available |
@@ -116,6 +120,7 @@ plasticity-only recovery must pass its own separately frozen R1 and R2 gates.
| 2026-07-22 / `15c60d0` dynamic projection D3 | The sole frozen full validation record passes 19/19 checks at 91.18% versus BP's 91.62% and clean KP's 91.26%, with 0.9994 early alignment, zero queries, and 1.326x BP MACs | 5 → 6 | Resolves the primary full-scale objection and opens the already frozen D4 panel; one validation seed and narrow biological scope prevent a stronger recommendation |
| 2026-07-22 / `32122d0` oral-B recovery freeze | Before any recovery task endpoint, the untouched 30-record R2 protocol, task-cluster uncertainty, original B1/B2 signatures, plasticity lesion, digest binding, and immutable analyzer are executable | 6 → 6 | Removes a preregistration gap but supplies no empirical evidence; R1 is still sealed behind D4 and confirmation seeds remain untouched |
| 2026-07-22 / `1853620` oral-A recovery freeze | Before any new standard-depth endpoint, a D4-reusing 60-cell ResNet-20/32/56 panel and exact runner/analyzer contract are executable, with BP/DFA/KP controls, positive depth gain, alignment, projection, query, MAC, and memory gates | 6 → 6 | Precommits the final oral claim without bypassing the required sequence; no experiment opens unless D4 reaches 7 and oral-B R2 reaches 8 |
+| 2026-07-23 / `2fed62a` dynamic projection D4 | All ten untouched paired test records pass the frozen gate: dynamic 91.584% versus clean KP 91.388%, clean-minus-dynamic upper bound 0.131 points, 0.999687 early alignment, and zero invariant failures | 6 → 7 | Establishes the strict accept bar with independent ResNet-20 robustness/noninferiority; opens oral-B R1 but does not support added-depth or desired-velocity claims |
Future rows are appended only after an audited frozen stage. A score staying flat is informative:
engineering, theory exposition, or visualization may make the paper more defensible without
diff --git a/ROADMAP.md b/ROADMAP.md
index da17bb8..d10db78 100644
--- a/ROADMAP.md
+++ b/ROADMAP.md
@@ -573,6 +573,18 @@ one validation and zero test evaluations occur. Per the frozen rule, the
strict reviewer score rises from 5 to 6 and the already committed D4 paired
five-seed test panel is now open.
+**Dynamic neutral projection D4: passed.** All ten untouched test records at
+source revision `0008f2c` pass the frozen audit. Across seeds 10--14, dynamic
+innovation reaches `91.584%` mean test accuracy versus clean KP's `91.388%`;
+the mean paired clean-minus-dynamic deficit is `-0.196` points and its
+one-sided 95% upper bound is `0.131` points. Every dynamic seed reaches at
+least `91.51%`, mean early alignment is `0.999687`, and all finite-trajectory,
+feedback-tracking, projection, leakage, query, cost, memory, hardware,
+provenance, split, and exactly-once test checks pass. Per the frozen rule the
+strict reviewer score rises from 6 to 7, establishing the accept bar and
+opening oral-B recovery R1. Oral-A remains sealed until the separately frozen
+oral-B R2 confirmation passes.
+
Prepare convolutional local-update primitives and ResNet-20/32/56 protocols early. Queue frozen
runs opportunistically on authorized idle GPUs. Because BurstCCN already reports CIFAR-10 and
ImageNet scaling, dataset scale alone is not novel. The oral-level target is a memorable joint