From 524c0704029c4b831d9f19413995e8cff1508b8b Mon Sep 17 00:00:00 2001 From: YurenHao0426 Date: Wed, 22 Jul 2026 16:31:15 -0500 Subject: docs: close mixed-traffic path after short failure --- KP_BASELINE.md | 4 +++- MIXED_TRAFFIC.md | 10 ++++++++++ PAPER_PLAN.md | 7 +++++++ README.md | 9 +++++++-- RESULTS.md | 14 ++++++++++++++ REVIEW_SCORECARD.md | 14 ++++++++------ ROADMAP.md | 17 +++++++++-------- 7 files changed, 58 insertions(+), 17 deletions(-) diff --git a/KP_BASELINE.md b/KP_BASELINE.md index 27fd5ff..4b7ac86 100644 --- a/KP_BASELINE.md +++ b/KP_BASELINE.md @@ -120,4 +120,6 @@ Every epoch loss and tracking diagnostic is finite. The run uses 1080. It is 0.36 accuracy points below matched BP while using no reverse-mode gradient or loss query to train its feedback path. The result opens MT-1 but does not change the 5/10 reviewer score: reciprocal KP is inherited substrate, -not evidence that somato-dendritic innovation is load-bearing. +not evidence that somato-dendritic innovation is load-bearing. The subsequently +completed MT-1 panel fails under all three traffic rules, so this opening does +not advance to full validation or test confirmation. diff --git a/MIXED_TRAFFIC.md b/MIXED_TRAFFIC.md index 0f14e05..4852ef0 100644 --- a/MIXED_TRAFFIC.md +++ b/MIXED_TRAFFIC.md @@ -119,6 +119,16 @@ The accuracy gaps and same-state alignment rotation are conjunctive: clipping an unstable raw update is not enough. MT-1 failure closes the branch without a weaker traffic ratio. A pass opens MT-2 but leaves the reviewer score at 5/10. +MT-1 fails. All three clean-revision records become nonfinite during epoch 1 +and end at 10.00% validation accuracy with NaN loss. Thus innovation does not +separate from either raw control at the frozen four-to-one intervention. The +initial layer-ratio error is only `4.77e-7`, neutral warmup leaves a traffic +residual ratio of `0.139942`, all conditions use zero task-loss queries, and +each costs `1.3271x` the epoch-20 BP MAC reference. Calibration, warmup, query, +and cost invariants therefore pass, but they do not rescue the failed finite, +accuracy, alignment, tracking, or norm checks. MT-2 and MT-3 remain untouched; +the branch closes without changing traffic strength or predictor schedule. + ## MT-2: frozen full validation gate Copy MT-1 exactly for 200 epochs. The three jobs remain seed 0 and do not touch diff --git a/PAPER_PLAN.md b/PAPER_PLAN.md index baf5ebe..8e85aae 100644 --- a/PAPER_PLAN.md +++ b/PAPER_PLAN.md @@ -213,3 +213,10 @@ If it passes, revise the abstract and contribution language narrowly: If MT-1, MT-2, or MT-3 fails, preserve the current mechanism-only narrative and the score of 5/10. No lower traffic ratio, deleted seed, or replacement confirmation panel is permitted. + +This stop rule has now fired at MT-1. All three signal conditions become +nonfinite in epoch 1 and end at chance despite passing the calibration, +predictor-warmup, query, and cost invariants. MT-2 and MT-3 remain untouched. +The paper therefore retains the mechanism-focused title and 5/10 assessment; +the controlled standard-ResNet recovery is a disclosed negative result, not an +active acceptance claim. diff --git a/README.md b/README.md index f855473..5af05cf 100644 --- a/README.md +++ b/README.md @@ -59,7 +59,9 @@ and scaling behavior. See `NOVELTY.md` for the exact prior-art boundary. reaches 0.999998, exposing final alignment as an inadequate certificate for an intermittent tracker. The separate reciprocal KP short gate reaches 82.66%, and its frozen full gate reaches 91.26% versus matched BP's 91.62% - with 0.9997 late feedback cosine. It is the active strong substrate. + with 0.9997 late feedback cosine. The subsequent frozen mixed-traffic screen + nevertheless makes raw, norm-matched raw, and innovation all nonfinite in + epoch 1, so no full or confirmation panel opens. - Native author-code fidelity is complete. BurstCCN reaches `80.10%` at its validation-selected epoch versus published `82.97 +/- 0.21%`; Dual Prop reaches `92.46%` versus published `92.41 +/- 0.07%`. Their audited walls are @@ -162,7 +164,10 @@ The stronger reciprocal Kolen--Pollack substrate is audited separately in feedback loss queries. Before that full endpoint, `MIXED_TRAFFIC.md` froze the actual Harnett-specific test: raw, norm-matched raw, and innovation under identical four-times-RMS soma-predictable apical traffic and predictor cost. -Mechanics are green and MT-1 task access is now open. +Mechanics pass, but the complete MT-1 panel fails: all three signals become +nonfinite in epoch 1 and end at 10.00% validation accuracy. MT-2 and MT-3 +remain untouched, closing this standard-ResNet recovery without a weaker +traffic intervention. The subsequent V3 mechanism estimates the required A/G matrix statistics directly by perturbing the vectorizer parameter subspace. It remains diff --git a/RESULTS.md b/RESULTS.md index dfcb352..743e5bd 100644 --- a/RESULTS.md +++ b/RESULTS.md @@ -923,6 +923,20 @@ establishes a stable, near-BP standard-ResNet substrate and opens the frozen mixed-traffic screen. It remains prior-art evidence and therefore leaves the reviewer score at 5/10 until innovation itself wins the controlled ablation. +The frozen MT-1 load-bearing screen fails decisively. Raw, norm-matched raw, +and innovation all become nonfinite during epoch 1 and end at exactly 10.00% +validation accuracy with NaN loss. They share clean source revision `412314e`, +the same data order, initialization, four-times-RMS traffic, neutral predictor +schedule, and `1.3271x` BP affine-MAC cost. + +The failure is not explained by the pre-endpoint mechanical invariants: the +maximum initial traffic-ratio error is `4.77e-7`, the maximum post-warmup +traffic-residual ratio is `0.139942` versus the frozen `0.25` ceiling, and all +conditions use zero task-loss queries. Nevertheless all finite, accuracy, +alignment, feedback-tracking, and usable matched-norm checks fail. MT-2 and the +five-seed test confirmation remain untouched. This closes the controlled +standard-ResNet accept recovery and leaves the reviewer estimate at 5/10. + ## How to run `experiments/run.py --mode {bp,fa,dfa,sdil} --dataset {mnist,fmnist,cifar10} --depth D --residual {0,1} --act {tanh,gelu,silu,relu}` Batteries: `experiments/run_v2.sh "" "" `. diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md index 73fb424..bbc5bd3 100644 --- a/REVIEW_SCORECARD.md +++ b/REVIEW_SCORECARD.md @@ -25,7 +25,7 @@ Every formal result report records: | Soundness | 3/4 | Theory, local-gradient checks, causal diagnostics, cost accounting, and frozen stop rules are unusually careful. The learned apical vectorizer remains an unresolved failure mode. | | Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. | | Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, but the current flattened-CIFAR task does not benefit from depth and the frozen standard-ResNet recipe failed. | -| Empirical support | 2/4 | Five-depth scaling and the residualization ablation are strong. Useful-depth C2, broad endogenous C1, oral-B, and full ResNet A3 gates failed; the untouched A4 panel correctly remained sealed. | +| Empirical support | 2/4 | Five-depth scaling and the residualization ablation are strong. Useful-depth C2, broad endogenous C1, oral-B, full ResNet A3, and KP mixed-traffic MT-1 gates failed; untouched confirmation panels correctly remained sealed. | | Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. | | **Overall** | **5/10** | **Borderline reject: a strong core result without a completed standard-scale or useful-depth demonstration.** | | Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. | @@ -47,8 +47,9 @@ Every formal result report records: 3. Innovation was not uniformly beneficial for arbitrary endogenous top-down traffic, so the supported mechanism is narrower than the initial claim. 4. The frozen oral-B screen falsified the desired-velocity and Harnett error-derivative claims. -5. The frozen ResNet-20 A3 run became nonfinite and ended at chance. The failed gate forbids the - A4 test panel; completed native baselines improve fairness but do not supply SDIL scale evidence. +5. The frozen ResNet-20 A3 run became nonfinite and ended at chance; the later KP mixed-traffic + MT-1 panel also became nonfinite under all three signals. A4, MT-2, and MT-3 remain sealed, so + completed native baselines improve fairness but do not supply SDIL scale evidence. ## Score trajectory and prospective gates @@ -69,9 +70,9 @@ Every formal result report records: | Residual response mirror full gate fails | 5 | RRM ends at 10% with NaN validation loss although endpoint Q/W cosine is 0.999998 | Closes intermittent mirroring and proves endpoint alignment is an inadequate trajectory certificate; KP remains the separate strong substrate | | Modified KP short gate passes | 5 | KP reaches 82.66% versus BP's epoch-20 81.02%, with 0.886 early alignment and 1.326x BP MACs | Strongly solves the substrate at useful scale, but all credit belongs to inherited reciprocal plasticity until innovation is load-bearing | | Modified KP full gate passes | 5 | KP reaches 91.26% versus matched BP's 91.62%, with 0.9994 early alignment, 0.9997 late feedback cosine, zero queries, and 1.326x BP MACs | Establishes a stable near-BP ResNet substrate and opens MT-1, but inherited KP evidence cannot raise the SDIL score | -| Mixed-traffic MT-1 | 5 if passed | A frozen short raw/matched/innovation screen must preserve clean-KP utility and show accuracy plus same-state directional gains | One seed and 20 epochs can only open full validation | -| Mixed-traffic MT-2 | 5 if passed | A full seed-0 panel must reach 88%, stay within 3 points of clean KP, and beat both controls under audited cost | A single development seed still cannot establish acceptance | -| Mixed-traffic MT-3 | 6 if passed | Five untouched network/traffic draws must retain 88% mean test accuracy and positive confidence-bounded paired gains | Establishes the controlled ResNet innovation claim; natural traffic and broad depth/architecture scaling remain open | +| Mixed-traffic MT-1 fails | 5 | Raw, norm-matched raw, and innovation all become nonfinite in epoch 1 and end at 10%; calibration and predictor warmup still pass | Closes the controlled standard-ResNet innovation path; MT-2/MT-3 remain untouched and no weaker traffic rescue is allowed | +| Mixed-traffic MT-2 | not opened | The prerequisite MT-1 gate failed | No full seed-0 validation claim is available | +| Mixed-traffic MT-3 | not opened | MT-2 was never opened; all five confirmation seeds remain untouched | No controlled ResNet innovation confirmation claim is available | | Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available | These are conditional reviewer forecasts, not promised scores. A failed stage leaves its negative @@ -100,6 +101,7 @@ retroactively reopened by success on standard vision benchmarks. | 2026-07-22 / modified KP-1 | One clean constant-LR ResNet record reaches 82.66%, above BP's epoch-20 81.02%, with 0.8856 early alignment, 0.8546 train-period feedback cosine, zero queries, and 1.326x BP MACs | 5 → 5 | Establishes a stable strong substrate under a gate that observes the training trajectory, but reciprocal Kolen--Pollack plasticity is prior art and adds no Harnett-specific evidence | | 2026-07-22 / `2ef7f94` modified KP-2 | One clean full ResNet record reaches 91.26%, 0.36 points below matched BP, with 0.9994 early alignment, 0.9997 late feedback cosine, zero queries, and 1.326x BP MACs | 5 → 5 | Removes substrate stability as the immediate blocker and opens the load-bearing innovation experiment, but prior-art reciprocal plasticity earns no novelty credit | | 2026-07-22 / mixed-traffic MT-0 | Exact raw/matched/innovation mechanics pass; synthetic ResNet-20 matched batch peaks at 0.857 GB allocated on GTX 1080 without task data | 5 → 5 | Removes implementation, graph, and memory objections before endpoint access; supplies no evidence yet that innovation is useful | +| 2026-07-22 / `f310ad5` mixed-traffic MT-1 | All three frozen conditions become nonfinite in epoch 1 and end at 10%; ratio calibration, predictor warmup, zero-query, and cost checks pass | 5 → 5 | Fails to make innovation load-bearing on a standard ResNet and closes MT-2/MT-3; the prior mechanism-only evidence survives but empirical support does not improve | Future rows are appended only after an audited frozen stage. A score staying flat is informative: engineering, theory exposition, or visualization may make the paper more defensible without diff --git a/ROADMAP.md b/ROADMAP.md index 1bf6e32..c425356 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -481,8 +481,8 @@ zero task-loss queries, and cost remains `1.3261x` BP. This opens MT-1 without raising the reviewer score because the result is entirely inherited KP evidence. -**KP mixed-traffic MT-0 status: mechanics passed; MT-1--MT-3 frozen before -endpoints.** `MIXED_TRAFFIC.md` fixes a four-to-one, initialization-calibrated +**KP mixed-traffic status: MT-0 passed; MT-1 failed.** `MIXED_TRAFFIC.md` fixes +a four-to-one, initialization-calibrated soma-predictable traffic intervention and crosses raw apical activity, per-example norm-matched raw activity, and neutral-period innovation on the same reciprocal KP substrate. The zero-traffic limit, exact-predictor limit, @@ -494,12 +494,13 @@ MACs, and every control pays the predictor schedule. The synthetic batch-128 matched path is finite on a GTX 1080 and peaks at 0.857 GB allocated after reset; no task endpoint was used for this hardware check. -MT-1 is now open after the KP-2 pass. A 20-epoch pass opens one full seed-0 -validation panel but cannot raise the reviewer score; a full pass opens the -already frozen five-seed, all-50k, one-test-evaluation confirmation. Only the -complete MT-3 confirmation can move the score from 5/10 to 6/10. This is the -current accept-bar path. It is a controlled predictable-traffic experiment, -not a retroactive rescue of the failed natural/top-down C1 or Oral-B gates. +The complete MT-1 panel then becomes nonfinite in epoch 1 under all three +signals and ends at 10.00% validation accuracy. Calibration error remains only +`4.77e-7`, predictor warmup reaches residual ratio `0.139942`, cost is +`1.3271x` BP, and feedback uses zero task-loss queries, so those mechanical +checks pass while all performance, finite, alignment, and tracking checks +fail. Per the frozen stop rule, MT-2 and MT-3 remain untouched, no weaker +traffic ratio is allowed, and this standard-ResNet accept path closes at 5/10. Prepare convolutional local-update primitives and ResNet-20/32/56 protocols early. Queue frozen runs opportunistically on authorized idle GPUs. Because BurstCCN already reports CIFAR-10 and -- cgit v1.2.3