summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
authorYurenHao0426 <Blackhao0426@gmail.com>2026-07-22 16:31:15 -0500
committerYurenHao0426 <Blackhao0426@gmail.com>2026-07-22 16:31:15 -0500
commit524c0704029c4b831d9f19413995e8cff1508b8b (patch)
tree4df653547478723fa61df58b37af423bc3709583
parentf310ad55b76ab96de488cdb4f2d05861ed4616f2 (diff)
docs: close mixed-traffic path after short failure
-rw-r--r--KP_BASELINE.md4
-rw-r--r--MIXED_TRAFFIC.md10
-rw-r--r--PAPER_PLAN.md7
-rw-r--r--README.md9
-rw-r--r--RESULTS.md14
-rw-r--r--REVIEW_SCORECARD.md14
-rw-r--r--ROADMAP.md17
7 files changed, 58 insertions, 17 deletions
diff --git a/KP_BASELINE.md b/KP_BASELINE.md
index 27fd5ff..4b7ac86 100644
--- a/KP_BASELINE.md
+++ b/KP_BASELINE.md
@@ -120,4 +120,6 @@ Every epoch loss and tracking diagnostic is finite. The run uses
1080. It is 0.36 accuracy points below matched BP while using no reverse-mode
gradient or loss query to train its feedback path. The result opens MT-1 but
does not change the 5/10 reviewer score: reciprocal KP is inherited substrate,
-not evidence that somato-dendritic innovation is load-bearing.
+not evidence that somato-dendritic innovation is load-bearing. The subsequently
+completed MT-1 panel fails under all three traffic rules, so this opening does
+not advance to full validation or test confirmation.
diff --git a/MIXED_TRAFFIC.md b/MIXED_TRAFFIC.md
index 0f14e05..4852ef0 100644
--- a/MIXED_TRAFFIC.md
+++ b/MIXED_TRAFFIC.md
@@ -119,6 +119,16 @@ The accuracy gaps and same-state alignment rotation are conjunctive: clipping
an unstable raw update is not enough. MT-1 failure closes the branch without a
weaker traffic ratio. A pass opens MT-2 but leaves the reviewer score at 5/10.
+MT-1 fails. All three clean-revision records become nonfinite during epoch 1
+and end at 10.00% validation accuracy with NaN loss. Thus innovation does not
+separate from either raw control at the frozen four-to-one intervention. The
+initial layer-ratio error is only `4.77e-7`, neutral warmup leaves a traffic
+residual ratio of `0.139942`, all conditions use zero task-loss queries, and
+each costs `1.3271x` the epoch-20 BP MAC reference. Calibration, warmup, query,
+and cost invariants therefore pass, but they do not rescue the failed finite,
+accuracy, alignment, tracking, or norm checks. MT-2 and MT-3 remain untouched;
+the branch closes without changing traffic strength or predictor schedule.
+
## MT-2: frozen full validation gate
Copy MT-1 exactly for 200 epochs. The three jobs remain seed 0 and do not touch
diff --git a/PAPER_PLAN.md b/PAPER_PLAN.md
index baf5ebe..8e85aae 100644
--- a/PAPER_PLAN.md
+++ b/PAPER_PLAN.md
@@ -213,3 +213,10 @@ If it passes, revise the abstract and contribution language narrowly:
If MT-1, MT-2, or MT-3 fails, preserve the current mechanism-only narrative and
the score of 5/10. No lower traffic ratio, deleted seed, or replacement
confirmation panel is permitted.
+
+This stop rule has now fired at MT-1. All three signal conditions become
+nonfinite in epoch 1 and end at chance despite passing the calibration,
+predictor-warmup, query, and cost invariants. MT-2 and MT-3 remain untouched.
+The paper therefore retains the mechanism-focused title and 5/10 assessment;
+the controlled standard-ResNet recovery is a disclosed negative result, not an
+active acceptance claim.
diff --git a/README.md b/README.md
index f855473..5af05cf 100644
--- a/README.md
+++ b/README.md
@@ -59,7 +59,9 @@ and scaling behavior. See `NOVELTY.md` for the exact prior-art boundary.
reaches 0.999998, exposing final alignment as an inadequate certificate for
an intermittent tracker. The separate reciprocal KP short gate reaches
82.66%, and its frozen full gate reaches 91.26% versus matched BP's 91.62%
- with 0.9997 late feedback cosine. It is the active strong substrate.
+ with 0.9997 late feedback cosine. The subsequent frozen mixed-traffic screen
+ nevertheless makes raw, norm-matched raw, and innovation all nonfinite in
+ epoch 1, so no full or confirmation panel opens.
- Native author-code fidelity is complete. BurstCCN reaches `80.10%` at its
validation-selected epoch versus published `82.97 +/- 0.21%`; Dual Prop
reaches `92.46%` versus published `92.41 +/- 0.07%`. Their audited walls are
@@ -162,7 +164,10 @@ The stronger reciprocal Kolen--Pollack substrate is audited separately in
feedback loss queries. Before that full endpoint, `MIXED_TRAFFIC.md` froze the
actual Harnett-specific test: raw, norm-matched raw, and innovation under
identical four-times-RMS soma-predictable apical traffic and predictor cost.
-Mechanics are green and MT-1 task access is now open.
+Mechanics pass, but the complete MT-1 panel fails: all three signals become
+nonfinite in epoch 1 and end at 10.00% validation accuracy. MT-2 and MT-3
+remain untouched, closing this standard-ResNet recovery without a weaker
+traffic intervention.
The subsequent V3 mechanism estimates the required A/G matrix statistics
directly by perturbing the vectorizer parameter subspace. It remains
diff --git a/RESULTS.md b/RESULTS.md
index dfcb352..743e5bd 100644
--- a/RESULTS.md
+++ b/RESULTS.md
@@ -923,6 +923,20 @@ establishes a stable, near-BP standard-ResNet substrate and opens the frozen
mixed-traffic screen. It remains prior-art evidence and therefore leaves the
reviewer score at 5/10 until innovation itself wins the controlled ablation.
+The frozen MT-1 load-bearing screen fails decisively. Raw, norm-matched raw,
+and innovation all become nonfinite during epoch 1 and end at exactly 10.00%
+validation accuracy with NaN loss. They share clean source revision `412314e`,
+the same data order, initialization, four-times-RMS traffic, neutral predictor
+schedule, and `1.3271x` BP affine-MAC cost.
+
+The failure is not explained by the pre-endpoint mechanical invariants: the
+maximum initial traffic-ratio error is `4.77e-7`, the maximum post-warmup
+traffic-residual ratio is `0.139942` versus the frozen `0.25` ceiling, and all
+conditions use zero task-loss queries. Nevertheless all finite, accuracy,
+alignment, feedback-tracking, and usable matched-norm checks fail. MT-2 and the
+five-seed test confirmation remain untouched. This closes the controlled
+standard-ResNet accept recovery and leaves the reviewer estimate at 5/10.
+
## How to run
`experiments/run.py --mode {bp,fa,dfa,sdil} --dataset {mnist,fmnist,cifar10} --depth D --residual {0,1} --act {tanh,gelu,silu,relu}`
Batteries: `experiments/run_v2.sh <ds> "<depths>" <res> <act> "<seeds>" <ep> <pfx>`.
diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md
index 73fb424..bbc5bd3 100644
--- a/REVIEW_SCORECARD.md
+++ b/REVIEW_SCORECARD.md
@@ -25,7 +25,7 @@ Every formal result report records:
| Soundness | 3/4 | Theory, local-gradient checks, causal diagnostics, cost accounting, and frozen stop rules are unusually careful. The learned apical vectorizer remains an unresolved failure mode. |
| Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. |
| Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, but the current flattened-CIFAR task does not benefit from depth and the frozen standard-ResNet recipe failed. |
-| Empirical support | 2/4 | Five-depth scaling and the residualization ablation are strong. Useful-depth C2, broad endogenous C1, oral-B, and full ResNet A3 gates failed; the untouched A4 panel correctly remained sealed. |
+| Empirical support | 2/4 | Five-depth scaling and the residualization ablation are strong. Useful-depth C2, broad endogenous C1, oral-B, full ResNet A3, and KP mixed-traffic MT-1 gates failed; untouched confirmation panels correctly remained sealed. |
| Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. |
| **Overall** | **5/10** | **Borderline reject: a strong core result without a completed standard-scale or useful-depth demonstration.** |
| Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. |
@@ -47,8 +47,9 @@ Every formal result report records:
3. Innovation was not uniformly beneficial for arbitrary endogenous top-down traffic, so the
supported mechanism is narrower than the initial claim.
4. The frozen oral-B screen falsified the desired-velocity and Harnett error-derivative claims.
-5. The frozen ResNet-20 A3 run became nonfinite and ended at chance. The failed gate forbids the
- A4 test panel; completed native baselines improve fairness but do not supply SDIL scale evidence.
+5. The frozen ResNet-20 A3 run became nonfinite and ended at chance; the later KP mixed-traffic
+ MT-1 panel also became nonfinite under all three signals. A4, MT-2, and MT-3 remain sealed, so
+ completed native baselines improve fairness but do not supply SDIL scale evidence.
## Score trajectory and prospective gates
@@ -69,9 +70,9 @@ Every formal result report records:
| Residual response mirror full gate fails | 5 | RRM ends at 10% with NaN validation loss although endpoint Q/W cosine is 0.999998 | Closes intermittent mirroring and proves endpoint alignment is an inadequate trajectory certificate; KP remains the separate strong substrate |
| Modified KP short gate passes | 5 | KP reaches 82.66% versus BP's epoch-20 81.02%, with 0.886 early alignment and 1.326x BP MACs | Strongly solves the substrate at useful scale, but all credit belongs to inherited reciprocal plasticity until innovation is load-bearing |
| Modified KP full gate passes | 5 | KP reaches 91.26% versus matched BP's 91.62%, with 0.9994 early alignment, 0.9997 late feedback cosine, zero queries, and 1.326x BP MACs | Establishes a stable near-BP ResNet substrate and opens MT-1, but inherited KP evidence cannot raise the SDIL score |
-| Mixed-traffic MT-1 | 5 if passed | A frozen short raw/matched/innovation screen must preserve clean-KP utility and show accuracy plus same-state directional gains | One seed and 20 epochs can only open full validation |
-| Mixed-traffic MT-2 | 5 if passed | A full seed-0 panel must reach 88%, stay within 3 points of clean KP, and beat both controls under audited cost | A single development seed still cannot establish acceptance |
-| Mixed-traffic MT-3 | 6 if passed | Five untouched network/traffic draws must retain 88% mean test accuracy and positive confidence-bounded paired gains | Establishes the controlled ResNet innovation claim; natural traffic and broad depth/architecture scaling remain open |
+| Mixed-traffic MT-1 fails | 5 | Raw, norm-matched raw, and innovation all become nonfinite in epoch 1 and end at 10%; calibration and predictor warmup still pass | Closes the controlled standard-ResNet innovation path; MT-2/MT-3 remain untouched and no weaker traffic rescue is allowed |
+| Mixed-traffic MT-2 | not opened | The prerequisite MT-1 gate failed | No full seed-0 validation claim is available |
+| Mixed-traffic MT-3 | not opened | MT-2 was never opened; all five confirmation seeds remain untouched | No controlled ResNet innovation confirmation claim is available |
| Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available |
These are conditional reviewer forecasts, not promised scores. A failed stage leaves its negative
@@ -100,6 +101,7 @@ retroactively reopened by success on standard vision benchmarks.
| 2026-07-22 / modified KP-1 | One clean constant-LR ResNet record reaches 82.66%, above BP's epoch-20 81.02%, with 0.8856 early alignment, 0.8546 train-period feedback cosine, zero queries, and 1.326x BP MACs | 5 → 5 | Establishes a stable strong substrate under a gate that observes the training trajectory, but reciprocal Kolen--Pollack plasticity is prior art and adds no Harnett-specific evidence |
| 2026-07-22 / `2ef7f94` modified KP-2 | One clean full ResNet record reaches 91.26%, 0.36 points below matched BP, with 0.9994 early alignment, 0.9997 late feedback cosine, zero queries, and 1.326x BP MACs | 5 → 5 | Removes substrate stability as the immediate blocker and opens the load-bearing innovation experiment, but prior-art reciprocal plasticity earns no novelty credit |
| 2026-07-22 / mixed-traffic MT-0 | Exact raw/matched/innovation mechanics pass; synthetic ResNet-20 matched batch peaks at 0.857 GB allocated on GTX 1080 without task data | 5 → 5 | Removes implementation, graph, and memory objections before endpoint access; supplies no evidence yet that innovation is useful |
+| 2026-07-22 / `f310ad5` mixed-traffic MT-1 | All three frozen conditions become nonfinite in epoch 1 and end at 10%; ratio calibration, predictor warmup, zero-query, and cost checks pass | 5 → 5 | Fails to make innovation load-bearing on a standard ResNet and closes MT-2/MT-3; the prior mechanism-only evidence survives but empirical support does not improve |
Future rows are appended only after an audited frozen stage. A score staying flat is informative:
engineering, theory exposition, or visualization may make the paper more defensible without
diff --git a/ROADMAP.md b/ROADMAP.md
index 1bf6e32..c425356 100644
--- a/ROADMAP.md
+++ b/ROADMAP.md
@@ -481,8 +481,8 @@ zero task-loss queries, and cost remains `1.3261x` BP. This opens MT-1 without
raising the reviewer score because the result is entirely inherited KP
evidence.
-**KP mixed-traffic MT-0 status: mechanics passed; MT-1--MT-3 frozen before
-endpoints.** `MIXED_TRAFFIC.md` fixes a four-to-one, initialization-calibrated
+**KP mixed-traffic status: MT-0 passed; MT-1 failed.** `MIXED_TRAFFIC.md` fixes
+a four-to-one, initialization-calibrated
soma-predictable traffic intervention and crosses raw apical activity,
per-example norm-matched raw activity, and neutral-period innovation on the
same reciprocal KP substrate. The zero-traffic limit, exact-predictor limit,
@@ -494,12 +494,13 @@ MACs, and every control pays the predictor schedule. The synthetic batch-128
matched path is finite on a GTX 1080 and peaks at 0.857 GB allocated after
reset; no task endpoint was used for this hardware check.
-MT-1 is now open after the KP-2 pass. A 20-epoch pass opens one full seed-0
-validation panel but cannot raise the reviewer score; a full pass opens the
-already frozen five-seed, all-50k, one-test-evaluation confirmation. Only the
-complete MT-3 confirmation can move the score from 5/10 to 6/10. This is the
-current accept-bar path. It is a controlled predictable-traffic experiment,
-not a retroactive rescue of the failed natural/top-down C1 or Oral-B gates.
+The complete MT-1 panel then becomes nonfinite in epoch 1 under all three
+signals and ends at 10.00% validation accuracy. Calibration error remains only
+`4.77e-7`, predictor warmup reaches residual ratio `0.139942`, cost is
+`1.3271x` BP, and feedback uses zero task-loss queries, so those mechanical
+checks pass while all performance, finite, alignment, and tracking checks
+fail. Per the frozen stop rule, MT-2 and MT-3 remain untouched, no weaker
+traffic ratio is allowed, and this standard-ResNet accept path closes at 5/10.
Prepare convolutional local-update primitives and ResNet-20/32/56 protocols early. Queue frozen
runs opportunistically on authorized idle GPUs. Because BurstCCN already reports CIFAR-10 and