From c8bfa25f597dc2760855b3f71a59549753c9e4a2 Mon Sep 17 00:00:00 2001 From: YurenHao0426 Date: Thu, 23 Jul 2026 07:03:22 -0500 Subject: docs: record oral B recovery boundary --- RESULTS.md | 44 ++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 44 insertions(+) (limited to 'RESULTS.md') diff --git a/RESULTS.md b/RESULTS.md index 436740f..c511d1e 100644 --- a/RESULTS.md +++ b/RESULTS.md @@ -184,6 +184,50 @@ learning, outcome decoding, and residual decorrelation do not imply the cell-role-specific error-derivative signature observed by Harnett and colleagues. +## Temporal-difference oral-B recovery (R1/R2) + +The separately frozen recovery changes the structural target rather than +reinterpreting the failed screen. It learns a per-cell causal-role coefficient +from scalar antithetic cursor observations and multiplies that role by +within-episode performance innovation +`|e_(t-1)|-|e_t|`; `kappa=0`, so the branch tests plasticity rather than online +control. R1 uses only development task seeds 0--2. The complete two-rate +screen selects `eta=0.1`: all three seed-level ten-check gates pass, final +success is `98.05%`, `98.83%`, and `99.61%`, sign inversion is +`0.0651--0.0675`, and learned-role cosine is `0.945--0.973`. The lower +`eta=0.03` rate retains the mechanism checks but fails final success at +`36.72--46.09%`. + +R2 then evaluates the selected rate once on untouched task seeds 10--15 and +model seeds 0--4, treating the six tasks--not the 30 models--as independent +units. Its main task-cluster results are: + +| frozen R2 quantity | mean | one-sided 95% relevant bound | gate | +|:--|--:|--:|:--:| +| intact final success | 99.531% | lower 99.364% | pass | +| early-to-late gain | 90.451 points | lower 89.222 | pass | +| gap over fixed vectorizer | 99.531 points | lower 99.364 | pass | +| learned-role cosine | 0.9395 | lower 0.9336 | pass | +| absolute residual-soma correlation | 0.0609 | upper 0.0625 | pass | +| raw-minus-residual correlation | 0.9365 | lower 0.9350 | pass | +| surrounding-state balanced accuracy | 54.508% | lower 54.431% | **fail mean >=55%** | +| decoder-distance residual correlation | 0.1559 | lower 0.1520 | pass | +| residual outcome balanced accuracy | 47.325% | lower 43.067% | **fail** | +| residual-minus-soma outcome accuracy | -4.322 points | lower -9.101 | **fail** | +| causal-role sign inversion | 0.0712; 30/30 positive | lower 0.0703 | pass | +| velocity-minus-error correlation advantage | 0.6389 | lower 0.6376 | pass | +| early-residual/late-soma correlation | -0.0132 | lower -0.1786 | **fail** | + +All 30 records are finite and pass pairing, neutral-predictor, cursor-cost, +role-learning, plasticity-lesion, oracle-ceiling, decorrelation, and +provenance checks. However, the complete frozen R2 gate fails seven aggregate +population-vectorization/longitudinal checks. It therefore does not establish +oral-B innovation-guided plasticity under the predeclared joint claim, does +not raise the reviewer score above 7, and does not unlock oral-A. The strongly +positive learning, plasticity-lesion, sign-inversion, and velocity-dominance +results remain useful mechanistic evidence, but they cannot be reported as a +passed Harnett-signature panel. + ## Frozen endogenous mixed-traffic confirmation (MNIST, d3/w256, 15 epochs) The 95-run C1 panel crossed five paired model seeds, soma and endogenous top-down traffic, two -- cgit v1.2.3