diff options
| author | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-23 07:03:22 -0500 |
|---|---|---|
| committer | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-23 07:03:22 -0500 |
| commit | c8bfa25f597dc2760855b3f71a59549753c9e4a2 (patch) | |
| tree | a8c1ecbad2a8d17e8c0651531519d76bb41b167e /RESULTS.md | |
| parent | 03c94a1728d9a45f653a815a5f8d6eace9dd3d02 (diff) | |
docs: record oral B recovery boundary
Diffstat (limited to 'RESULTS.md')
| -rw-r--r-- | RESULTS.md | 44 |
1 files changed, 44 insertions, 0 deletions
@@ -184,6 +184,50 @@ learning, outcome decoding, and residual decorrelation do not imply the cell-role-specific error-derivative signature observed by Harnett and colleagues. +## Temporal-difference oral-B recovery (R1/R2) + +The separately frozen recovery changes the structural target rather than +reinterpreting the failed screen. It learns a per-cell causal-role coefficient +from scalar antithetic cursor observations and multiplies that role by +within-episode performance innovation +`|e_(t-1)|-|e_t|`; `kappa=0`, so the branch tests plasticity rather than online +control. R1 uses only development task seeds 0--2. The complete two-rate +screen selects `eta=0.1`: all three seed-level ten-check gates pass, final +success is `98.05%`, `98.83%`, and `99.61%`, sign inversion is +`0.0651--0.0675`, and learned-role cosine is `0.945--0.973`. The lower +`eta=0.03` rate retains the mechanism checks but fails final success at +`36.72--46.09%`. + +R2 then evaluates the selected rate once on untouched task seeds 10--15 and +model seeds 0--4, treating the six tasks--not the 30 models--as independent +units. Its main task-cluster results are: + +| frozen R2 quantity | mean | one-sided 95% relevant bound | gate | +|:--|--:|--:|:--:| +| intact final success | 99.531% | lower 99.364% | pass | +| early-to-late gain | 90.451 points | lower 89.222 | pass | +| gap over fixed vectorizer | 99.531 points | lower 99.364 | pass | +| learned-role cosine | 0.9395 | lower 0.9336 | pass | +| absolute residual-soma correlation | 0.0609 | upper 0.0625 | pass | +| raw-minus-residual correlation | 0.9365 | lower 0.9350 | pass | +| surrounding-state balanced accuracy | 54.508% | lower 54.431% | **fail mean >=55%** | +| decoder-distance residual correlation | 0.1559 | lower 0.1520 | pass | +| residual outcome balanced accuracy | 47.325% | lower 43.067% | **fail** | +| residual-minus-soma outcome accuracy | -4.322 points | lower -9.101 | **fail** | +| causal-role sign inversion | 0.0712; 30/30 positive | lower 0.0703 | pass | +| velocity-minus-error correlation advantage | 0.6389 | lower 0.6376 | pass | +| early-residual/late-soma correlation | -0.0132 | lower -0.1786 | **fail** | + +All 30 records are finite and pass pairing, neutral-predictor, cursor-cost, +role-learning, plasticity-lesion, oracle-ceiling, decorrelation, and +provenance checks. However, the complete frozen R2 gate fails seven aggregate +population-vectorization/longitudinal checks. It therefore does not establish +oral-B innovation-guided plasticity under the predeclared joint claim, does +not raise the reviewer score above 7, and does not unlock oral-A. The strongly +positive learning, plasticity-lesion, sign-inversion, and velocity-dominance +results remain useful mechanistic evidence, but they cannot be reported as a +passed Harnett-signature panel. + ## Frozen endogenous mixed-traffic confirmation (MNIST, d3/w256, 15 epochs) The 95-run C1 panel crossed five paired model seeds, soma and endogenous top-down traffic, two |
