summaryrefslogtreecommitdiff
path: root/RESULTS.md
diff options
context:
space:
mode:
authorYurenHao0426 <Blackhao0426@gmail.com>2026-07-23 07:03:22 -0500
committerYurenHao0426 <Blackhao0426@gmail.com>2026-07-23 07:03:22 -0500
commitc8bfa25f597dc2760855b3f71a59549753c9e4a2 (patch)
treea8c1ecbad2a8d17e8c0651531519d76bb41b167e /RESULTS.md
parent03c94a1728d9a45f653a815a5f8d6eace9dd3d02 (diff)
docs: record oral B recovery boundary
Diffstat (limited to 'RESULTS.md')
-rw-r--r--RESULTS.md44
1 files changed, 44 insertions, 0 deletions
diff --git a/RESULTS.md b/RESULTS.md
index 436740f..c511d1e 100644
--- a/RESULTS.md
+++ b/RESULTS.md
@@ -184,6 +184,50 @@ learning, outcome decoding, and residual decorrelation do not imply the
cell-role-specific error-derivative signature observed by Harnett and
colleagues.
+## Temporal-difference oral-B recovery (R1/R2)
+
+The separately frozen recovery changes the structural target rather than
+reinterpreting the failed screen. It learns a per-cell causal-role coefficient
+from scalar antithetic cursor observations and multiplies that role by
+within-episode performance innovation
+`|e_(t-1)|-|e_t|`; `kappa=0`, so the branch tests plasticity rather than online
+control. R1 uses only development task seeds 0--2. The complete two-rate
+screen selects `eta=0.1`: all three seed-level ten-check gates pass, final
+success is `98.05%`, `98.83%`, and `99.61%`, sign inversion is
+`0.0651--0.0675`, and learned-role cosine is `0.945--0.973`. The lower
+`eta=0.03` rate retains the mechanism checks but fails final success at
+`36.72--46.09%`.
+
+R2 then evaluates the selected rate once on untouched task seeds 10--15 and
+model seeds 0--4, treating the six tasks--not the 30 models--as independent
+units. Its main task-cluster results are:
+
+| frozen R2 quantity | mean | one-sided 95% relevant bound | gate |
+|:--|--:|--:|:--:|
+| intact final success | 99.531% | lower 99.364% | pass |
+| early-to-late gain | 90.451 points | lower 89.222 | pass |
+| gap over fixed vectorizer | 99.531 points | lower 99.364 | pass |
+| learned-role cosine | 0.9395 | lower 0.9336 | pass |
+| absolute residual-soma correlation | 0.0609 | upper 0.0625 | pass |
+| raw-minus-residual correlation | 0.9365 | lower 0.9350 | pass |
+| surrounding-state balanced accuracy | 54.508% | lower 54.431% | **fail mean >=55%** |
+| decoder-distance residual correlation | 0.1559 | lower 0.1520 | pass |
+| residual outcome balanced accuracy | 47.325% | lower 43.067% | **fail** |
+| residual-minus-soma outcome accuracy | -4.322 points | lower -9.101 | **fail** |
+| causal-role sign inversion | 0.0712; 30/30 positive | lower 0.0703 | pass |
+| velocity-minus-error correlation advantage | 0.6389 | lower 0.6376 | pass |
+| early-residual/late-soma correlation | -0.0132 | lower -0.1786 | **fail** |
+
+All 30 records are finite and pass pairing, neutral-predictor, cursor-cost,
+role-learning, plasticity-lesion, oracle-ceiling, decorrelation, and
+provenance checks. However, the complete frozen R2 gate fails seven aggregate
+population-vectorization/longitudinal checks. It therefore does not establish
+oral-B innovation-guided plasticity under the predeclared joint claim, does
+not raise the reviewer score above 7, and does not unlock oral-A. The strongly
+positive learning, plasticity-lesion, sign-inversion, and velocity-dominance
+results remain useful mechanistic evidence, but they cannot be reported as a
+passed Harnett-signature panel.
+
## Frozen endogenous mixed-traffic confirmation (MNIST, d3/w256, 15 epochs)
The 95-run C1 panel crossed five paired model seeds, soma and endogenous top-down traffic, two