1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
|
# Oral-B-v2 cold-start recovery: dense innovation plus target psychometrics
This branch begins only after the complete, frozen
`oral_b_v2_development_v1` grid failed. That failure is retained at
`results/bci_v2_dev_gate.json`: none of 24 records was eligible, no candidate
exceeded 0.78% full-horizon evaluation success, and confirmation task seeds
`30--35` were not touched. The recovery does not edit that protocol, result,
or gate.
## Development-only diagnosis
The failure localized cleanly. Learned-role cosine was about `0.995`, the
neutral residual remained decorrelated from soma, and the surrounding-state
and velocity-over-error diagnostics passed, but even the exact-role oracle
could not learn. Before the first success, the critic can observe only the
dense performance-change term. Frozen v2 scaled that term by `0.25`, so the
largest actor rate was effectively one quarter of the already validated
temporal-difference rule and the terminal reward was unreachable.
On already used development task seed 20, a causal postmortem changed only
the dense velocity scale from `0.25` to `1.0`: final success changed from
about 1% to 100%. With critic enabled, success rose sharply on days 7--8;
without critic it did not rise until days 13--14. At scale 1, the diagnostic
also recovered learned-role cosine `0.983`, residual--soma correlation
`0.058`, surrounding decoder `54.4%`, velocity advantage `0.643`, and an
acute-outcome-lesion loss of `0.386` in role-aligned terminal separation.
These are exploratory postmortem values and are not R1 or R2 evidence.
That postmortem exposed a second structural issue: the learned fixed-target
policy reaches target `0.8` within four steps on every trial, so no
horizon-only assay can produce an unrewarded class. A development-only
uncensored probe placed the maximum cursor transition between targets 1.55
and 1.80. The recovery therefore replaces the failed horizon ladder with a
fixed target psychometric ladder. This is disclosed development tuning, not
claimed as preregistration against task seed 20.
## Frozen recovery
The mechanism, local updates, six paired conditions, neutral warmup, cost
accounting, full-horizon performance assay, and all other parameters remain
as in `ORAL_B_V2.md`, except:
```text
velocity_reward_scale = 1.0
forward_eta = 0.1
gamma = 0.8
critic_eta = 0.03
```
There is no recovery hyperparameter grid. Endpoint-free tests establish that
the nonterminal TD and actor update are exactly four times the failed v2
cold-start drive when critic and terminal reward are absent. They also bind
the ordered target ladder.
The signature assay uses six targets
`{1.55, 1.60, 1.65, 1.70, 1.75, 1.80}`, 128 independent full-28-step episodes
per target, and no target selection or reweighting. All targets are pooled.
The same trained intact model is replayed under intact, acute critic lesion,
and acute terminal-outcome lesion. This is a psychometric generalization
assay: ordinary performance remains measured at the trained target `0.8`.
The target ladder is intended to create rewarded and censored timeout trials
without defining outcome directly from a balanced random label.
## R1 independent development validation
R1 uses previously untouched development-validation task seeds `23--25` and
model seed 0. Full performance seed is `460000 + task_seed`; target-ladder
seed for zero-based target index `k` is
`470000 + 1000 task_seed + k`. There is one fixed candidate and exactly three
records.
Every seed must pass every check:
- intact final success at least 70%, early-to-late gain at least 10 points,
fixed-role gap at least 20 points, oracle deficit at most 10 points,
plasticity lesion at most half the intact gain, and role cosine at least
`0.80`;
- pooled target-ladder success fraction in `[0.15, 0.85]`;
- terminal residual outcome balanced accuracy at least 75%, role-aligned
outcome separation at least `0.20`, acute outcome lesion separation drop at
least `0.20`, critic expectedness contribution at least `0.05`, and paired
critic contribution/value correlation at least `0.95`;
- residual--soma correlation at most `0.10`, raw-minus-residual correlation
gap at least `0.20`, surrounding-event accuracy at least 52%, and
decoder-distance correlation at least `0.02`; and
- sign inversion at least `0.01` and velocity-over-error advantage at least
`0.05`.
R1 cannot change the score. Any failure closes this recovery and leaves
confirmation untouched.
## R2 untouched confirmation
Only a complete R1 pass permits the fixed candidate on task seeds `30--35`
and model seeds `0--4`, exactly 30 records with no replacement or further
selection. Full performance seed is `520000 + task_seed`; target-ladder seed
is `530000 + 1000 task_seed + k`. The task seed is the uncertainty unit: the
five models are averaged within each task, then one-sided 95% Student-t
bounds use five degrees of freedom.
Learning/plasticity requirements are unchanged from the prior frozen R2:
mean final at least 70%, every task mean and lower bound at least 60%; mean
gain at least 10 points and lower bound 5; mean fixed-role gap at least 20
points and lower bound 10; mean oracle deficit at most 10 points and upper
bound 20; nonnegative lower bound for
`0.5 intact_gain - lesion_gain`; mean role cosine at least `0.80` and every
record at least `0.70`.
Innovation requirements are: mean residual--soma correlation at most `0.10`
with upper bound `0.12`; raw-minus-residual gap at least `0.20` with lower
bound `0.15`; surrounding accuracy at least 52% with lower bound 50%;
decoder-distance correlation at least `0.05` with nonnegative lower bound;
positive sign in at least 25/30 records and every task mean; velocity
advantage at least `0.05` with nonnegative lower bound.
Outcome-surprise requirements are: every pooled success fraction in
`[0.10, 0.90]`; terminal accuracy mean at least 80% and lower bound 75%;
role-aligned separation mean at least `0.20` and lower bound `0.15`; acute
outcome-lesion drop mean at least `0.20` and lower bound `0.15`; critic
expectedness mean at least `0.05` and lower bound `0.02`; every paired
critic/value correlation at least `0.95`.
All records require a clean tracked tree, exact filenames, and SHA-256 binding
to both preserved failure gates, the D4 gate, protocol, dynamics, metrics,
runners, analyzers, and R1 gate where applicable. A complete R2 pass raises
the formal repository milestone from 7 to 8 and permits a separately frozen
oral-A-v2 protocol. It does not erase either failure, open the old oral-A
gate, or establish online apical control.
|