summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
authorYurenHao0426 <Blackhao0426@gmail.com>2026-07-23 08:07:50 -0500
committerYurenHao0426 <Blackhao0426@gmail.com>2026-07-23 08:07:50 -0500
commitdae172b1425c646f0441609c28fc5156b50f820d (patch)
tree8227325805d28f854cc2606395301268882bb88e
parenteb021a6beac9d4b55a595171f149d669868c403b (diff)
protocol: freeze oral-B-v2 cold-start recovery
-rw-r--r--ORAL_B_V2_RECOVERY.md128
-rw-r--r--experiments/analyze_bci_v2_recovery_confirmation.py498
-rw-r--r--experiments/analyze_bci_v2_recovery_development.py284
-rw-r--r--experiments/bci_v2_confirmation.sh2
-rw-r--r--experiments/bci_v2_development_screen.sh2
-rw-r--r--experiments/bci_v2_recovery_confirmation.py141
-rw-r--r--experiments/bci_v2_recovery_confirmation.sh16
-rw-r--r--experiments/bci_v2_recovery_development.sh12
-rw-r--r--experiments/bci_v2_recovery_run.py377
-rw-r--r--experiments/bci_v2_recovery_smoke.py104
-rw-r--r--sdil/bci_v2_recovery_metrics.py44
11 files changed, 1606 insertions, 2 deletions
diff --git a/ORAL_B_V2_RECOVERY.md b/ORAL_B_V2_RECOVERY.md
new file mode 100644
index 0000000..d929755
--- /dev/null
+++ b/ORAL_B_V2_RECOVERY.md
@@ -0,0 +1,128 @@
+# Oral-B-v2 cold-start recovery: dense innovation plus target psychometrics
+
+This branch begins only after the complete, frozen
+`oral_b_v2_development_v1` grid failed. That failure is retained at
+`results/bci_v2_dev_gate.json`: none of 24 records was eligible, no candidate
+exceeded 0.78% full-horizon evaluation success, and confirmation task seeds
+`30--35` were not touched. The recovery does not edit that protocol, result,
+or gate.
+
+## Development-only diagnosis
+
+The failure localized cleanly. Learned-role cosine was about `0.995`, the
+neutral residual remained decorrelated from soma, and the surrounding-state
+and velocity-over-error diagnostics passed, but even the exact-role oracle
+could not learn. Before the first success, the critic can observe only the
+dense performance-change term. Frozen v2 scaled that term by `0.25`, so the
+largest actor rate was effectively one quarter of the already validated
+temporal-difference rule and the terminal reward was unreachable.
+
+On already used development task seed 20, a causal postmortem changed only
+the dense velocity scale from `0.25` to `1.0`: final success changed from
+about 1% to 100%. With critic enabled, success rose sharply on days 7--8;
+without critic it did not rise until days 13--14. At scale 1, the diagnostic
+also recovered learned-role cosine `0.983`, residual--soma correlation
+`0.058`, surrounding decoder `54.4%`, velocity advantage `0.643`, and an
+acute-outcome-lesion loss of `0.386` in role-aligned terminal separation.
+These are exploratory postmortem values and are not R1 or R2 evidence.
+
+That postmortem exposed a second structural issue: the learned fixed-target
+policy reaches target `0.8` within four steps on every trial, so no
+horizon-only assay can produce an unrewarded class. A development-only
+uncensored probe placed the maximum cursor transition between targets 1.55
+and 1.80. The recovery therefore replaces the failed horizon ladder with a
+fixed target psychometric ladder. This is disclosed development tuning, not
+claimed as preregistration against task seed 20.
+
+## Frozen recovery
+
+The mechanism, local updates, six paired conditions, neutral warmup, cost
+accounting, full-horizon performance assay, and all other parameters remain
+as in `ORAL_B_V2.md`, except:
+
+```text
+velocity_reward_scale = 1.0
+forward_eta = 0.1
+gamma = 0.8
+critic_eta = 0.03
+```
+
+There is no recovery hyperparameter grid. Endpoint-free tests establish that
+the nonterminal TD and actor update are exactly four times the failed v2
+cold-start drive when critic and terminal reward are absent. They also bind
+the ordered target ladder.
+
+The signature assay uses six targets
+`{1.55, 1.60, 1.65, 1.70, 1.75, 1.80}`, 128 independent full-28-step episodes
+per target, and no target selection or reweighting. All targets are pooled.
+The same trained intact model is replayed under intact, acute critic lesion,
+and acute terminal-outcome lesion. This is a psychometric generalization
+assay: ordinary performance remains measured at the trained target `0.8`.
+The target ladder is intended to create rewarded and censored timeout trials
+without defining outcome directly from a balanced random label.
+
+## R1 independent development validation
+
+R1 uses previously untouched development-validation task seeds `23--25` and
+model seed 0. Full performance seed is `460000 + task_seed`; target-ladder
+seed for zero-based target index `k` is
+`470000 + 1000 task_seed + k`. There is one fixed candidate and exactly three
+records.
+
+Every seed must pass every check:
+
+- intact final success at least 70%, early-to-late gain at least 10 points,
+ fixed-role gap at least 20 points, oracle deficit at most 10 points,
+ plasticity lesion at most half the intact gain, and role cosine at least
+ `0.80`;
+- pooled target-ladder success fraction in `[0.15, 0.85]`;
+- terminal residual outcome balanced accuracy at least 75%, role-aligned
+ outcome separation at least `0.20`, acute outcome lesion separation drop at
+ least `0.20`, critic expectedness contribution at least `0.05`, and paired
+ critic contribution/value correlation at least `0.95`;
+- residual--soma correlation at most `0.10`, raw-minus-residual correlation
+ gap at least `0.20`, surrounding-event accuracy at least 52%, and
+ decoder-distance correlation at least `0.02`; and
+- sign inversion at least `0.01` and velocity-over-error advantage at least
+ `0.05`.
+
+R1 cannot change the score. Any failure closes this recovery and leaves
+confirmation untouched.
+
+## R2 untouched confirmation
+
+Only a complete R1 pass permits the fixed candidate on task seeds `30--35`
+and model seeds `0--4`, exactly 30 records with no replacement or further
+selection. Full performance seed is `520000 + task_seed`; target-ladder seed
+is `530000 + 1000 task_seed + k`. The task seed is the uncertainty unit: the
+five models are averaged within each task, then one-sided 95% Student-t
+bounds use five degrees of freedom.
+
+Learning/plasticity requirements are unchanged from the prior frozen R2:
+mean final at least 70%, every task mean and lower bound at least 60%; mean
+gain at least 10 points and lower bound 5; mean fixed-role gap at least 20
+points and lower bound 10; mean oracle deficit at most 10 points and upper
+bound 20; nonnegative lower bound for
+`0.5 intact_gain - lesion_gain`; mean role cosine at least `0.80` and every
+record at least `0.70`.
+
+Innovation requirements are: mean residual--soma correlation at most `0.10`
+with upper bound `0.12`; raw-minus-residual gap at least `0.20` with lower
+bound `0.15`; surrounding accuracy at least 52% with lower bound 50%;
+decoder-distance correlation at least `0.05` with nonnegative lower bound;
+positive sign in at least 25/30 records and every task mean; velocity
+advantage at least `0.05` with nonnegative lower bound.
+
+Outcome-surprise requirements are: every pooled success fraction in
+`[0.10, 0.90]`; terminal accuracy mean at least 80% and lower bound 75%;
+role-aligned separation mean at least `0.20` and lower bound `0.15`; acute
+outcome-lesion drop mean at least `0.20` and lower bound `0.15`; critic
+expectedness mean at least `0.05` and lower bound `0.02`; every paired
+critic/value correlation at least `0.95`.
+
+All records require a clean tracked tree, exact filenames, and SHA-256 binding
+to both preserved failure gates, the D4 gate, protocol, dynamics, metrics,
+runners, analyzers, and R1 gate where applicable. A complete R2 pass raises
+the formal repository milestone from 7 to 8 and permits a separately frozen
+oral-A-v2 protocol. It does not erase either failure, open the old oral-A
+gate, or establish online apical control.
diff --git a/experiments/analyze_bci_v2_recovery_confirmation.py b/experiments/analyze_bci_v2_recovery_confirmation.py
new file mode 100644
index 0000000..94143be
--- /dev/null
+++ b/experiments/analyze_bci_v2_recovery_confirmation.py
@@ -0,0 +1,498 @@
+#!/usr/bin/env python3
+"""Audit untouched oral-B-v2 cold-start recovery confirmation."""
+import argparse
+import glob
+import hashlib
+import json
+import math
+import os
+import statistics
+
+from experiments.bci_v2_recovery_confirmation import (
+ MODEL_SEEDS,
+ TASK_SEEDS,
+)
+from experiments.bci_v2_recovery_run import (
+ CHALLENGE_EPISODES,
+ FIXED_CONFIG,
+ TARGETS,
+ build_recovery_config,
+ source_paths,
+)
+from experiments.bci_v2_run import CONDITIONS
+
+
+ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
+RUNNER_PATH = os.path.join(
+ ROOT, "experiments", "bci_v2_recovery_confirmation.py"
+)
+T_CRITICAL_ONE_SIDED_95_DF5 = 2.015048373
+
+
+def require(condition, message):
+ if not condition:
+ raise ValueError(message)
+
+
+def sha256(path):
+ digest = hashlib.sha256()
+ with open(path, "rb") as handle:
+ for chunk in iter(lambda: handle.read(1024 * 1024), b""):
+ digest.update(chunk)
+ return digest.hexdigest()
+
+
+def finite_tree(value):
+ if isinstance(value, dict):
+ return all(finite_tree(item) for item in value.values())
+ if isinstance(value, (list, tuple)):
+ return all(finite_tree(item) for item in value)
+ if isinstance(value, (int, float)):
+ return math.isfinite(value)
+ return True
+
+
+def lower_bound(values):
+ return (
+ statistics.mean(values)
+ - T_CRITICAL_ONE_SIDED_95_DF5
+ * statistics.stdev(values)
+ / math.sqrt(len(values))
+ )
+
+
+def upper_bound(values):
+ return (
+ statistics.mean(values)
+ + T_CRITICAL_ONE_SIDED_95_DF5
+ * statistics.stdev(values)
+ / math.sqrt(len(values))
+ )
+
+
+def summarize(values):
+ return {
+ "by_task_seed": values,
+ "mean": statistics.mean(values),
+ "one_sided_95pct_lower": lower_bound(values),
+ "one_sided_95pct_upper": upper_bound(values),
+ }
+
+
+def task_clusters(records, getter):
+ return [
+ statistics.mean(
+ getter(records[(task_seed, model_seed)])
+ for model_seed in MODEL_SEEDS
+ )
+ for task_seed in TASK_SEEDS
+ ]
+
+
+def confirmation_paths(r1_gate):
+ paths = source_paths()
+ paths["development_runner"] = paths["runner"]
+ paths["runner"] = RUNNER_PATH
+ paths["r1_gate"] = os.path.abspath(r1_gate)
+ return paths
+
+
+def validate(row, path, digests):
+ require(row.get("schema_version") == 3, f"{path}: schema")
+ args = row.get("args", {})
+ task_seed = args.get("task_seed")
+ model_seed = args.get("model_seed")
+ require(
+ task_seed in TASK_SEEDS and model_seed in MODEL_SEEDS,
+ f"{path}: seeds",
+ )
+ protocol = row.get("protocol", {})
+ require(
+ protocol.get("name")
+ == "oral_b_v2_cold_start_recovery_confirmation_v1"
+ and protocol.get("split") == "untouched_confirmation"
+ and protocol.get("training_task_seed") == task_seed
+ and protocol.get("model_seed") == model_seed
+ and protocol.get("fixed_config") == FIXED_CONFIG
+ and protocol.get("no_further_selection") is True
+ and protocol.get("confirmation_grid_size") == 30
+ and protocol.get("protocol_sha256") == digests["protocol"]
+ and protocol.get("r1_gate_sha256") == digests["r1_gate"],
+ f"{path}: protocol",
+ )
+ provenance = row.get("provenance", {})
+ require(
+ provenance.get("git_tracked_dirty") is False
+ and provenance.get("tracked_inputs")
+ and all(provenance["tracked_inputs"].values())
+ and set(provenance["tracked_inputs"]) == set(digests)
+ and provenance.get("input_sha256") == digests,
+ f"{path}: provenance",
+ )
+ require(
+ row.get("config") == vars(build_recovery_config()),
+ f"{path}: config",
+ )
+ require(
+ row.get("split") == "untouched_confirmation"
+ and row.get("finite") is True
+ and finite_tree(row),
+ f"{path}: finite/split",
+ )
+ require(
+ set(row.get("conditions", {})) == set(CONDITIONS),
+ f"{path}: conditions",
+ )
+ for name in CONDITIONS:
+ warmup = row["warmup"][name]
+ condition = row["conditions"][name]
+ cost = condition["cost"]
+ require(
+ warmup["batches"] == 100
+ and warmup["examples"] == 6400
+ and warmup["instruction_present"] is False
+ and warmup["role_cursor_scalar_observations"] == 12800
+ and warmup["predictor_max_abs_error"] <= 1e-5
+ and len(condition["daily_success"]) == 14
+ and cost["maximum_state_episode_steps"] == 25088
+ and cost["cursor_scalar_observations"]
+ == 2 * cost["role_probe_examples"]
+ and cost["task_loss_queries"] == 0
+ and cost["reverse_mode_calls"] == 0,
+ f"{path}: condition {name}",
+ )
+ challenge = row["assays"]["challenge"]
+ require(
+ challenge["targets"] == list(TARGETS)
+ and challenge["episodes_per_target"] == CHALLENGE_EPISODES
+ and challenge["selection_over_targets"] is False
+ and challenge["maximum_steps_per_episode"] == 28
+ and row["signatures"]["challenge_episodes"]
+ == len(TARGETS) * CHALLENGE_EPISODES,
+ f"{path}: target ladder",
+ )
+
+
+def checks(metrics, clustered, all_values):
+ return {
+ "all_records_finite_paired_and_cost_audited": True,
+ "learning_and_plasticity": {
+ "mean_intact_final_at_least_0p70":
+ metrics["intact_final"]["mean"] >= 0.70,
+ "every_task_mean_final_at_least_0p60":
+ min(clustered["intact_final"]) >= 0.60,
+ "intact_final_lower_bound_at_least_0p60":
+ metrics["intact_final"][
+ "one_sided_95pct_lower"
+ ] >= 0.60,
+ "mean_learning_gain_at_least_0p10":
+ metrics["intact_gain"]["mean"] >= 0.10,
+ "learning_gain_lower_bound_at_least_0p05":
+ metrics["intact_gain"][
+ "one_sided_95pct_lower"
+ ] >= 0.05,
+ "mean_fixed_role_gap_at_least_0p20":
+ metrics["fixed_role_gap"]["mean"] >= 0.20,
+ "fixed_role_gap_lower_bound_at_least_0p10":
+ metrics["fixed_role_gap"][
+ "one_sided_95pct_lower"
+ ] >= 0.10,
+ "mean_oracle_deficit_at_most_0p10":
+ metrics["oracle_deficit"]["mean"] <= 0.10,
+ "oracle_deficit_upper_bound_at_most_0p20":
+ metrics["oracle_deficit"][
+ "one_sided_95pct_upper"
+ ] <= 0.20,
+ "plasticity_half_margin_lower_bound_nonnegative":
+ metrics["plasticity_half_margin"][
+ "one_sided_95pct_lower"
+ ] >= 0.0,
+ "mean_role_cosine_at_least_0p80":
+ metrics["role_cosine"]["mean"] >= 0.80,
+ "every_role_cosine_at_least_0p70":
+ min(all_values["role_cosine"]) >= 0.70,
+ },
+ "innovation_and_network_prediction": {
+ "mean_residual_soma_corr_at_most_0p10":
+ metrics["residual_soma_corr"]["mean"] <= 0.10,
+ "residual_soma_corr_upper_bound_at_most_0p12":
+ metrics["residual_soma_corr"][
+ "one_sided_95pct_upper"
+ ] <= 0.12,
+ "mean_raw_residual_corr_gap_at_least_0p20":
+ metrics["raw_residual_corr_gap"]["mean"] >= 0.20,
+ "raw_residual_corr_gap_lower_bound_at_least_0p15":
+ metrics["raw_residual_corr_gap"][
+ "one_sided_95pct_lower"
+ ] >= 0.15,
+ "mean_surrounding_accuracy_at_least_0p52":
+ metrics["surrounding_accuracy"]["mean"] >= 0.52,
+ "surrounding_accuracy_lower_bound_at_least_0p50":
+ metrics["surrounding_accuracy"][
+ "one_sided_95pct_lower"
+ ] >= 0.50,
+ "mean_decoder_corr_at_least_0p05":
+ metrics["decoder_corr"]["mean"] >= 0.05,
+ "decoder_corr_lower_bound_nonnegative":
+ metrics["decoder_corr"][
+ "one_sided_95pct_lower"
+ ] >= 0.0,
+ "positive_sign_in_at_least_25_of_30":
+ sum(
+ value > 0 for value in all_values["sign_inversion"]
+ ) >= 25,
+ "positive_sign_in_every_task_cluster":
+ min(clustered["sign_inversion"]) > 0.0,
+ "mean_velocity_advantage_at_least_0p05":
+ metrics["velocity_advantage"]["mean"] >= 0.05,
+ "velocity_advantage_lower_bound_nonnegative":
+ metrics["velocity_advantage"][
+ "one_sided_95pct_lower"
+ ] >= 0.0,
+ },
+ "target_ladder_outcome_surprise": {
+ "every_challenge_fraction_between_0p10_and_0p90":
+ min(all_values["challenge_fraction"]) >= 0.10
+ and max(all_values["challenge_fraction"]) <= 0.90,
+ "mean_terminal_accuracy_at_least_0p80":
+ metrics["terminal_accuracy"]["mean"] >= 0.80,
+ "terminal_accuracy_lower_bound_at_least_0p75":
+ metrics["terminal_accuracy"][
+ "one_sided_95pct_lower"
+ ] >= 0.75,
+ "mean_terminal_separation_at_least_0p20":
+ metrics["terminal_separation"]["mean"] >= 0.20,
+ "terminal_separation_lower_bound_at_least_0p15":
+ metrics["terminal_separation"][
+ "one_sided_95pct_lower"
+ ] >= 0.15,
+ "mean_outcome_lesion_drop_at_least_0p20":
+ metrics["outcome_lesion_drop"]["mean"] >= 0.20,
+ "outcome_lesion_drop_lower_bound_at_least_0p15":
+ metrics["outcome_lesion_drop"][
+ "one_sided_95pct_lower"
+ ] >= 0.15,
+ "mean_critic_expectedness_at_least_0p05":
+ metrics["critic_expectedness"]["mean"] >= 0.05,
+ "critic_expectedness_lower_bound_at_least_0p02":
+ metrics["critic_expectedness"][
+ "one_sided_95pct_lower"
+ ] >= 0.02,
+ "every_critic_value_corr_at_least_0p95":
+ min(all_values["critic_value_corr"]) >= 0.95,
+ },
+ }
+
+
+def main():
+ parser = argparse.ArgumentParser()
+ parser.add_argument(
+ "--results",
+ default="results/bci_v2_recovery_confirmation",
+ )
+ parser.add_argument(
+ "--r1-gate",
+ default="results/bci_v2_recovery_dev_gate.json",
+ )
+ parser.add_argument(
+ "--out",
+ default="results/bci_v2_recovery_confirmation_gate.json",
+ )
+ args = parser.parse_args()
+ with open(args.r1_gate) as handle:
+ r1 = json.load(handle)
+ require(
+ r1.get("protocol")
+ == "oral_b_v2_cold_start_recovery_development_v1"
+ and r1.get("status") == "passed"
+ and r1.get("complete_grid") is True
+ and r1.get("fixed_config") == FIXED_CONFIG
+ and r1.get("confirmation_seeds_touched") is False
+ and r1.get("recovery_confirmation_opened") is True,
+ "recovery R1 gate",
+ )
+ development_digests = {
+ name: sha256(path)
+ for name, path in source_paths().items()
+ }
+ require(
+ r1.get("input_sha256") == development_digests,
+ "R1 source binding",
+ )
+ digests = {
+ name: sha256(path)
+ for name, path in confirmation_paths(args.r1_gate).items()
+ }
+ expected = {
+ f"bci_v2_recovery_confirm_t{task_seed}_m{model_seed}.json"
+ for task_seed in TASK_SEEDS
+ for model_seed in MODEL_SEEDS
+ }
+ observed = {
+ os.path.basename(path)
+ for path in glob.glob(os.path.join(args.results, "*.json"))
+ }
+ require(
+ observed == expected,
+ (
+ f"confirmation drift: missing={sorted(expected-observed)}, "
+ f"extra={sorted(observed-expected)}"
+ ),
+ )
+ records = {}
+ commits = set()
+ source_sha256 = {}
+ for task_seed in TASK_SEEDS:
+ for model_seed in MODEL_SEEDS:
+ path = os.path.join(
+ args.results,
+ (
+ f"bci_v2_recovery_confirm_t{task_seed}"
+ f"_m{model_seed}.json"
+ ),
+ )
+ with open(path) as handle:
+ row = json.load(handle)
+ validate(row, path, digests)
+ records[(task_seed, model_seed)] = row
+ commits.add(row["provenance"]["git_commit"])
+ source_sha256[path] = sha256(path)
+ require(len(commits) == 1, "confirmation source revision drift")
+
+ getters = {
+ "intact_final": lambda row:
+ row["conditions"]["intact"]["final_success"],
+ "intact_gain": lambda row:
+ row["conditions"]["intact"]["learning_gain"],
+ "fixed_role_gap": lambda row: (
+ row["conditions"]["intact"]["final_success"]
+ - row["conditions"]["fixed_role"]["final_success"]
+ ),
+ "oracle_deficit": lambda row: (
+ row["conditions"]["oracle_role"]["final_success"]
+ - row["conditions"]["intact"]["final_success"]
+ ),
+ "plasticity_half_margin": lambda row: (
+ 0.5 * row["conditions"]["intact"]["learning_gain"]
+ - row["conditions"]["plasticity_lesion"]["learning_gain"]
+ ),
+ "role_cosine": lambda row:
+ row["conditions"]["intact"]["role_cosine_after_training"],
+ "critic_training_gap": lambda row: (
+ row["conditions"]["intact"]["final_success"]
+ - row["conditions"]["critic_training_lesion"][
+ "final_success"
+ ]
+ ),
+ "outcome_training_gap": lambda row: (
+ row["conditions"]["intact"]["final_success"]
+ - row["conditions"]["outcome_training_lesion"][
+ "final_success"
+ ]
+ ),
+ "residual_soma_corr": lambda row:
+ row["signatures"]["mean_abs_residual_soma_corr"],
+ "raw_residual_corr_gap": lambda row:
+ row["signatures"][
+ "raw_minus_residual_abs_soma_corr"
+ ],
+ "surrounding_accuracy": lambda row:
+ row["signatures"][
+ "surrounding_event_decoder_balanced_acc"
+ ],
+ "decoder_corr": lambda row:
+ row["signatures"]["decoder_distance_residual_corr"],
+ "sign_inversion": lambda row:
+ row["signatures"][
+ "causal_role_sign_inversion_index"
+ ],
+ "velocity_advantage": lambda row:
+ row["signatures"][
+ "velocity_minus_error_abs_cv_corr"
+ ],
+ "challenge_fraction": lambda row:
+ row["signatures"]["challenge_success_fraction"],
+ "terminal_accuracy": lambda row:
+ row["signatures"][
+ "terminal_residual_outcome_balanced_acc"
+ ],
+ "terminal_separation": lambda row:
+ row["signatures"][
+ "terminal_role_aligned_outcome_separation"
+ ],
+ "outcome_lesion_drop": lambda row:
+ row["signatures"][
+ "terminal_outcome_separation_drop_under_acute_lesion"
+ ],
+ "critic_expectedness": lambda row:
+ row["signatures"][
+ "mean_critic_expectedness_contribution"
+ ],
+ "critic_value_corr": lambda row:
+ row["signatures"][
+ "critic_contribution_value_prediction_corr"
+ ],
+ }
+ clustered = {
+ name: task_clusters(records, getter)
+ for name, getter in getters.items()
+ }
+ metrics = {
+ name: summarize(values)
+ for name, values in clustered.items()
+ }
+ all_values = {
+ name: [getter(row) for row in records.values()]
+ for name, getter in getters.items()
+ }
+ gate_checks = checks(metrics, clustered, all_values)
+ passed = all(
+ value
+ for category in gate_checks.values()
+ for value in (
+ [category]
+ if isinstance(category, bool)
+ else category.values()
+ )
+ )
+ output = {
+ "protocol":
+ "oral_b_v2_cold_start_recovery_confirmation_v1",
+ "status": "passed" if passed else "failed",
+ "complete_grid": True,
+ "fixed_config": FIXED_CONFIG,
+ "checks": gate_checks,
+ "metrics": metrics,
+ "positive_sign_count": sum(
+ value > 0 for value in all_values["sign_inversion"]
+ ),
+ "source_commit": next(iter(commits)),
+ "source_sha256": source_sha256,
+ "input_sha256": digests,
+ "oral_b_v2_outcome_surprise_established": passed,
+ "old_r2_status_preserved": "failed",
+ "failed_v2_development_status_preserved": "failed",
+ "old_oral_a_gate_remains_closed": True,
+ "new_oral_a_v2_protocol_may_be_frozen": passed,
+ "review_score_before": 7,
+ "review_score_after": 8 if passed else 7,
+ "score_change_rule": (
+ "only a complete untouched recovery R2 pass establishes "
+ "role-vectorized TD outcome surprise; both prior failures "
+ "and the old oral-A gate remain unchanged"
+ ),
+ }
+ os.makedirs(os.path.dirname(os.path.abspath(args.out)), exist_ok=True)
+ if os.path.exists(args.out):
+ with open(args.out) as handle:
+ existing = json.load(handle)
+ require(existing == output, "existing recovery R2 gate differs")
+ else:
+ with open(args.out, "w") as handle:
+ json.dump(output, handle, indent=2, sort_keys=True)
+ handle.write("\n")
+ print(json.dumps(output, indent=2))
+
+
+if __name__ == "__main__":
+ main()
diff --git a/experiments/analyze_bci_v2_recovery_development.py b/experiments/analyze_bci_v2_recovery_development.py
new file mode 100644
index 0000000..a38582b
--- /dev/null
+++ b/experiments/analyze_bci_v2_recovery_development.py
@@ -0,0 +1,284 @@
+#!/usr/bin/env python3
+"""Audit the frozen oral-B-v2 cold-start recovery development validation."""
+import argparse
+import glob
+import hashlib
+import json
+import os
+
+from experiments.bci_v2_run import CONDITIONS
+from experiments.bci_v2_recovery_run import (
+ CHALLENGE_EPISODES,
+ FIXED_CONFIG,
+ MODEL_SEEDS,
+ TARGETS,
+ TASK_SEEDS,
+ build_recovery_config,
+ source_paths,
+)
+
+
+def require(condition, message):
+ if not condition:
+ raise ValueError(message)
+
+
+def sha256(path):
+ digest = hashlib.sha256()
+ with open(path, "rb") as handle:
+ for chunk in iter(lambda: handle.read(1024 * 1024), b""):
+ digest.update(chunk)
+ return digest.hexdigest()
+
+
+def checks(row):
+ conditions = row["conditions"]
+ signatures = row["signatures"]
+ intact = conditions["intact"]
+ return {
+ "final_success_at_least_0p70":
+ intact["final_success"] >= 0.70,
+ "learning_gain_at_least_0p10":
+ intact["learning_gain"] >= 0.10,
+ "fixed_role_gap_at_least_0p20": (
+ intact["final_success"]
+ - conditions["fixed_role"]["final_success"]
+ ) >= 0.20,
+ "oracle_deficit_at_most_0p10": (
+ conditions["oracle_role"]["final_success"]
+ - intact["final_success"]
+ ) <= 0.10,
+ "plasticity_lesion_retains_at_most_half_gain":
+ conditions["plasticity_lesion"]["learning_gain"]
+ <= 0.5 * intact["learning_gain"],
+ "role_cosine_at_least_0p80":
+ intact["role_cosine_after_training"] >= 0.80,
+ "challenge_fraction_between_0p15_and_0p85":
+ 0.15 <= signatures["challenge_success_fraction"] <= 0.85,
+ "terminal_outcome_accuracy_at_least_0p75":
+ signatures[
+ "terminal_residual_outcome_balanced_acc"
+ ] >= 0.75,
+ "terminal_role_separation_at_least_0p20":
+ signatures[
+ "terminal_role_aligned_outcome_separation"
+ ] >= 0.20,
+ "outcome_lesion_separation_drop_at_least_0p20":
+ signatures[
+ "terminal_outcome_separation_drop_under_acute_lesion"
+ ] >= 0.20,
+ "critic_expectedness_at_least_0p05":
+ signatures[
+ "mean_critic_expectedness_contribution"
+ ] >= 0.05,
+ "critic_contribution_value_corr_at_least_0p95":
+ signatures[
+ "critic_contribution_value_prediction_corr"
+ ] >= 0.95,
+ "residual_soma_corr_at_most_0p10":
+ signatures["mean_abs_residual_soma_corr"] <= 0.10,
+ "raw_residual_corr_gap_at_least_0p20":
+ signatures[
+ "raw_minus_residual_abs_soma_corr"
+ ] >= 0.20,
+ "surrounding_event_accuracy_at_least_0p52":
+ signatures[
+ "surrounding_event_decoder_balanced_acc"
+ ] >= 0.52,
+ "decoder_distance_corr_at_least_0p02":
+ signatures[
+ "decoder_distance_residual_corr"
+ ] >= 0.02,
+ "sign_inversion_at_least_0p01":
+ signatures[
+ "causal_role_sign_inversion_index"
+ ] >= 0.01,
+ "velocity_advantage_at_least_0p05":
+ signatures[
+ "velocity_minus_error_abs_cv_corr"
+ ] >= 0.05,
+ }
+
+
+def validate(row, path, task_seed, digests):
+ require(row.get("schema_version") == 3, f"{path}: schema")
+ require(
+ row.get("args", {}).get("task_seed") == task_seed
+ and row["args"].get("model_seed") in MODEL_SEEDS,
+ f"{path}: seeds",
+ )
+ protocol = row.get("protocol", {})
+ require(
+ protocol.get("name")
+ == "oral_b_v2_cold_start_recovery_development_v1"
+ and protocol.get("selection_split") == "development_validation"
+ and protocol.get("hyperparameter_selection") is False
+ and protocol.get("confirmation_seeds_touched") is False
+ and protocol.get("fixed_config") == FIXED_CONFIG,
+ f"{path}: protocol",
+ )
+ require(
+ protocol.get("protocol_sha256") == digests["protocol"]
+ and protocol.get("d4_gate_sha256") == digests["d4_gate"]
+ and protocol.get("old_r2_gate_sha256")
+ == digests["old_r2_gate"]
+ and protocol.get("failed_v2_gate_sha256")
+ == digests["failed_v2_gate"],
+ f"{path}: parent digest",
+ )
+ provenance = row.get("provenance", {})
+ require(
+ provenance.get("git_tracked_dirty") is False
+ and provenance.get("tracked_inputs")
+ and all(provenance["tracked_inputs"].values())
+ and set(provenance["tracked_inputs"]) == set(digests)
+ and provenance.get("input_sha256") == digests,
+ f"{path}: provenance",
+ )
+ require(
+ row.get("config") == vars(build_recovery_config()),
+ f"{path}: config",
+ )
+ require(
+ row.get("split") == "development"
+ and row.get("finite") is True,
+ f"{path}: finite/split",
+ )
+ require(
+ set(row.get("conditions", {})) == set(CONDITIONS),
+ f"{path}: conditions",
+ )
+ for name in CONDITIONS:
+ warmup = row["warmup"][name]
+ require(
+ warmup["batches"] == 100
+ and warmup["examples"] == 6400
+ and warmup["instruction_present"] is False
+ and warmup["role_cursor_scalar_observations"] == 12800
+ and warmup["predictor_max_abs_error"] <= 1e-5,
+ f"{path}: warmup {name}",
+ )
+ cost = row["conditions"][name]["cost"]
+ require(
+ len(row["conditions"][name]["daily_success"]) == 14
+ and cost["maximum_state_episode_steps"] == 25088
+ and cost["cursor_scalar_observations"]
+ == 2 * cost["role_probe_examples"]
+ and cost["task_loss_queries"] == 0
+ and cost["reverse_mode_calls"] == 0,
+ f"{path}: condition {name}",
+ )
+ challenge = row["assays"]["challenge"]
+ require(
+ challenge["targets"] == list(TARGETS)
+ and challenge["episodes_per_target"] == CHALLENGE_EPISODES
+ and challenge["selection_over_targets"] is False
+ and challenge["maximum_steps_per_episode"] == 28,
+ f"{path}: target ladder",
+ )
+ require(
+ row["signatures"]["challenge_episodes"]
+ == len(TARGETS) * CHALLENGE_EPISODES,
+ f"{path}: challenge count",
+ )
+
+
+def main():
+ parser = argparse.ArgumentParser()
+ parser.add_argument(
+ "--results", default="results/bci_v2_recovery_dev"
+ )
+ parser.add_argument(
+ "--out", default="results/bci_v2_recovery_dev_gate.json"
+ )
+ args = parser.parse_args()
+ digests = {
+ name: sha256(path)
+ for name, path in source_paths().items()
+ }
+ expected = {
+ f"bci_v2_recovery_t{task_seed}_m0.json"
+ for task_seed in TASK_SEEDS
+ }
+ observed = {
+ os.path.basename(path)
+ for path in glob.glob(os.path.join(args.results, "*.json"))
+ }
+ require(
+ observed == expected,
+ (
+ f"recovery grid drift: missing={sorted(expected-observed)}, "
+ f"extra={sorted(observed-expected)}"
+ ),
+ )
+ records = {}
+ commits = set()
+ source_sha256 = {}
+ for task_seed in TASK_SEEDS:
+ path = os.path.join(
+ args.results,
+ f"bci_v2_recovery_t{task_seed}_m0.json",
+ )
+ with open(path) as handle:
+ row = json.load(handle)
+ validate(row, path, task_seed, digests)
+ records[task_seed] = row
+ commits.add(row["provenance"]["git_commit"])
+ source_sha256[path] = sha256(path)
+ require(len(commits) == 1, "recovery source revision drift")
+ checks_by_seed = {
+ str(seed): checks(records[seed]) for seed in TASK_SEEDS
+ }
+ passed = all(
+ value for seed_checks in checks_by_seed.values()
+ for value in seed_checks.values()
+ )
+ output = {
+ "protocol":
+ "oral_b_v2_cold_start_recovery_development_v1",
+ "status": "passed" if passed else "failed",
+ "complete_grid": True,
+ "fixed_config": FIXED_CONFIG,
+ "checks_by_task_seed": checks_by_seed,
+ "intact_final_by_task_seed": [
+ records[seed]["conditions"]["intact"]["final_success"]
+ for seed in TASK_SEEDS
+ ],
+ "challenge_success_fraction_by_task_seed": [
+ records[seed]["signatures"]["challenge_success_fraction"]
+ for seed in TASK_SEEDS
+ ],
+ "terminal_outcome_accuracy_by_task_seed": [
+ records[seed]["signatures"][
+ "terminal_residual_outcome_balanced_acc"
+ ]
+ for seed in TASK_SEEDS
+ ],
+ "confirmation_seeds_touched": False,
+ "recovery_confirmation_opened": passed,
+ "source_commit": next(iter(commits)),
+ "source_sha256": source_sha256,
+ "input_sha256": digests,
+ "review_score_before": 7,
+ "review_score_after": 7,
+ "score_change_rule": (
+ "development validation never changes the formal score"
+ ),
+ }
+ os.makedirs(os.path.dirname(os.path.abspath(args.out)), exist_ok=True)
+ if os.path.exists(args.out):
+ with open(args.out) as handle:
+ existing = json.load(handle)
+ require(
+ existing == output,
+ "existing recovery development gate differs",
+ )
+ else:
+ with open(args.out, "w") as handle:
+ json.dump(output, handle, indent=2, sort_keys=True)
+ handle.write("\n")
+ print(json.dumps(output, indent=2))
+
+
+if __name__ == "__main__":
+ main()
diff --git a/experiments/bci_v2_confirmation.sh b/experiments/bci_v2_confirmation.sh
index da7328c..a071526 100644
--- a/experiments/bci_v2_confirmation.sh
+++ b/experiments/bci_v2_confirmation.sh
@@ -13,4 +13,4 @@ for task_seed in 30 31 32 33 34 35; do
done
done
-python experiments/analyze_bci_v2_confirmation.py
+PYTHONPATH=. python experiments/analyze_bci_v2_confirmation.py
diff --git a/experiments/bci_v2_development_screen.sh b/experiments/bci_v2_development_screen.sh
index 8612cae..b2c5141 100644
--- a/experiments/bci_v2_development_screen.sh
+++ b/experiments/bci_v2_development_screen.sh
@@ -19,4 +19,4 @@ for forward_eta in 0.03 0.1; do
done
done
-python experiments/analyze_bci_v2_development.py
+PYTHONPATH=. python experiments/analyze_bci_v2_development.py
diff --git a/experiments/bci_v2_recovery_confirmation.py b/experiments/bci_v2_recovery_confirmation.py
new file mode 100644
index 0000000..9d45626
--- /dev/null
+++ b/experiments/bci_v2_recovery_confirmation.py
@@ -0,0 +1,141 @@
+#!/usr/bin/env python3
+"""Run one untouched oral-B-v2 cold-start recovery confirmation cell."""
+import argparse
+import json
+import os
+import sys
+
+sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
+from experiments.bci_v2_recovery_run import (
+ FIXED_CONFIG,
+ PROTOCOL_PATH,
+ finite_tree,
+ provenance,
+ require_parent_gates,
+ run_cell,
+ sha256,
+ source_paths,
+)
+
+
+ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
+R1_GATE_PATH = os.path.join(
+ ROOT, "results", "bci_v2_recovery_dev_gate.json"
+)
+TASK_SEEDS = tuple(range(30, 36))
+MODEL_SEEDS = tuple(range(5))
+
+
+def require_r1_gate(r1):
+ digests = {
+ name: sha256(path)
+ for name, path in source_paths().items()
+ }
+ if not (
+ r1.get("protocol")
+ == "oral_b_v2_cold_start_recovery_development_v1"
+ and r1.get("status") == "passed"
+ and r1.get("complete_grid") is True
+ and r1.get("fixed_config") == FIXED_CONFIG
+ and r1.get("confirmation_seeds_touched") is False
+ and r1.get("recovery_confirmation_opened") is True
+ and r1.get("review_score_after") == 7
+ and r1.get("input_sha256") == digests
+ ):
+ raise ValueError(
+ "recovery R2 requires the complete eligible R1 gate"
+ )
+
+
+def main():
+ parser = argparse.ArgumentParser()
+ parser.add_argument(
+ "--task-seed", type=int, choices=TASK_SEEDS, required=True
+ )
+ parser.add_argument(
+ "--model-seed", type=int, choices=MODEL_SEEDS, required=True
+ )
+ parser.add_argument(
+ "--r1-gate",
+ default="results/bci_v2_recovery_dev_gate.json",
+ )
+ parser.add_argument(
+ "--outdir", default="results/bci_v2_recovery_confirmation"
+ )
+ args = parser.parse_args()
+ require_parent_gates()
+ with open(args.r1_gate) as handle:
+ r1 = json.load(handle)
+ require_r1_gate(r1)
+ source = provenance({
+ "development_runner": os.path.join(
+ ROOT, "experiments", "bci_v2_recovery_run.py"
+ ),
+ "runner": os.path.abspath(__file__),
+ "r1_gate": os.path.abspath(args.r1_gate),
+ })
+ if (
+ source["git_tracked_dirty"]
+ or not all(source["tracked_inputs"].values())
+ ):
+ raise RuntimeError(
+ "recovery R2 requires clean, tracked, frozen inputs"
+ )
+ cell = run_cell(
+ args.task_seed,
+ args.model_seed,
+ split="untouched_confirmation",
+ performance_seed_offset=520_000,
+ challenge_seed_offset=530_000,
+ )
+ result = {
+ "schema_version": 3,
+ "protocol": {
+ "name":
+ "oral_b_v2_cold_start_recovery_confirmation_v1",
+ "split": "untouched_confirmation",
+ "training_task_seed": args.task_seed,
+ "model_seed": args.model_seed,
+ "fixed_config": FIXED_CONFIG,
+ "no_further_selection": True,
+ "confirmation_grid_size": (
+ len(TASK_SEEDS) * len(MODEL_SEEDS)
+ ),
+ "protocol_sha256": sha256(PROTOCOL_PATH),
+ "r1_gate_sha256": sha256(args.r1_gate),
+ },
+ "args": vars(args),
+ "provenance": source,
+ **cell,
+ }
+ result["finite"] = finite_tree(result)
+ if not result["finite"]:
+ raise RuntimeError("non-finite recovery confirmation record")
+ os.makedirs(args.outdir, exist_ok=True)
+ path = os.path.join(
+ args.outdir,
+ (
+ f"bci_v2_recovery_confirm_t{args.task_seed}"
+ f"_m{args.model_seed}.json"
+ ),
+ )
+ if os.path.exists(path):
+ raise FileExistsError(f"refusing to overwrite {path}")
+ with open(path, "w") as handle:
+ json.dump(result, handle, indent=2, sort_keys=True)
+ handle.write("\n")
+ print(json.dumps({
+ "path": path,
+ "intact_final": result["conditions"]["intact"]["final_success"],
+ "challenge_success_fraction": result["signatures"][
+ "challenge_success_fraction"
+ ],
+ "terminal_outcome_accuracy": result["signatures"][
+ "terminal_residual_outcome_balanced_acc"
+ ],
+ "finite": result["finite"],
+ }, indent=2))
+
+
+if __name__ == "__main__":
+ main()
diff --git a/experiments/bci_v2_recovery_confirmation.sh b/experiments/bci_v2_recovery_confirmation.sh
new file mode 100644
index 0000000..16534b1
--- /dev/null
+++ b/experiments/bci_v2_recovery_confirmation.sh
@@ -0,0 +1,16 @@
+#!/usr/bin/env bash
+# Frozen v2 recovery R2: 6 untouched task seeds x 5 model seeds.
+set -euo pipefail
+
+export OMP_NUM_THREADS=1
+export MKL_NUM_THREADS=1
+
+for task_seed in 30 31 32 33 34 35; do
+ for model_seed in 0 1 2 3 4; do
+ python experiments/bci_v2_recovery_confirmation.py \
+ --task-seed "$task_seed" \
+ --model-seed "$model_seed"
+ done
+done
+
+PYTHONPATH=. python experiments/analyze_bci_v2_recovery_confirmation.py
diff --git a/experiments/bci_v2_recovery_development.sh b/experiments/bci_v2_recovery_development.sh
new file mode 100644
index 0000000..68f22f5
--- /dev/null
+++ b/experiments/bci_v2_recovery_development.sh
@@ -0,0 +1,12 @@
+#!/usr/bin/env bash
+# Frozen v2 recovery R1: one fixed mechanism x 3 new development seeds.
+set -euo pipefail
+
+export OMP_NUM_THREADS=1
+export MKL_NUM_THREADS=1
+
+for task_seed in 23 24 25; do
+ python experiments/bci_v2_recovery_run.py --task-seed "$task_seed"
+done
+
+PYTHONPATH=. python experiments/analyze_bci_v2_recovery_development.py
diff --git a/experiments/bci_v2_recovery_run.py b/experiments/bci_v2_recovery_run.py
new file mode 100644
index 0000000..29507b3
--- /dev/null
+++ b/experiments/bci_v2_recovery_run.py
@@ -0,0 +1,377 @@
+#!/usr/bin/env python3
+"""Run one frozen oral-B-v2 cold-start recovery development cell."""
+import argparse
+from dataclasses import replace
+import hashlib
+import json
+import os
+import platform
+import resource
+import subprocess
+import sys
+import time
+
+import torch
+
+sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
+from experiments.bci_v2_run import (
+ ACUTE_MODES,
+ CONDITIONS,
+ build_config,
+ evaluate_performance,
+ finite_tree,
+ neutral_warmup,
+ train,
+)
+from sdil.bci import generate_trajectories
+from sdil.bci_v2 import BCIV2, run_day_v2
+from sdil.bci_v2_recovery_metrics import (
+ annotate_target_events,
+ recovery_signature_metrics,
+)
+
+
+ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
+PROTOCOL_PATH = os.path.join(ROOT, "ORAL_B_V2_RECOVERY.md")
+D4_GATE_PATH = os.path.join(
+ ROOT, "results", "kp_dynamic_projection_confirmation_gate.json"
+)
+OLD_R2_GATE_PATH = os.path.join(
+ ROOT, "results", "bci_td_confirmation_gate.json"
+)
+FAILED_V2_GATE_PATH = os.path.join(
+ ROOT, "results", "bci_v2_dev_gate.json"
+)
+TASK_SEEDS = (23, 24, 25)
+MODEL_SEEDS = (0,)
+TARGETS = (1.55, 1.60, 1.65, 1.70, 1.75, 1.80)
+CHALLENGE_EPISODES = 128
+FIXED_CONFIG = {
+ "forward_eta": 0.1,
+ "gamma": 0.8,
+ "critic_eta": 0.03,
+ "velocity_reward_scale": 1.0,
+}
+
+
+def sha256(path):
+ digest = hashlib.sha256()
+ with open(path, "rb") as handle:
+ for chunk in iter(lambda: handle.read(1024 * 1024), b""):
+ digest.update(chunk)
+ return digest.hexdigest()
+
+
+def source_paths():
+ return {
+ "runner": os.path.abspath(__file__),
+ "development_analyzer": os.path.join(
+ ROOT,
+ "experiments",
+ "analyze_bci_v2_recovery_development.py",
+ ),
+ "confirmation_runner": os.path.join(
+ ROOT, "experiments", "bci_v2_recovery_confirmation.py"
+ ),
+ "confirmation_analyzer": os.path.join(
+ ROOT,
+ "experiments",
+ "analyze_bci_v2_recovery_confirmation.py",
+ ),
+ "base_dynamics": os.path.join(ROOT, "sdil", "bci.py"),
+ "v2_dynamics": os.path.join(ROOT, "sdil", "bci_v2.py"),
+ "v2_metrics": os.path.join(ROOT, "sdil", "bci_v2_metrics.py"),
+ "recovery_metrics": os.path.join(
+ ROOT, "sdil", "bci_v2_recovery_metrics.py"
+ ),
+ "protocol": PROTOCOL_PATH,
+ "d4_gate": D4_GATE_PATH,
+ "old_r2_gate": OLD_R2_GATE_PATH,
+ "failed_v2_gate": FAILED_V2_GATE_PATH,
+ }
+
+
+def provenance(extra_paths=None):
+ def run(command):
+ return subprocess.run(
+ command,
+ cwd=ROOT,
+ check=True,
+ capture_output=True,
+ text=True,
+ ).stdout.strip()
+
+ paths = source_paths()
+ if extra_paths:
+ paths.update(extra_paths)
+ relative = {
+ name: os.path.relpath(path, ROOT)
+ for name, path in paths.items()
+ }
+ tracked = {
+ name: subprocess.run(
+ ["git", "ls-files", "--error-unmatch", path],
+ cwd=ROOT,
+ capture_output=True,
+ ).returncode == 0
+ for name, path in relative.items()
+ }
+ return {
+ "git_commit": run(["git", "rev-parse", "HEAD"]),
+ "git_tracked_dirty": bool(run([
+ "git", "status", "--porcelain", "--untracked-files=no"
+ ])),
+ "tracked_inputs": tracked,
+ "input_sha256": {
+ name: sha256(path) for name, path in paths.items()
+ },
+ }
+
+
+def require_parent_gates():
+ with open(D4_GATE_PATH) as handle:
+ d4 = json.load(handle)
+ with open(OLD_R2_GATE_PATH) as handle:
+ old_r2 = json.load(handle)
+ with open(FAILED_V2_GATE_PATH) as handle:
+ failed_v2 = json.load(handle)
+ if not (
+ d4.get("status") == "passed"
+ and d4.get("review_score_after") == 7
+ and old_r2.get("status") == "failed"
+ and old_r2.get("review_score_after") == 7
+ and failed_v2.get("protocol") == "oral_b_v2_development_v1"
+ and failed_v2.get("status") == "failed"
+ and failed_v2.get("complete_grid") is True
+ and failed_v2.get("oral_b_v2_confirmation_opened") is False
+ and failed_v2.get("review_score_after") == 7
+ ):
+ raise ValueError(
+ "cold-start recovery requires D4 and both preserved failures"
+ )
+
+
+def build_recovery_config():
+ return replace(
+ build_config(
+ FIXED_CONFIG["forward_eta"],
+ FIXED_CONFIG["gamma"],
+ FIXED_CONFIG["critic_eta"],
+ ),
+ velocity_reward_scale=FIXED_CONFIG[
+ "velocity_reward_scale"
+ ],
+ )
+
+
+def evaluate_target_ladder(model, task_seed, seed_offset):
+ mode_events = {name: [] for name in ACUTE_MODES}
+ costs = {name: 0 for name in ACUTE_MODES}
+ seeds = {}
+ for target_index, target in enumerate(TARGETS, start=1):
+ challenge_cfg = replace(model.cfg, target=target)
+ trajectory_seed = (
+ seed_offset + 1_000 * task_seed + target_index - 1
+ )
+ seeds[str(target)] = trajectory_seed
+ trajectories = generate_trajectories(
+ challenge_cfg,
+ trajectory_seed,
+ days=1,
+ episodes=CHALLENGE_EPISODES,
+ )
+ for mode in ACUTE_MODES:
+ settings = {
+ "intact": {},
+ "acute_critic_lesion": {"critic_enabled": False},
+ "acute_outcome_lesion": {
+ "terminal_outcome_enabled": False
+ },
+ }[mode]
+ assay_model = model.clone()
+ assay_model.cfg = challenge_cfg
+ report = run_day_v2(
+ assay_model,
+ trajectories,
+ 0,
+ horizon=challenge_cfg.steps_per_episode,
+ plasticity_gain=0.0,
+ learn_role=False,
+ probe_role=False,
+ learn_predictor=False,
+ learn_critic=False,
+ collect=True,
+ **settings,
+ )
+ episode_offset = (
+ 2_000_000
+ + (target_index - 1) * CHALLENGE_EPISODES
+ )
+ annotate_target_events(
+ report["events"],
+ 0,
+ report["success"],
+ episode_offset,
+ target_index,
+ )
+ mode_events[mode].extend(report["events"])
+ costs[mode] += report["active_transitions"]
+ return mode_events, {
+ "targets": list(TARGETS),
+ "episodes_per_target": CHALLENGE_EPISODES,
+ "trajectory_seeds": seeds,
+ "active_state_episode_steps_by_mode": costs,
+ "selection_over_targets": False,
+ "maximum_steps_per_episode": model.cfg.steps_per_episode,
+ }
+
+
+def run_cell(
+ task_seed,
+ model_seed,
+ *,
+ split,
+ performance_seed_offset,
+ challenge_seed_offset,
+):
+ cfg = build_recovery_config()
+ trajectories = generate_trajectories(cfg, task_seed)
+ performance_seed = performance_seed_offset + task_seed
+ performance_trajectories = generate_trajectories(
+ cfg, performance_seed, days=1, episodes=256
+ )
+ initial = BCIV2(cfg, model_seed)
+ conditions = {}
+ trained = {}
+ warmups = {}
+ training_events = None
+ started = time.perf_counter()
+ for name in CONDITIONS:
+ model, warmup = neutral_warmup(
+ initial, task_seed, model_seed, name
+ )
+ events, report = train(
+ model, trajectories, name, collect=name == "intact"
+ )
+ warmups[name] = warmup
+ conditions[name] = report
+ trained[name] = model
+ if name == "intact":
+ training_events = events
+ for name in CONDITIONS:
+ report = evaluate_performance(
+ trained[name], performance_trajectories
+ )
+ conditions[name]["final_success"] = report["success_rate"]
+ conditions[name]["evaluation_active_state_episode_steps"] = (
+ report["active_transitions"]
+ )
+ challenge_events, challenge_cost = evaluate_target_ladder(
+ trained["intact"], task_seed, challenge_seed_offset
+ )
+ signatures = recovery_signature_metrics(
+ training_events,
+ challenge_events,
+ cfg,
+ trained["intact"].role,
+ TARGETS,
+ )
+ return {
+ "config": vars(cfg),
+ "warmup": warmups,
+ "conditions": conditions,
+ "signatures": signatures,
+ "assays": {
+ "performance_evaluation_seed": performance_seed,
+ "performance_evaluation_episodes": 256,
+ "challenge": challenge_cost,
+ },
+ "wall_s": time.perf_counter() - started,
+ "peak_rss_mib": (
+ resource.getrusage(resource.RUSAGE_SELF).ru_maxrss / 1024
+ ),
+ "hardware": {
+ "device": "cpu",
+ "platform": platform.platform(),
+ "torch_version": torch.__version__,
+ "threads": torch.get_num_threads(),
+ },
+ "split": split,
+ }
+
+
+def main():
+ parser = argparse.ArgumentParser()
+ parser.add_argument(
+ "--task-seed", type=int, choices=TASK_SEEDS, required=True
+ )
+ parser.add_argument(
+ "--model-seed", type=int, choices=MODEL_SEEDS, default=0
+ )
+ parser.add_argument(
+ "--outdir", default="results/bci_v2_recovery_dev"
+ )
+ args = parser.parse_args()
+ require_parent_gates()
+ source = provenance()
+ if (
+ source["git_tracked_dirty"]
+ or not all(source["tracked_inputs"].values())
+ ):
+ raise RuntimeError(
+ "v2 recovery R1 requires clean, tracked, frozen inputs"
+ )
+ cell = run_cell(
+ args.task_seed,
+ args.model_seed,
+ split="development",
+ performance_seed_offset=460_000,
+ challenge_seed_offset=470_000,
+ )
+ result = {
+ "schema_version": 3,
+ "protocol": {
+ "name": "oral_b_v2_cold_start_recovery_development_v1",
+ "selection_split": "development_validation",
+ "hyperparameter_selection": False,
+ "confirmation_seeds_touched": False,
+ "fixed_config": FIXED_CONFIG,
+ "protocol_sha256": source["input_sha256"]["protocol"],
+ "d4_gate_sha256": source["input_sha256"]["d4_gate"],
+ "old_r2_gate_sha256": source["input_sha256"]["old_r2_gate"],
+ "failed_v2_gate_sha256": source["input_sha256"][
+ "failed_v2_gate"
+ ],
+ },
+ "args": vars(args),
+ "provenance": source,
+ **cell,
+ }
+ result["finite"] = finite_tree(result)
+ if not result["finite"]:
+ raise RuntimeError("non-finite v2 recovery development record")
+ os.makedirs(args.outdir, exist_ok=True)
+ path = os.path.join(
+ args.outdir,
+ f"bci_v2_recovery_t{args.task_seed}_m0.json",
+ )
+ if os.path.exists(path):
+ raise FileExistsError(f"refusing to overwrite {path}")
+ with open(path, "w") as handle:
+ json.dump(result, handle, indent=2, sort_keys=True)
+ handle.write("\n")
+ print(json.dumps({
+ "path": path,
+ "intact_final": result["conditions"]["intact"]["final_success"],
+ "challenge_success_fraction": result["signatures"][
+ "challenge_success_fraction"
+ ],
+ "terminal_outcome_accuracy": result["signatures"][
+ "terminal_residual_outcome_balanced_acc"
+ ],
+ "finite": result["finite"],
+ }, indent=2))
+
+
+if __name__ == "__main__":
+ main()
diff --git a/experiments/bci_v2_recovery_smoke.py b/experiments/bci_v2_recovery_smoke.py
new file mode 100644
index 0000000..39055c1
--- /dev/null
+++ b/experiments/bci_v2_recovery_smoke.py
@@ -0,0 +1,104 @@
+#!/usr/bin/env python3
+"""Endpoint-free checks for the v2 cold-start recovery."""
+from dataclasses import replace
+import os
+import sys
+
+import torch
+
+sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
+from experiments.bci_v2_recovery_run import (
+ TARGETS,
+ build_recovery_config,
+)
+from sdil.bci_v2 import BCIV2
+
+
+def copy_state(source, target):
+ for name in ("W", "A", "coupling", "P", "P_bias", "critic"):
+ setattr(target, name, getattr(source, name).clone())
+
+
+def check_dense_velocity_cold_start():
+ recovered_cfg = replace(
+ build_recovery_config(),
+ days=2,
+ episodes_per_day=8,
+ steps_per_episode=2,
+ target=10.0,
+ inertia=0.0,
+ process_noise=0.0,
+ forward_eta=0.1,
+ perturb_every=99,
+ )
+ failed_cfg = replace(recovered_cfg, velocity_reward_scale=0.25)
+ recovered = BCIV2(recovered_cfg, model_seed=1)
+ recovered.P.copy_(recovered.coupling)
+ recovered.A.copy_(recovered.role)
+ recovered.critic.zero_()
+ failed = BCIV2(failed_cfg, model_seed=1)
+ copy_state(recovered, failed)
+ context = torch.zeros(8, recovered_cfg.context_dim)
+ context[:, 0] = 1.0
+ noise = torch.zeros(8, recovered_cfg.n_neurons)
+ noise[:, :recovered_cfg.n_plus] = torch.linspace(
+ 0.1, 0.8, 8
+ ).unsqueeze(1)
+ perturbations = torch.ones_like(noise)
+ recovered_initial_w = recovered.W.clone()
+ failed_initial_w = failed.W.clone()
+ recovered_state = recovered.initial_episode_state(8)
+ failed_state = failed.initial_episode_state(8)
+ _, recovered_record = recovered.step(
+ context,
+ noise,
+ perturbations,
+ recovered_state,
+ global_step=1,
+ final_step=False,
+ learn_role=False,
+ probe_role=False,
+ learn_predictor=False,
+ learn_critic=False,
+ critic_enabled=False,
+ )
+ _, failed_record = failed.step(
+ context,
+ noise,
+ perturbations,
+ failed_state,
+ global_step=1,
+ final_step=False,
+ learn_role=False,
+ probe_role=False,
+ learn_predictor=False,
+ learn_critic=False,
+ critic_enabled=False,
+ )
+ assert torch.allclose(
+ recovered_record["td_innovation"],
+ 4.0 * failed_record["td_innovation"],
+ )
+ assert torch.allclose(
+ recovered.W - recovered_initial_w,
+ 4.0 * (failed.W - failed_initial_w),
+ atol=1e-7,
+ )
+ print("recovery dense cold-start drive is exactly 4x failed v2")
+
+
+def check_fixed_target_ladder():
+ assert TARGETS == (1.55, 1.60, 1.65, 1.70, 1.75, 1.80)
+ maxima = torch.tensor([1.58, 1.63, 1.68, 1.73, 1.78])
+ fractions = torch.tensor([
+ (maxima >= target).float().mean() for target in TARGETS
+ ])
+ assert torch.all(fractions[:-1] >= fractions[1:])
+ assert bool((fractions > 0).any()) and bool((fractions < 1).any())
+ print("recovery target ladder is fixed, ordered, and monotone")
+
+
+if __name__ == "__main__":
+ check_dense_velocity_cold_start()
+ check_fixed_target_ladder()
+ print("ALL V2 COLD-START RECOVERY MECHANICS CHECKS PASSED")
diff --git a/sdil/bci_v2_recovery_metrics.py b/sdil/bci_v2_recovery_metrics.py
new file mode 100644
index 0000000..78d3449
--- /dev/null
+++ b/sdil/bci_v2_recovery_metrics.py
@@ -0,0 +1,44 @@
+"""Target-ladder wrapper for the oral-B-v2 cold-start recovery."""
+
+from sdil.bci_v2_metrics import (
+ annotate_events,
+ signature_metrics_v2,
+)
+
+
+def annotate_target_events(
+ events,
+ day,
+ final_success,
+ episode_offset,
+ target_index,
+):
+ """Reuse paired grouping while encoding the fixed target-ladder index."""
+ return annotate_events(
+ events,
+ day,
+ final_success,
+ episode_offset,
+ target_index,
+ )
+
+
+def recovery_signature_metrics(
+ train_events,
+ challenge_mode_events,
+ cfg,
+ role,
+ targets,
+):
+ values = signature_metrics_v2(
+ train_events,
+ challenge_mode_events,
+ cfg,
+ role,
+ )
+ by_index = values.pop("horizon_success_fraction")
+ values["target_success_fraction"] = {
+ str(target): by_index[str(index)]
+ for index, target in enumerate(targets, start=1)
+ }
+ return values