summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
authorYurenHao0426 <Blackhao0426@gmail.com>2026-07-23 08:24:43 -0500
committerYurenHao0426 <Blackhao0426@gmail.com>2026-07-23 08:24:43 -0500
commit6ebf1f1858590b53b47f67708cbd79aa57f79a67 (patch)
tree07ffa9397133e2f696ddca92c0fe1ac276820ab5
parent75a64888cb3e117ee6d232adc7b48974647fbc1e (diff)
paper: integrate calibrated oral-B-v2 evidence
-rw-r--r--README.md33
-rw-r--r--REVIEW_SCORECARD.md40
-rw-r--r--ROADMAP.md25
-rw-r--r--experiments/audit_manuscript.py2
-rwxr-xr-xexperiments/finalize_accept.sh50
-rw-r--r--paper/CLAIM_LEDGER.json127
-rw-r--r--paper/MANUSCRIPT.md150
-rw-r--r--paper/manuscript_audit.json125
8 files changed, 502 insertions, 50 deletions
diff --git a/README.md b/README.md
index b6782cc..b5ec4d4 100644
--- a/README.md
+++ b/README.md
@@ -82,7 +82,7 @@ and scaling behavior. See `NOVELTY.md` for the exact prior-art boundary.
`0.999687`, and all trajectory, projection, leakage, query, hardware, MAC,
memory, split, and test-isolation checks pass. This untouched confirmation
establishes the strict 7/10 accept bar; it does not establish that added
- standard-network depth is useful or repair the failed biological signatures.
+ standard-network depth is useful.
- Native author-code fidelity is complete. BurstCCN reaches `80.10%` at its
validation-selected epoch versus published `82.97 +/- 0.21%`; Dual Prop
reaches `92.46%` versus published `92.41 +/- 0.07%`. Their audited walls are
@@ -108,6 +108,20 @@ works in this synthetic task, but the broader Harnett-like population
vectorization signature is not established. The strict score remains 7/10
and the oral-A depth panel stays sealed.
+An independent oral-B-v2 then adds an explicit terminal reward/timeout phase,
+a local linear TD critic, and temporal eligibility traces. Its first frozen
+grid is retained as a cold-start failure, and a fixed-target recovery is
+retained after one new seed remains at ceiling. The final algorithm is frozen
+before a label-free psychometric calibration: cursor-max quantiles from 512
+outcome-free trials define targets for independent rewarded/timeout trials.
+The complete untouched 6-task by 5-model confirmation passes every clustered
+gate: 100% task success, 0.976 mean learned-role cosine, 0.063
+residual--soma correlation, 54.37% surrounding-state decoding, 50.09%
+rewarded trials, 99.83% terminal outcome decoding, a 0.400 acute outcome-lesion
+drop in role separation, and 30/30 positive causal signs. This raises the
+formal milestone to 8/10 and permits only a separately frozen oral-A-v2
+protocol; the old depth panel remains closed.
+
## Publication-facing artifacts
- `RESULTS.md`: audited positive and negative results;
@@ -120,6 +134,9 @@ and the oral-A depth panel stays sealed.
frozen D1--D4 gates;
- `ORAL_B_RECOVERY.md`: structural diagnosis and mechanics-only temporal-
difference recovery boundary;
+- `ORAL_B_V2.md`, `ORAL_B_V2_RECOVERY.md`, and
+ `ORAL_B_V2_CALIBRATED_RECOVERY.md`: retained v2 failures and the passed
+ label-free calibrated outcome-surprise protocol;
- `ORAL_A.md`: frozen standard CIFAR ResNet funnel;
- `ORAL_A_V2.md`: frozen post-failure representable-subspace funnel;
- `ORAL_A_V3.md`: frozen vectorizer-space causal-calibration funnel;
@@ -129,8 +146,8 @@ and the oral-A depth panel stays sealed.
from manuscript numbers, gate statuses, figures, and claim boundaries to
their audited source files;
- `results/figs/`: deterministic PDF/PNG main figures, captions, and a source
- hash manifest, including the untouched D4 ResNet-20 confirmation, plus
- audited RRM failure and dynamic-stability supplements.
+ hash manifest, including the untouched D4 ResNet-20 and oral-B-v2
+ confirmations, plus audited RRM failure and dynamic-stability supplements.
The current main figures show the local-method Pareto frontier, credit
assignment versus depth, the load-bearing innovation ablation, and the
@@ -234,7 +251,9 @@ accounting, and final finiteness. Development, validation, and untouched
confirmation results are never pooled. Failed gates close their branch instead
of triggering seed deletion or post-hoc threshold changes.
-The current strict reviewer estimate is `5/10` (borderline reject, confidence
-`4/5`): the mechanism and controlled depth-preservation results are strong,
-but the frozen standard-useful-scale attempt failed. The score changes only
-after an audited frozen stage, not after a pilot or a presentation improvement.
+The current formal milestone is `8/10` (accept, confidence `4/5`) after the
+untouched D4 and oral-B-v2 confirmations. A conservative external-review
+forecast is `7/10`: positive added-depth utility, original-data biological
+validation, and a credit pathway novel beyond inherited perturbation/KP
+mechanisms remain open. Scores change only after an audited frozen stage, not
+after a pilot or presentation improvement.
diff --git a/REVIEW_SCORECARD.md b/REVIEW_SCORECARD.md
index 1b23968..36da020 100644
--- a/REVIEW_SCORECARD.md
+++ b/REVIEW_SCORECARD.md
@@ -22,14 +22,20 @@ Every formal result report records:
| Dimension | Score | Strict reviewer assessment |
|:--|--:|:--|
-| Soundness | 3/4 | Theory, local-gradient checks, causal diagnostics, cost accounting, and frozen stop rules are unusually careful. The learned apical vectorizer remains an unresolved failure mode. |
+| Soundness | 4/4 | Theory, exact local-update checks, causal lesions, clustered uncertainty, cost accounting, and retained frozen failures make the implemented claims unusually well identified. |
| Novelty | 2/4 | Learned node-perturbation feedback is prior art. The defensible novelty is the per-cell innovation operation under mixed apical traffic, together with its causal and scaling analysis. |
-| Significance | 3/4 | Near-flat performance over 12x depth while DFA alignment collapses is potentially important, and dynamic innovation now reaches near-BP accuracy on a standard ResNet-20. Added-depth utility and broader biological generality remain absent. |
-| Empirical support | 4/4 | Five-depth scaling, residual necessity, the frozen 91.18% ResNet-20 validation endpoint, and an untouched five-seed paired test confirmation at 91.584% are strong. Useful-depth C2, broad endogenous C1, oral-B, A3, and the original MT-1 remain disclosed failures. |
+| Significance | 3/4 | Dynamic innovation reaches near-BP ResNet-20 accuracy and a separate local actor--critic reproduces role-vectorized outcome surprise. Added-depth utility and original-data biological validation remain absent. |
+| Empirical support | 4/4 | Five-depth preservation, residual necessity, untouched ResNet-20 confirmation, and a complete 30-record task-clustered BCI confirmation are strong. Useful-depth C2, broad endogenous C1, old oral-B, A3, and the original MT-1 remain disclosed failures. |
| Reproducibility | 4/4 | Code, exact provenance, seed panels, costs, failed branches, frozen selectors, and staged test-access rules are retained in git. |
-| **Overall** | **7/10** | **Weak accept: untouched five-seed test confirmation establishes that dynamic innovation is robust and noninferior to strong clean KP on ResNet-20. Positive added-depth utility and the biological signature remain oral-level gaps.** |
+| **Overall** | **8/10** | **Internal accept milestone: untouched confirmations establish both load-bearing ResNet-20 innovation and role-vectorized TD outcome surprise under the declared synthetic BCI paradigm.** |
| Confidence | 4/5 | High confidence in the assessment because the positive and negative branches are both extensively audited. |
+The conservative external-review forecast is **7/10**, not 8: a reviewer can
+reasonably discount the synthetic BCI because terminal reward is supplied and
+the psychometric target range is calibrated per policy. The 8/10 value is the
+repository's predeclared evidence milestone; the external forecast is the
+recommendation I would actually submit as a reviewer today.
+
### Evidence already carrying the paper
- On flattened CIFAR-10, SDIL changes by only `-0.214 +/- 0.349` accuracy points from depth 5 to
@@ -39,6 +45,10 @@ Every formal result report records:
- On untouched CIFAR-10 test endpoints, dynamic innovation reaches `91.584%` versus clean KP's
`91.388%` over five paired ResNet-20 seeds; its paired deficit upper bound is only `0.131`
points and every frozen mechanism/cost invariant passes.
+- In the untouched six-task by five-model BCI panel, final task success is
+ `100%`, terminal residual outcome decoding is `99.83%`, the acute
+ outcome-lesion separation drop is `0.400`, critic expectedness is `0.319`,
+ and all 30 causal signs are positive after label-free calibration.
- The local update has a proved descent condition and an explicit query/MAC/memory audit; direct
node perturbation isolates the learned vectorizer as the useful-depth bottleneck.
@@ -49,9 +59,10 @@ Every formal result report records:
but cost `68.4x` ordinary forward-equivalent work.
3. Innovation was not uniformly beneficial for arbitrary endogenous top-down traffic, so the
supported mechanism is narrower than the initial claim.
-4. The temporal-difference recovery solves the task, passes the plasticity lesion, and yields
- 30/30 positive sign inversions, but its untouched R2 panel fails residual outcome advantage and
- longitudinal prediction. The broader Harnett-like population signature remains unsupported.
+4. The passed BCI is synthetic: reward is directly supplied, causal roles are
+ experimenter-defined for diagnostics, and target quantiles are calibrated
+ on a separate cursor split. Longitudinal prediction remains failed, and no
+ original Francioni/Harnett event-level data are tested.
5. The standard-network result inherits reciprocal KP and pays for a paired neutral microphase.
D4 establishes the innovation operation, not a new credit-transport mechanism or a literal
cortical implementation.
@@ -60,7 +71,7 @@ Every formal result report records:
| Checkpoint | Overall | What changed | Remaining ceiling |
|:--|--:|:--|:--|
-| Current audited package | 7 | The untouched D4 panel reaches 91.584% dynamic versus 91.388% clean KP over five paired test seeds, with every mechanism and cost invariant passing | Added-depth utility and oral-B biology remain absent |
+| Current audited package | 8 | D4 confirms load-bearing ResNet-20 innovation; calibrated oral-B-v2 confirms role-vectorized TD outcome surprise over 30 untouched records with all cluster bounds passing | Added-depth utility and original-data biological validation remain absent |
| Native baselines complete | 5 | BurstCCN is below its published endpoint; Dual Prop reproduces 92.46% versus 92.41%, with strict provenance and cost semantics | Fairness objection narrows, but SDIL gains no standard-scale evidence |
| Oral-A A1/A2 | 5 | BP reached 91.62%; short channel-gated SDIL reached 41.98% versus tuned DFA at 37.16% | Development screening alone cannot raise the score |
| Oral-A A3 fails | 5 | Full ResNet-20 SDIL became nonfinite at epoch 89 and ended at 10%; DFA ended finite at 33.06% | Standard-scale and oral-A claims are closed; A4 remains untouched |
@@ -83,14 +94,17 @@ Every formal result report records:
| Dynamic projection D3 | 6 | All 19 frozen checks pass at 91.18%, within 0.44/0.08 points of BP/clean KP, with 0.9994 early alignment and 1.326x BP MACs | One validation seed cannot establish robustness |
| Dynamic projection D4 | 7 | All ten untouched records pass: dynamic 91.584% versus clean KP 91.388%, paired upper deficit bound 0.131 points, early alignment 0.999687, and no invariant failures | Establishes ResNet-20 robustness/noninferiority, not positive depth utility |
| Oral-B recovery R1/R2 | failed at R2 | R1 selects eta 0.1 with 98.05% worst-task success; untouched R2 retains 99.53% mean success and 30/30 positive signs but fails outcome-vectorization and longitudinal gates | Score remains 7; the joint oral-B claim is not established |
+| Oral-B-v2 initial grid | failed at development | All 24 records preserve role learning and residual identification but fail from a one-quarter dense-signal cold start | Failure retained; no confirmation touched |
+| Oral-B-v2 fixed-target recovery | failed at development | Two seeds pass 18/18; the third passes 17/18 but its fixed target ladder has 98.96% success | Mechanism works, absolute assay scale does not generalize |
+| Oral-B-v2 calibrated R1/R2 | 8 | Three fresh development seeds pass, then all 30 untouched records and every task-cluster bound pass under independent label-free calibration/evaluation splits | Establishes synthetic outcome surprise; does not establish cortex or added-depth utility |
| Oral-A dynamic depth recovery | closed | A 60-cell ResNet-20/32/56 BP/DFA/clean-KP/dynamic panel was frozen before any new endpoint | Its oral-B R2 prerequisite failed, so none of the 50 new cells may run |
| Oral-A A4 | not opened | The prerequisite A3 gate failed | No oral-A confirmation claim is available |
These are conditional reviewer forecasts, not promised scores. A failed stage leaves its negative
result in the record and can lower the score if it invalidates a current claim. The original
-oral-B branch remains failed and cannot be retroactively reopened by vision results. The
-plasticity-only recovery passes R1 but fails its separately frozen R2 joint gate, so oral-A
-remains closed.
+oral-B branch and both v2 development failures remain failed. The calibrated
+v2 pass does not reopen the old oral-A panel; it permits only a new
+independently frozen oral-A-v2 protocol.
## Evidence-to-score log
@@ -127,6 +141,10 @@ remains closed.
| 2026-07-23 / `03c94a1` oral-B recovery R2 | Thirty untouched records retain 99.53% mean success, 90.45-point gain, 30/30 positive signs, and strong decorrelation, but fail seven population-vectorization/longitudinal checks | 7 → 7 | The recovery fixes learning, causal role, and sign but not the broader Harnett-like signature; oral-B and oral-A close without threshold repair |
| 2026-07-23 / `2a6f72e` audited D4 main figure | The strict renderer independently rechecks the ten D4 records and visualizes paired test accuracy, layerwise raw-versus-innovation direction, all 200 tracking epochs, and explicit neutral/MAC/memory/wall costs | 7 → 7 | Makes the accept evidence reviewable without adding or selecting data; presentation improves, but visualization alone cannot repair oral-B or justify score inflation |
| 2026-07-23 / `87cfb93` evidence-bound manuscript | A 3,238-word working draft binds 34 central numbers and all four figures to source manifests, retains the passed D4 gate and all seven failed R2 checks, and is re-audited by the accept finalizer | 7 → 7 | Substantially improves submission readiness and guards against claim drift; it adds no empirical evidence, so soundness and recommendation do not inflate |
+| 2026-07-23 / `eb021a6` oral-B-v2 initial R1 | The complete 24-record grid learns causal roles and identifies innovations but reaches at most 0.78% evaluation success because terminal reward remains unreachable | 7 → 7 | Localizes a cold-start created by scaling the only pre-reward drive to one quarter; confirmation remains untouched |
+| 2026-07-23 / `378e68d` fixed-target recovery R1 | Unit dense velocity restores 100% task learning and 17--18 biological checks per seed, but one fresh seed has 98.96% challenge success | 7 → 7 | Confirms the algorithmic repair while falsifying an absolute target ladder as a model-independent assay |
+| 2026-07-23 / `70e180c`, `9a8c057` calibrated oral-B-v2 R1/R2 | Three new development seeds pass all 18 gates; all 30 untouched confirmation records then pass every clustered learning, innovation, decoder, lesion, and expectedness bound | 7 → 8 | Establishes role-vectorized TD outcome surprise in the synthetic paradigm and raises the formal milestone; ecological validity and added depth remain the external-review ceiling |
+| 2026-07-23 / `75a6488` audited oral-B-v2 figure | The strict renderer visualizes all 30 records, task-cluster learning, residualization, independent psychometrics, and acute lesions | 8 → 8 | Improves reviewability without adding evidence or inflating the score |
Future rows are appended only after an audited frozen stage. A score staying flat is informative:
engineering, theory exposition, or visualization may make the paper more defensible without
diff --git a/ROADMAP.md b/ROADMAP.md
index c40efc5..39b78f5 100644
--- a/ROADMAP.md
+++ b/ROADMAP.md
@@ -55,6 +55,31 @@ outcome accuracy is only `47.33%`, residuals trail soma outcome decoding by
oral-A stays sealed, and no threshold repair or replacement confirmation is
permitted. Because `kappa=0`, neither R1 nor R2 could support online control.
+**Oral-B-v2 development path: two failures retained.** `ORAL_B_V2.md` adds an
+explicit terminal reward/timeout phase, a local TD critic, and eligibility
+traces without changing the old R2. Its complete 24-record development grid
+fails from a cold start: the dense performance innovation was scaled to
+one-quarter of the validated rule and no candidate learns. A separately frozen
+unit-scale recovery restores 100% task performance and passes every mechanism
+gate on task seeds 23 and 24; task seed 25 passes 17/18 gates but succeeds on
+98.96% of a fixed absolute target ladder. That class-balance failure is also
+retained and does not open confirmation.
+
+**Oral-B-v2 calibrated R1/R2 status: passed.** The final branch changes no
+learning parameter. Five target levels are fixed by cursor-maximum quantiles
+on 512 separate outcome-free calibration trials, then evaluated on independent
+trajectories. Fresh development seeds 26--28 pass all 18/18 checks. The
+untouched 6-task by 5-model confirmation then passes every learning,
+innovation, network-prediction, outcome, lesion, and task-cluster confidence
+gate: final success is `100%`, learning gain `98.80` points, fixed-role gap
+`99.97` points, role cosine `0.9761`, residual--soma correlation `0.0631`,
+surrounding accuracy `54.37%` (lower bound `54.26%`), velocity advantage
+`0.6384`, terminal outcome accuracy `99.83%` (lower bound `99.68%`), acute
+outcome-lesion separation drop `0.4000` (lower bound `0.3886`), critic
+expectedness `0.3187` (lower bound `0.2814`), and 30/30 positive signs. The
+formal milestone rises from 7 to 8. The old oral-A gate remains closed; only a
+new independently frozen oral-A-v2 protocol may now run.
+
## Frozen accept claims and gates
### C1. Innovation is necessary under naturally mixed apical traffic
diff --git a/experiments/audit_manuscript.py b/experiments/audit_manuscript.py
index 5743f24..555a50e 100644
--- a/experiments/audit_manuscript.py
+++ b/experiments/audit_manuscript.py
@@ -15,7 +15,7 @@ EXPECTED_SECTIONS = [
"## 3. What residualization guarantees—and what it does not",
"## 4. Experimental protocol",
"## 5. Results",
- "## 6. Biological-signature test and negative evidence",
+ "## 6. Biological-signature test and outcome-surprise evidence",
"## 7. Related work",
"## 8. Limitations and discussion",
"## 9. Reproducibility statement",
diff --git a/experiments/finalize_accept.sh b/experiments/finalize_accept.sh
index b39a714..db3f9bd 100755
--- a/experiments/finalize_accept.sh
+++ b/experiments/finalize_accept.sh
@@ -16,14 +16,28 @@ experiments/finalize_claims.sh
/home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \
experiments/bci_td_protocol_smoke.py
/home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \
+ experiments/bci_v2_smoke.py
+/home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \
+ experiments/bci_v2_recovery_smoke.py
+/home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \
+ experiments/bci_v2_calibrated_smoke.py
+/home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \
experiments/oral_a_dynamic_scaling_smoke.py
/home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 -m py_compile \
experiments/bci_td_run.py experiments/analyze_bci_td_development.py \
experiments/bci_td_confirmation.py \
experiments/analyze_bci_td_confirmation.py \
+ experiments/bci_v2_run.py experiments/analyze_bci_v2_development.py \
+ experiments/bci_v2_recovery_run.py \
+ experiments/analyze_bci_v2_recovery_development.py \
+ experiments/bci_v2_calibrated_run.py \
+ experiments/analyze_bci_v2_calibrated_development.py \
+ experiments/bci_v2_calibrated_confirmation.py \
+ experiments/analyze_bci_v2_calibrated_confirmation.py \
experiments/oral_a_dynamic_scaling.py \
experiments/analyze_oral_a_dynamic_scaling.py \
experiments/plot_resnet_confirmation.py \
+ experiments/plot_bci_v2_confirmation.py \
experiments/audit_manuscript.py
/home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \
experiments/analyze_kp_dynamic_projection.py >/dev/null
@@ -58,6 +72,38 @@ if [ -f results/bci_td_confirmation_gate.json ]; then
/home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \
experiments/analyze_bci_td_confirmation.py >/dev/null
fi
+if [ -f results/bci_v2_dev_gate.json ]; then
+ PYTHONPATH=. /home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \
+ experiments/analyze_bci_v2_development.py >/dev/null
+fi
+if [ -f results/bci_v2_recovery_dev_gate.json ]; then
+ PYTHONPATH=. /home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \
+ experiments/analyze_bci_v2_recovery_development.py >/dev/null
+fi
+if [ -f results/bci_v2_calibrated_dev_gate.json ]; then
+ PYTHONPATH=. /home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \
+ experiments/analyze_bci_v2_calibrated_development.py >/dev/null
+fi
+if [ -f results/bci_v2_calibrated_confirmation_gate.json ]; then
+ PYTHONPATH=. /home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \
+ experiments/analyze_bci_v2_calibrated_confirmation.py >/dev/null
+ /home/yurenh2/miniconda3/bin/python \
+ experiments/plot_bci_v2_confirmation.py >/dev/null
+ jq -e '
+ .strict == true and
+ .gate_status == "passed" and
+ .task_seeds == [30, 31, 32, 33, 34, 35] and
+ .model_seeds == [0, 1, 2, 3, 4] and
+ .record_count == 30 and
+ .statistics.intact_final_mean >= 0.99 and
+ .statistics.residual_soma_corr_mean <= 0.10 and
+ .statistics.surrounding_accuracy_lower >= 0.50 and
+ .statistics.terminal_accuracy_lower >= 0.75 and
+ .statistics.outcome_lesion_drop_lower >= 0.15 and
+ .statistics.critic_expectedness_lower >= 0.02 and
+ .statistics.positive_sign_count == 30
+ ' results/figs/figure5_bci_v2_manifest.json >/dev/null
+fi
if [ -f results/oral_a_dynamic_scaling_gate.json ]; then
/home/yurenh2/miniconda3/envs/ep_pascal/bin/python3 \
experiments/analyze_oral_a_dynamic_scaling.py >/dev/null
@@ -75,8 +121,8 @@ fi
jq -e '
.strict == true and
.status == "passed" and
- (.audited_numeric_claims | length) == 34 and
- (.gates | map(select(.status == "passed")) | length) == 1 and
+ (.audited_numeric_claims | length) == 51 and
+ (.gates | map(select(.status == "passed")) | length) == 2 and
(.gates | map(select(.status == "failed")) | length) == 1 and
(.gates[] | select(.status == "failed") | .false_checks | length) == 7
' paper/manuscript_audit.json >/dev/null
diff --git a/paper/CLAIM_LEDGER.json b/paper/CLAIM_LEDGER.json
index f687938..93fcc90 100644
--- a/paper/CLAIM_LEDGER.json
+++ b/paper/CLAIM_LEDGER.json
@@ -3,7 +3,8 @@
"../results/figs/figure1_pareto.png",
"../results/figs/figure2_scaling.png",
"../results/figs/figure3_innovation.png",
- "../results/figs/figure4_resnet_confirmation.png"
+ "../results/figs/figure4_resnet_confirmation.png",
+ "../results/figs/figure5_bci_v2.png"
],
"gate_statuses": [
{
@@ -13,6 +14,10 @@
{
"expected": "failed",
"source": "results/bci_td_confirmation_gate.json"
+ },
+ {
+ "expected": "passed",
+ "source": "results/bci_v2_calibrated_confirmation_gate.json"
}
],
"manuscript": "paper/MANUSCRIPT.md",
@@ -254,6 +259,125 @@
"pointer": "/metrics/longitudinal_prediction/mean",
"source": "results/bci_td_confirmation_gate.json",
"token": "-0.013"
+ },
+ {
+ "format": "percent3_percent",
+ "id": "bci_v2_final_performance",
+ "pointer": "/statistics/intact_final_mean",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "100.000%"
+ },
+ {
+ "format": "percent3",
+ "id": "bci_v2_learning_gain",
+ "pointer": "/statistics/learning_gain_mean",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "98.802"
+ },
+ {
+ "format": "percent3",
+ "id": "bci_v2_fixed_role_gap",
+ "pointer": "/statistics/fixed_role_gap_mean",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "99.974"
+ },
+ {
+ "format": "fixed4",
+ "id": "bci_v2_role_cosine",
+ "pointer": "/statistics/role_cosine_mean",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "0.9761"
+ },
+ {
+ "format": "fixed3",
+ "id": "bci_v2_residual_soma_corr",
+ "pointer": "/statistics/residual_soma_corr_mean",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "0.063"
+ },
+ {
+ "format": "fixed3",
+ "id": "bci_v2_raw_residual_gap",
+ "pointer": "/statistics/raw_residual_corr_gap_mean",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "0.936"
+ },
+ {
+ "format": "percent2_percent",
+ "id": "bci_v2_surrounding_accuracy",
+ "pointer": "/statistics/surrounding_accuracy_mean",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "54.37%"
+ },
+ {
+ "format": "percent2_percent",
+ "id": "bci_v2_surrounding_accuracy_lower",
+ "pointer": "/statistics/surrounding_accuracy_lower",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "54.26%"
+ },
+ {
+ "format": "fixed3",
+ "id": "bci_v2_velocity_advantage",
+ "pointer": "/statistics/velocity_advantage_mean",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "0.638"
+ },
+ {
+ "format": "percent3_percent",
+ "id": "bci_v2_challenge_fraction",
+ "pointer": "/statistics/challenge_fraction_mean",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "50.094%"
+ },
+ {
+ "format": "percent3_percent",
+ "id": "bci_v2_terminal_accuracy",
+ "pointer": "/statistics/terminal_accuracy_mean",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "99.833%"
+ },
+ {
+ "format": "percent3_percent",
+ "id": "bci_v2_terminal_accuracy_lower",
+ "pointer": "/statistics/terminal_accuracy_lower",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "99.680%"
+ },
+ {
+ "format": "percent2_percent",
+ "id": "bci_v2_soma_accuracy",
+ "pointer": "/statistics/terminal_previous_soma_accuracy_mean",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "73.50%"
+ },
+ {
+ "format": "fixed3",
+ "id": "bci_v2_outcome_lesion_drop",
+ "pointer": "/statistics/outcome_lesion_drop_mean",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "0.400"
+ },
+ {
+ "format": "fixed3",
+ "id": "bci_v2_outcome_lesion_drop_lower",
+ "pointer": "/statistics/outcome_lesion_drop_lower",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "0.389"
+ },
+ {
+ "format": "fixed3",
+ "id": "bci_v2_critic_expectedness",
+ "pointer": "/statistics/critic_expectedness_mean",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "0.319"
+ },
+ {
+ "format": "fixed3",
+ "id": "bci_v2_critic_expectedness_lower",
+ "pointer": "/statistics/critic_expectedness_lower",
+ "source": "results/figs/figure5_bci_v2_manifest.json",
+ "token": "0.281"
}
],
"required_boundary_text": [
@@ -262,6 +386,7 @@
"we do not claim arbitrary top-down traffic removal;",
"It does not support a ResNet-20-to-56 depth claim.",
"The joint biological gate nevertheless fails.",
+ "It does not establish that the same plasticity rule operates in cortex.",
"we do not call this variant single-phase."
],
"required_reference_urls": [
diff --git a/paper/MANUSCRIPT.md b/paper/MANUSCRIPT.md
index b4b735a..733284e 100644
--- a/paper/MANUSCRIPT.md
+++ b/paper/MANUSCRIPT.md
@@ -29,13 +29,16 @@ dynamic paired-neutral innovation rule reaches 91.584% mean CIFAR-10 test
accuracy across five untouched seeds, versus 91.388% for clean reciprocal
credit; the one-sided 95% upper bound on its deficit is 0.131 points. It uses
zero task-loss queries and 1.326 times the matched BP MAC estimate, but pays for
-one instruction-off neutral observation per training example. A preregistered
-synthetic BCI confirmation learns the task and causal sign yet fails outcome
-vectorization and longitudinal prediction. The supported conclusion is
-therefore algorithmic and narrow: neutral somato-dendritic innovation can
-protect local credit from soma-predictable traffic and remain stable on
-ResNet-20, without establishing a cortical learning rule or positive utility
-from added standard-network depth.
+one instruction-off neutral observation per training example. In a separate
+six-task by five-model synthetic BCI confirmation, a local actor--critic
+innovation reaches 100.000% task success, 99.833% terminal outcome decoding,
+and 30/30 predicted causal-role signs. An acute outcome lesion reduces
+role-aligned separation by 0.400. The supported conclusion remains
+algorithmic: somato-dendritic innovation can protect local credit from
+soma-predictable traffic, remain stable on ResNet-20, and multiplex performance
+change with outcome surprise in a controlled dynamical task. These results do
+not establish a cortical learning rule or positive utility from added
+standard-network depth.
## 1. Introduction
@@ -86,10 +89,9 @@ Our contributions are:
3. frozen evidence that the innovation operation is load-bearing under
soma-predictable traffic, including an independently confirmed ResNet-20
endpoint; and
-4. retained negative results showing where the proposal does not work:
- arbitrary top-down traffic, an unstable static ResNet predictor, a failed
- desired-velocity/online-control interpretation, and an untouched
- population-signature confirmation that fails its joint gate.
+4. an independently confirmed local actor--critic instantiation in which
+ learned causal roles vectorize performance change and outcome surprise,
+ together with retained failed protocols that delimit the result.
## 2. Somato-dendritic innovation learning
@@ -196,6 +198,37 @@ loss, downstream weight, or reverse pass. It is nevertheless a real
instruction-off microphase: every neutral observation and its elementwise
arithmetic are counted, and we do not call this variant single-phase.
+### 2.4 Temporal-difference outcome innovation
+
+The synthetic BCI branch uses the same innovation variable in a continuous
+dynamical task. A local linear critic reads the surrounding somatic population,
+
+\[
+V_t=v^\mathsf{T}[1,h_t],
+\qquad
+\delta_t=\rho_t+\gamma(1-z_t)V_{t+1}-V_t,
+\]
+
+where \(z_t\) marks a terminal event and
+\(\rho_t=|e_{t-1}|-|e_t|+\mathbb{1}\{\text{rewarded terminal}\}\).
+Sparse antithetic cursor probes estimate each cell's signed causal role
+\(m_i=\partial z/\partial h_i\); the instructional apical term is
+\(m_i\delta_t\). A local temporal eligibility trace,
+
+\[
+E_{ij,t}=0.8E_{ij,t-1}+
+(1-h_{i,t}^2)x_{j,t},
+\qquad
+\Delta W_{ij,t}=\eta r_{i,t}E_{ij,t},
+\]
+
+assigns the innovation to recent synaptic activity. Actor, critic, role
+estimator, and neutral predictor use manual local updates without autograd or
+task-loss queries. The terminal reward is explicitly supplied, so outcome
+encoding alone is not evidence for an emergent error code; the causal tests
+are residualization, learned role vectorization, critic expectedness, acute
+lesions, and untouched replication.
+
## 3. What residualization guarantees—and what it does not
Let \(n\) be neutral apical traffic and let
@@ -279,6 +312,25 @@ and dynamic innovation from scratch for untouched seeds 10--14, uses no
validation examples, and evaluates the 10,000-example test set once at the
endpoint. Both conditions use the same initialization/data seed pairing.
+### Synthetic BCI confirmation
+
+The biological-signature test uses 40 cells with known experimenter-only
+causal roles, 14 training days, 64 episodes per day, and 28 maximum steps per
+episode. Every condition receives the same instruction-off neutral warmup and
+paired scalar cursor probes. Fixed-role, plasticity-lesion, critic-lesion,
+outcome-lesion, and exact-role diagnostic conditions share all exogenous
+trajectories.
+
+Because independently learned policies have different cursor scales, a fixed
+absolute target produced a disclosed ceiling failure. The final assay freezes
+the algorithm, runs 512 separate outcome-free calibration trajectories, and
+uses fixed cursor-maximum quantiles
+\(\{0.20,0.35,0.50,0.65,0.80\}\) as target levels. Rewarded and timeout
+outcomes are then measured on independent trajectories; no evaluation label
+selects or reweights a target. Confirmation crosses six untouched task seeds
+and five model seeds. Uncertainty first averages models within task and then
+uses the six task seeds as independent clusters.
+
### Baselines and cost
In-repository controls include BP, FA, DFA, direct node perturbation,
@@ -371,7 +423,7 @@ time is 1.47 times paired clean KP. The result supports ResNet-20 robustness
under the audited traffic intervention. It does not support a ResNet-20-to-56
depth claim.
-## 6. Biological-signature test and negative evidence
+## 6. Biological-signature test and outcome-surprise evidence
The original online-control/desired-velocity screen fails its frozen causal
sign and acute-control gates. We therefore constructed a separate
@@ -395,14 +447,59 @@ the broader population outcome-vectorization and longitudinal signatures.
Because online control is disabled in this recovery, even a passed plasticity
gate would not establish an online desired-velocity controller.
+That failure exposed two structural mismatches rather than motivating a
+threshold repair. First, the old trajectory terminated without an explicit
+reward/timeout event, so the residual was never asked to encode the variable
+used by its outcome decoder. Second, near-ceiling task success left almost no
+unrewarded trials. A separately frozen v2 added an explicit actor--critic
+outcome phase. Its first development grid failed from a cold start when dense
+performance velocity was scaled to one quarter of the previously validated
+signal. Restoring unit scale recovered learning, but a fixed target ladder
+failed one new development seed because policy output scale varied. Both
+failures remain in the repository.
+
+The final recovery changes no learning parameter after that diagnosis. It uses
+the label-free calibration split described above and then opens a fully
+untouched 30-record confirmation.
+
+![Local actor--critic learning, residualization, psychometric calibration, and outcome lesions.](../results/figs/figure5_bci_v2.png)
+
+**Figure 5: Role-vectorized temporal-difference outcome surprise.** The
+renderer reads all 30 confirmation records and uses the task seed, not each
+model, as the uncertainty unit.
+
+All six task clusters reach 100.000% final success; mean learning gain is
+98.802 points and the fixed-role gap is 99.974 points. Learned-role cosine is
+0.9761. Nonterminal residual--soma correlation is 0.063, while subtracting the
+neutral prediction reduces absolute correlation by 0.936. The preceding
+surrounding population predicts causal-cell residual sign at 54.37% balanced
+accuracy, with a one-sided lower bound of 54.26%, and velocity has a 0.638
+absolute cross-validated correlation advantage over error magnitude. All 30
+records have the predicted positive P+/P- sign.
+
+Independent calibrated trials are balanced at 50.094% rewarded outcomes.
+Terminal residual outcome accuracy is 99.833% with a 99.680% lower bound,
+versus 73.50% from pre-outcome soma. Acute removal of terminal outcome input
+reduces role-aligned separation by 0.400 (lower bound 0.389). The learned
+critic contributes 0.319 of expectedness modulation (lower bound 0.281), and
+that contribution is paired with its stored value prediction.
+
+The result establishes the claimed signal within this synthetic paradigm:
+ordinary soma coupling is subtracted, recent performance change and terminal
+outcome are vectorized by learned cell-specific causal roles, the residual
+drives local eligibility-based plasticity, and the critic converts raw outcome
+into surprise. It does not establish that the same plasticity rule operates in
+cortex. Calibration makes rewarded and timeout trials statistically
+identifiable; it is an explicit psychometric phase and not a biological
+prediction by itself.
+
These negatives constrain interpretation:
- we do not infer that dendritic residuals directly drive cortical plasticity;
- we do not infer that cortex implements BP;
- we do not claim arbitrary top-down traffic removal;
- we do not claim positive utility from adding standard ResNet depth; and
-- we do not relabel task learning and sign inversion as a passed population
- signature.
+- we do not treat directly supplied terminal reward as an emergent error.
## 7. Related work
@@ -462,17 +559,19 @@ MAC overhead. Hardware implementations may price local elementwise operations,
state storage, and phases differently from GPUs; this is why we report several
resource axes rather than one scalar cost.
-The strongest standard result uses only ResNet-20. The separately frozen
-ResNet-20/32/56 panel remains unopened because its biological prerequisite
-failed. This preserves the declared accept-to-oral ordering but leaves positive
-added-depth utility unresolved.
+The strongest standard result uses only ResNet-20. The old separately frozen
+ResNet-20/32/56 panel remains unopened because its original biological
+prerequisite failed. The successful v2 gate permits only a new independently
+frozen depth protocol; it cannot retroactively open the old panel. Positive
+added-depth utility therefore remains unresolved.
-Finally, the synthetic BCI negatives are scientifically important. The
-residual operation can be algorithmically useful without reproducing every
-signature of cortical dendrites. A stronger biological paper needs a new
-prediction and mechanism frozen independently of the failed outcome and
-longitudinal metrics, ideally tested on the original event-level data rather
-than engineered to pass the present synthetic task.
+Finally, the synthetic BCI evidence has a hard ecological boundary. Outcome
+reward is supplied to the critic, the psychometric targets are calibrated to
+each trained policy on a separate split, and the task has experimenter-defined
+causal roles. The acute lesions show how the implemented signal is composed;
+they do not show that cortex uses the same decomposition. A stronger biological
+paper needs a prospective prediction tested on the original event-level data,
+and the still-failed longitudinal prediction should not be silently discarded.
## 9. Reproducibility statement
@@ -486,4 +585,5 @@ bash experiments/finalize_accept.sh
It rechecks the main figures, theoretical identities, local-rule mechanics,
baseline protocols, native-author records, standard-ResNet confirmation,
-failed biological confirmation, and the sealed standard-depth boundary.
+failed biological protocols, the passed calibrated BCI confirmation, and the
+sealed old standard-depth boundary.
diff --git a/paper/manuscript_audit.json b/paper/manuscript_audit.json
index 35cfcaa..55ee55a 100644
--- a/paper/manuscript_audit.json
+++ b/paper/manuscript_audit.json
@@ -203,6 +203,108 @@
"pointer": "/metrics/longitudinal_prediction/mean",
"rendered": "-0.013",
"source": "results/bci_td_confirmation_gate.json"
+ },
+ {
+ "id": "bci_v2_final_performance",
+ "pointer": "/statistics/intact_final_mean",
+ "rendered": "100.000%",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_learning_gain",
+ "pointer": "/statistics/learning_gain_mean",
+ "rendered": "98.802",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_fixed_role_gap",
+ "pointer": "/statistics/fixed_role_gap_mean",
+ "rendered": "99.974",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_role_cosine",
+ "pointer": "/statistics/role_cosine_mean",
+ "rendered": "0.9761",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_residual_soma_corr",
+ "pointer": "/statistics/residual_soma_corr_mean",
+ "rendered": "0.063",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_raw_residual_gap",
+ "pointer": "/statistics/raw_residual_corr_gap_mean",
+ "rendered": "0.936",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_surrounding_accuracy",
+ "pointer": "/statistics/surrounding_accuracy_mean",
+ "rendered": "54.37%",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_surrounding_accuracy_lower",
+ "pointer": "/statistics/surrounding_accuracy_lower",
+ "rendered": "54.26%",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_velocity_advantage",
+ "pointer": "/statistics/velocity_advantage_mean",
+ "rendered": "0.638",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_challenge_fraction",
+ "pointer": "/statistics/challenge_fraction_mean",
+ "rendered": "50.094%",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_terminal_accuracy",
+ "pointer": "/statistics/terminal_accuracy_mean",
+ "rendered": "99.833%",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_terminal_accuracy_lower",
+ "pointer": "/statistics/terminal_accuracy_lower",
+ "rendered": "99.680%",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_soma_accuracy",
+ "pointer": "/statistics/terminal_previous_soma_accuracy_mean",
+ "rendered": "73.50%",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_outcome_lesion_drop",
+ "pointer": "/statistics/outcome_lesion_drop_mean",
+ "rendered": "0.400",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_outcome_lesion_drop_lower",
+ "pointer": "/statistics/outcome_lesion_drop_lower",
+ "rendered": "0.389",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_critic_expectedness",
+ "pointer": "/statistics/critic_expectedness_mean",
+ "rendered": "0.319",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
+ },
+ {
+ "id": "bci_v2_critic_expectedness_lower",
+ "pointer": "/statistics/critic_expectedness_lower",
+ "rendered": "0.281",
+ "source": "results/figs/figure5_bci_v2_manifest.json"
}
],
"figures": [
@@ -221,6 +323,10 @@
{
"path": "results/figs/figure4_resnet_confirmation.png",
"sha256": "4011c9d31de4a53a7ae24bf4030693b22dabf6b7a7ed8a7eed2f6b5b91a97b78"
+ },
+ {
+ "path": "results/figs/figure5_bci_v2.png",
+ "sha256": "e43e07d56c28022e861682e7b7f68ef010e2754c64eb58cfc8079d8a59d57b75"
}
],
"gates": [
@@ -241,16 +347,21 @@
],
"source": "results/bci_td_confirmation_gate.json",
"status": "failed"
+ },
+ {
+ "false_checks": [],
+ "source": "results/bci_v2_calibrated_confirmation_gate.json",
+ "status": "passed"
}
],
"ledger": {
"path": "paper/CLAIM_LEDGER.json",
- "sha256": "5d779179ee8c1f6d2d00a34b2f4adaa483212996a54efe620a482115c1dcc689"
+ "sha256": "bd65f0747aba2dbe327f7150b7112dbc560c6e9cb984b058343679fd0dd89ab9"
},
"manuscript": {
"path": "paper/MANUSCRIPT.md",
- "sha256": "dc45c921be33b35c4a8629e308e118c1506719a0f018b807628be341c5a1da25",
- "word_count": 3238
+ "sha256": "c0633b8ff9f7b267e12eade47fac7b01429b6dc9b9255b71938e14169ca7822f",
+ "word_count": 4019
},
"sources": [
{
@@ -258,10 +369,18 @@
"sha256": "4f6f969ceae88afa2523e3472a3373522991f5ecaab69440830c07479d2d3597"
},
{
+ "path": "results/bci_v2_calibrated_confirmation_gate.json",
+ "sha256": "3b161c32727a2a7c4b50ca696dafff20a023d23e8cd008a0b8dddd662516b084"
+ },
+ {
"path": "results/figs/figure4_resnet_confirmation_manifest.json",
"sha256": "4aa3fbab16f41f608259b345283db7ef03ec8562f4e1f811ec6fe0a228dfa8c9"
},
{
+ "path": "results/figs/figure5_bci_v2_manifest.json",
+ "sha256": "dc24f1593b7e07b33b45e0e1dfec895cf7bdc5a953e60d933c5c3596f284738e"
+ },
+ {
"path": "results/figs/main_figure_manifest.json",
"sha256": "f5099b29d7f83ad7dc7783b927f682d7af43a90072e11fc7da2e8be5e7685450"
},