From 5dcd29000579d9dea626b47f84aa3fdc16700605 Mon Sep 17 00:00:00 2001 From: YurenHao0426 Date: Wed, 22 Jul 2026 12:41:18 -0500 Subject: results: close Oral-A v2 causal-capture gate --- ORAL_A_V2.md | 22 ++++++++++++++++++++++ 1 file changed, 22 insertions(+) (limited to 'ORAL_A_V2.md') diff --git a/ORAL_A_V2.md b/ORAL_A_V2.md index d4586f6..f77c5d2 100644 --- a/ORAL_A_V2.md +++ b/ORAL_A_V2.md @@ -103,3 +103,25 @@ mechanics and causal capture rather than standard-scale task success. A V2-2 pass can move the score from 5 only after its full audit is committed. A V2-3 multi-depth, multi-seed pass is required for an oral-level scaling claim. +## Audited outcome (2026-07-22) + +All six V2-1 records were finite and shared clean source commit `fc8fe99`. +The frozen selector chose `eta_A=0.01` for both estimators. Structured +calibration improved exact teaching/negative-gradient alignment: + +| estimator | early-third alignment | all-layer alignment | +|:--|--:|--:| +| unit targets | 0.001105 | 0.011664 | +| channel subspace | 0.007209 | 0.052740 | + +This is a real 6.5x early-layer and 4.5x all-layer improvement at identical +causal-query count, but it fails two frozen advancement checks: early-third +alignment is below `0.01`, and its absolute gain over unit targets is `0.006104` +rather than `0.01`. The all-layer check passes because the late blocks reach +substantially higher alignment; the earliest blocks remain the bottleneck. +The recorded target powers (`306.807` for the full hidden field and `0.270893` +for channel-basis moments) are intentionally not divided or compared: the two +estimators report different metric spaces. + +V2-1 status is **failed**. V2-2 was not launched, no test endpoint or +confirmation seed was touched, and the strict reviewer score remains 5/10. -- cgit v1.2.3