diff options
Diffstat (limited to 'ORAL_A_V2.md')
| -rw-r--r-- | ORAL_A_V2.md | 22 |
1 files changed, 22 insertions, 0 deletions
diff --git a/ORAL_A_V2.md b/ORAL_A_V2.md index d4586f6..f77c5d2 100644 --- a/ORAL_A_V2.md +++ b/ORAL_A_V2.md @@ -103,3 +103,25 @@ mechanics and causal capture rather than standard-scale task success. A V2-2 pass can move the score from 5 only after its full audit is committed. A V2-3 multi-depth, multi-seed pass is required for an oral-level scaling claim. +## Audited outcome (2026-07-22) + +All six V2-1 records were finite and shared clean source commit `fc8fe99`. +The frozen selector chose `eta_A=0.01` for both estimators. Structured +calibration improved exact teaching/negative-gradient alignment: + +| estimator | early-third alignment | all-layer alignment | +|:--|--:|--:| +| unit targets | 0.001105 | 0.011664 | +| channel subspace | 0.007209 | 0.052740 | + +This is a real 6.5x early-layer and 4.5x all-layer improvement at identical +causal-query count, but it fails two frozen advancement checks: early-third +alignment is below `0.01`, and its absolute gain over unit targets is `0.006104` +rather than `0.01`. The all-layer check passes because the late blocks reach +substantially higher alignment; the earliest blocks remain the bottleneck. +The recorded target powers (`306.807` for the full hidden field and `0.270893` +for channel-basis moments) are intentionally not divided or compared: the two +estimators report different metric spaces. + +V2-1 status is **failed**. V2-2 was not launched, no test endpoint or +confirmation seed was touched, and the strict reviewer score remains 5/10. |
