# Oral-A-v2: representable-subspace development protocol ## Status and claim boundary This is a transparent post-failure development branch. Frozen Oral-A-v1 failed when its SDIL run became nonfinite at epoch 89; it remains failed and is never relabeled. The executable diagnosis found that its unitwise K1 causal target was not captured by the translation-shared feedback model. The v2 change is restricted to that diagnosed interface: perturb the two basis fields that the channel-gated vectorizer can actually express, then apply the same expected full-field delta rule. Architecture, task update, causal-query count, cadence, sigma, data split, and v1 optimization schedule do not change. No v2 GPU endpoint was observed before this protocol and its selectors were committed. The CIFAR-10 test set and confirmation seeds 10--14 remain untouched. A v2 result cannot be pooled with v1 or described as preregistered before the v1 failure. ## V2-0: mechanics and mathematical identity `experiments/conv_local_smoke.py` must establish all of the following on CPU: - structured moment-estimator cosine above 0.985 and norm ratio in [0.90,1.10]; - antithetic directional derivative relative error below `2e-6` at a differentiable fixed ResNet point; - absolute error below `1e-14` between the implemented A/G update and the analytically expected full-field delta-rule update; - all legacy convolutional, BatchNorm, predictor, and local-eligibility checks remain green. This stage passed before the GPU funnel was opened: cosine `0.998217`, norm ratio `1.006989`, JVP relative error `1.69e-10`, and delta-rule absolute error `2.71e-20`. ## V2-1: frozen-forward causal-capture screen Use seed 0 ResNet-20, the frozen 45,000/5,000 split, the first 10,000 training examples, batch 128, and 400 feedback-only minibatches. Forward weights, BN state and affine parameters, and the readout remain bitwise fixed. The vectorizer is `channel_gated`, starts at exactly zero (`a_scale=0`), and uses K1 antithetic calibration with `sigma=0.01`. Cross: - estimator: legacy `unit_targets`, v2 `channel_subspace`; - apical rate: `0.01`, `0.1`, `1.0`. All candidates use identical data, perturbation seed 1000, and a fixed 64-example training-prefix exact-gradient audit. Autograd is used only after calibration to audit teaching alignment; it never updates a parameter. Within each estimator select maximum early-third teaching/negative-gradient cosine, then maximum all-layer cosine, then lower apical rate. V2 advances only if all records are finite and the selected structured estimator has: 1. early-third cosine at least `0.01`; 2. all-layer cosine at least `0.01`; 3. early-third cosine at least `0.01` above the best matched unit-target run. The estimator-specific prediction/target cosine is reported, but it is not compared across modes because one record is measured in the full hidden field and the other in its two channel-basis moments. No rate or threshold is added after results are observed. ## V2-2: full ResNet-20 validation gate Only if V2-1 passes, run one seed-0 SDIL model for 200 epochs on all 45,000 development-training examples. Copy Oral-A-v1 exactly: ResNet-20 width 16, batch 128, hidden LR `0.03`, output LR `0.1`, momentum `0.9`, weight decay `1e-4`, and 10x drops at epochs 100 and 150. Use the V2-1-selected apical rate, zero-initialized channel-gated feedback, 400 feedback-only warmup steps, and structured K1/every-4 calibration at `sigma=0.01`. The already frozen BP (`91.62%`) and DFA (`33.06%`) records remain the matched references. V2 passes only if: 1. every loss and metric is finite; 2. final validation accuracy is within 5 points of BP and at least 2 points above DFA; 3. early-third teaching alignment is at least `0.05`; 4. estimated total training MACs do not exceed BP. Failure closes v2; there is no LR, warmup, cadence, sigma, direction-count, or vectorizer recovery branch. ## V2-3: untouched depth confirmation Only after V2-2 passes, copy its complete hyperparameters without depth tuning to ResNet-20/32/56 and model/data-loader seeds 10--14. Refit on all 50,000 training examples. Each run evaluates CIFAR-10 test exactly once at the end; no intermediate test endpoint is logged. Exact BP and the v1-selected DFA are run on the identical panel. The confirmation gate is unchanged from Oral-A-v1: all 45 trajectories finite; SDIL within 2 points of BP at every depth; SDIL at least 2 points above DFA at depth 56; paired SDIL depth-20 to depth-56 change no worse than -2 points; depth-56 early-third alignment at least `0.05` and at least 30% of depth-20; and SDIL nondominated in accuracy versus estimated MACs, causal queries, and peak memory. Wall time is descriptive only. ## Reviewer-score rule V2-0 and V2-1 cannot raise the strict ICLR score because they establish mechanics and causal capture rather than standard-scale task success. A V2-2 pass can move the score from 5 only after its full audit is committed. A V2-3 multi-depth, multi-seed pass is required for an oral-level scaling claim. ## Audited outcome (2026-07-22) All six V2-1 records were finite and shared clean source commit `fc8fe99`. The frozen selector chose `eta_A=0.01` for both estimators. Structured calibration improved exact teaching/negative-gradient alignment: | estimator | early-third alignment | all-layer alignment | |:--|--:|--:| | unit targets | 0.001105 | 0.011664 | | channel subspace | 0.007209 | 0.052740 | This is a real 6.5x early-layer and 4.5x all-layer improvement at identical causal-query count, but it fails two frozen advancement checks: early-third alignment is below `0.01`, and its absolute gain over unit targets is `0.006104` rather than `0.01`. The all-layer check passes because the late blocks reach substantially higher alignment; the earliest blocks remain the bottleneck. The recorded target powers (`306.807` for the full hidden field and `0.270893` for channel-basis moments) are intentionally not divided or compared: the two estimators report different metric spaces. V2-1 status is **failed**. V2-2 was not launched, no test endpoint or confirmation seed was touched, and the strict reviewer score remains 5/10.