summaryrefslogtreecommitdiff
path: root/ORAL_A_V2.md
blob: f77c5d288faff3de9a6fabd3b560b5aaaac200ba (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
# Oral-A-v2: representable-subspace development protocol

## Status and claim boundary

This is a transparent post-failure development branch.  Frozen Oral-A-v1
failed when its SDIL run became nonfinite at epoch 89; it remains failed and is
never relabeled.  The executable diagnosis found that its unitwise K1 causal
target was not captured by the translation-shared feedback model.  The v2
change is restricted to that diagnosed interface: perturb the two basis fields
that the channel-gated vectorizer can actually express, then apply the same
expected full-field delta rule.  Architecture, task update, causal-query count,
cadence, sigma, data split, and v1 optimization schedule do not change.

No v2 GPU endpoint was observed before this protocol and its selectors were
committed.  The CIFAR-10 test set and confirmation seeds 10--14 remain
untouched.  A v2 result cannot be pooled with v1 or described as preregistered
before the v1 failure.

## V2-0: mechanics and mathematical identity

`experiments/conv_local_smoke.py` must establish all of the following on CPU:

- structured moment-estimator cosine above 0.985 and norm ratio in [0.90,1.10];
- antithetic directional derivative relative error below `2e-6` at a
  differentiable fixed ResNet point;
- absolute error below `1e-14` between the implemented A/G update and the
  analytically expected full-field delta-rule update;
- all legacy convolutional, BatchNorm, predictor, and local-eligibility checks
  remain green.

This stage passed before the GPU funnel was opened: cosine `0.998217`, norm
ratio `1.006989`, JVP relative error `1.69e-10`, and delta-rule absolute error
`2.71e-20`.

## V2-1: frozen-forward causal-capture screen

Use seed 0 ResNet-20, the frozen 45,000/5,000 split, the first 10,000 training
examples, batch 128, and 400 feedback-only minibatches.  Forward weights, BN
state and affine parameters, and the readout remain bitwise fixed.  The
vectorizer is `channel_gated`, starts at exactly zero (`a_scale=0`), and uses
K1 antithetic calibration with `sigma=0.01`.  Cross:

- estimator: legacy `unit_targets`, v2 `channel_subspace`;
- apical rate: `0.01`, `0.1`, `1.0`.

All candidates use identical data, perturbation seed 1000, and a fixed
64-example training-prefix exact-gradient audit.  Autograd is used only after
calibration to audit teaching alignment; it never updates a parameter.

Within each estimator select maximum early-third teaching/negative-gradient
cosine, then maximum all-layer cosine, then lower apical rate.  V2 advances
only if all records are finite and the selected structured estimator has:

1. early-third cosine at least `0.01`;
2. all-layer cosine at least `0.01`;
3. early-third cosine at least `0.01` above the best matched unit-target run.

The estimator-specific prediction/target cosine is reported, but it is not
compared across modes because one record is measured in the full hidden field
and the other in its two channel-basis moments.  No rate or threshold is added
after results are observed.

## V2-2: full ResNet-20 validation gate

Only if V2-1 passes, run one seed-0 SDIL model for 200 epochs on all 45,000
development-training examples.  Copy Oral-A-v1 exactly: ResNet-20 width 16,
batch 128, hidden LR `0.03`, output LR `0.1`, momentum `0.9`, weight decay
`1e-4`, and 10x drops at epochs 100 and 150.  Use the V2-1-selected apical
rate, zero-initialized channel-gated feedback, 400 feedback-only warmup steps,
and structured K1/every-4 calibration at `sigma=0.01`.

The already frozen BP (`91.62%`) and DFA (`33.06%`) records remain the matched
references.  V2 passes only if:

1. every loss and metric is finite;
2. final validation accuracy is within 5 points of BP and at least 2 points
   above DFA;
3. early-third teaching alignment is at least `0.05`;
4. estimated total training MACs do not exceed BP.

Failure closes v2; there is no LR, warmup, cadence, sigma, direction-count, or
vectorizer recovery branch.

## V2-3: untouched depth confirmation

Only after V2-2 passes, copy its complete hyperparameters without depth tuning
to ResNet-20/32/56 and model/data-loader seeds 10--14.  Refit on all 50,000
training examples.  Each run evaluates CIFAR-10 test exactly once at the end;
no intermediate test endpoint is logged.  Exact BP and the v1-selected DFA are
run on the identical panel.

The confirmation gate is unchanged from Oral-A-v1: all 45 trajectories finite;
SDIL within 2 points of BP at every depth; SDIL at least 2 points above DFA at
depth 56; paired SDIL depth-20 to depth-56 change no worse than -2 points;
depth-56 early-third alignment at least `0.05` and at least 30% of depth-20;
and SDIL nondominated in accuracy versus estimated MACs, causal queries, and
peak memory.  Wall time is descriptive only.

## Reviewer-score rule

V2-0 and V2-1 cannot raise the strict ICLR score because they establish
mechanics and causal capture rather than standard-scale task success.  A V2-2
pass can move the score from 5 only after its full audit is committed.  A V2-3
multi-depth, multi-seed pass is required for an oral-level scaling claim.

## Audited outcome (2026-07-22)

All six V2-1 records were finite and shared clean source commit `fc8fe99`.
The frozen selector chose `eta_A=0.01` for both estimators.  Structured
calibration improved exact teaching/negative-gradient alignment:

| estimator | early-third alignment | all-layer alignment |
|:--|--:|--:|
| unit targets | 0.001105 | 0.011664 |
| channel subspace | 0.007209 | 0.052740 |

This is a real 6.5x early-layer and 4.5x all-layer improvement at identical
causal-query count, but it fails two frozen advancement checks: early-third
alignment is below `0.01`, and its absolute gain over unit targets is `0.006104`
rather than `0.01`.  The all-layer check passes because the late blocks reach
substantially higher alignment; the earliest blocks remain the bottleneck.
The recorded target powers (`306.807` for the full hidden field and `0.270893`
for channel-basis moments) are intentionally not divided or compared: the two
estimators report different metric spaces.

V2-1 status is **failed**.  V2-2 was not launched, no test endpoint or
confirmation seed was touched, and the strict reviewer score remains 5/10.