summaryrefslogtreecommitdiff
path: root/CROSS_ARCHITECTURE_CROSSOVER.md
blob: eb3dfbe33571a702620e0a81d0a31b876b764f0a (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
# Cross-architecture local-learning crossover

## Purpose and claim boundary

This protocol tests two claims without selecting architectures or methods by
their observed endpoints:

1. SDIL lies on the accuracy--training-cost Pareto frontier among matched
   local-learning methods.
2. SDIL retains useful optimization as the size of a standard architecture
   family increases.

The primary development panel contains nine architecture--scale points and
nine training methods, hence 81 cells.  A failed, nonfinite, chance-level, or
timed-out cell is a result and remains in the panel.  A method cannot be
removed from larger models because it was weak on a smaller model, and a size
cannot be collapsed into an architecture-family average.  Smoke runs verify
mechanics only and never replace a complete cell.

This panel is a matched crossover, not a claim that all methods have equal
wall time or implement the same inference process.  Accuracy, wall time,
forward-equivalent work, peak memory, task-loss queries, relaxation steps, and
method-specific observations are reported separately.  Native author results
remain in `BASELINES.md` and are never silently mixed with matched cells.

## Frozen cell registry

The method axis is:

| ID | Method |
|---|---|
| `bp` | exact reverse-mode backpropagation |
| `fa` | ordinary layer-recursive feedback alignment |
| `dfa` | direct feedback alignment |
| `pepita` | PEPITA |
| `ff` | Forward-Forward |
| `ep` | equilibrium propagation |
| `dualprop` | Dual Propagation |
| `clean_kp` | reciprocal Kolen--Pollack learning without nuisance traffic |
| `sdil` | dynamic somato-dendritic innovation learning |

`fa` is not DFA and is not a learned feedback rule.  Its feedback tensors are
independent fixed random draws.  On a residual DAG it follows the actual
parent--child graph, including exact parameter-free shortcut and local
normalization/activation Jacobians, but it never reads or copies a downstream
forward convolution.  This is the existing hierarchical FA mechanism under
the paper-facing ordinary-FA label.  `clean_kp` and `sdil` use the same
reciprocal credit-transport substrate; SDIL alone receives the frozen
four-RMS soma-predictable traffic and paired-neutral subtractive innovation
operation.

The architecture--scale axis is:

| Family | ID | Frozen forward topology |
|---|---|---|
| plain CNN | `minicnn` | author miniCNN: C64-MP, C64-MP, linear-10 |
| plain CNN | `vgglike` | author VGGlike: C128-MP, C256-MP, C512-MP, C512-MP, linear-10 |
| plain CNN | `vgg16` | author CIFAR VGG16: 13 convolutions with 64/128/256/512 stages, five max-pools, dense 4096, dense 4096, linear-10 |
| residual CNN | `resnet20` | standard CIFAR `6n+2` ResNet, base width 16 |
| residual CNN | `resnet32` | standard CIFAR `6n+2` ResNet, base width 16 |
| residual CNN | `resnet56` | standard CIFAR `6n+2` ResNet, base width 16 |
| decoder Transformer | `transformer4` | four decoder blocks |
| decoder Transformer | `transformer8` | eight decoder blocks |
| decoder Transformer | `transformer12` | twelve decoder blocks |

The plain-CNN definitions are copied from the frozen Dual Propagation author
revision `7b2595b34421e1483a721dbfdeff8cdabda3a1ff`, not reconstructed from a
name.  Its VGG16 has convolutional widths
`64,64,128,128,256,256,256,512,512,512,512,512,512`, pooling after
convolutions 2, 4, 7, 10, and 13, then two 4096-unit ReLU layers.  The
residual-CNN definitions are the already audited `CIFARLocalResNet`
implementations.

The Transformer depth axis, task, width, context, algorithm adapters, and
validation-only selector are frozen in `TRANSFORMER_CROSSOVER.md`.  That
protocol specifies continuous-embedding PEPITA, autoregressive candidate
overlay for Forward-Forward, and relaxable block states for EP and Dual
Propagation rather than relabelling a generic local cross-entropy surrogate.
Its deterministic X0 gate is frozen at
`results/transformer_crossover/x0_equation_audit.json`.

Additional width or depth ladders may be added only as a new complete family
revision containing every method.  They do not replace, merge, or selectively
extend the 81 primary cells.

## Matched family rules

Plain CNN and residual CNN use CIFAR-10 with the same normalized images,
augmentation, train/validation split, minibatch order, initial forward
parameters, and evaluation batches within an architecture--seed pair.
Transformer methods use the same tokenized data bytes, sampled token
positions, forward initialization, and evaluation batches within a
depth--seed pair.  The following rules hold in every family:

- the forward topology and number of trainable forward parameters are matched;
- BP uses exact reverse mode, while non-BP methods may use autograd only to
  evaluate explicitly local Jacobian-vector products whose boundaries are
  detached and audited;
- FA/DFA random feedback never changes; KP feedback changes only through its
  reciprocal local correlation and never reads a forward weight or update;
- SDIL uses the same KP update plus the already frozen neutral observation,
  predictor, traffic, and innovation rules;
- every method receives the same number of ordinary training examples and the
  same endpoint-selection split; extra queries, forwards, perturbations, or
  inference iterations are charged rather than hidden;
- method-specific inference hyperparameters follow the cited paper or author
  implementation when available.  A bounded validation-only selection grid
  must be committed before its endpoints and applied at all three sizes;
- no test example is used for tuning, early stopping, divergence recovery, or
  choosing a method variant.

Forward-Forward, EP, and Dual Propagation necessarily have method-specific
state or example construction.  Their matched cells preserve the frozen
forward parameterization and task, while the native-author ledger separately
shows executable fidelity.  Any unavoidable semantic deviation is named in
the family protocol and table rather than described as exact reproduction.

## Stages and completeness

### X0: executable-rule audit

Before a formal endpoint, each family adapter must pass deterministic CPU or
one-minibatch tests that establish:

- identical forward outputs and parameter counts across method wrappers at
  initialization;
- exact BP agreement in an audit-only symmetric-feedback limit for FA/KP;
- independence of actual FA/DFA feedback from forward weights;
- absence of reverse-graph leakage across detached local boundaries;
- paper-equation agreement for PEPITA, Forward-Forward, EP, and Dual
  Propagation on a small deterministic network;
- clean-KP/SDIL identity when traffic is disabled and innovation prediction is
  zero;
- SDIL raw, matched-raw, and innovation signal identities from
  `KP_TEACHING_SIGNAL_ABLATION.md`;
- finite cost counters whose analytic values agree with enumerated operations.

One-epoch or short-model GPU smokes may check compilation, memory, and runtime.
They cannot select learning rates, inference steps, architecture cells, or
reported endpoints.

### X1: complete validation matrix

Every family freezes its optimizer, schedule, validation-only hyperparameter
grid, cap on inference/relaxation work, seed, split, hardware policy, and
failure policy before X1.  X1 is complete only when all 81 registered cells
have an immutable record from that frozen protocol.  The analyzer must reject
duplicates, missing cells, untracked or dirty source, architecture drift,
dataset drift, and method-dependent ordinary-example counts.

The primary metric is validation top-1 accuracy for CIFAR-10 and validation
negative log-likelihood/perplexity for autoregressive language modeling.
Each record also includes final and best-validation values, the selection
rule, wall time, peak allocated and reserved memory, total forward-equivalent
examples, affine MACs or FLOPs, task-loss queries, local observation counts,
relaxation passes, first nonfinite step, source revision, environment freeze,
dataset hashes, and physical GPU identity.

There is no numerical pass threshold for retaining a method in X1.  Timeout
caps are architecture- and method-independent functions of the frozen BP
runtime plus a predeclared multiplier; a timeout is plotted at its achieved
metric and charged elapsed cost.  Nonfinite cells are plotted at chance (or
their last finite language loss) with a failure marker, not omitted.

### X2: complete confirmation matrix

X2 uses untouched seeds and the X1-selected settings without a new grid.  The
number and identities of confirmation seeds are committed before any X2
endpoint.  Every one of the 81 cells receives the same seed set; a partial
SDIL-only or winner-only confirmation cannot support the crossover claim.
Test is evaluated once per completed confirmation record.  X1 and X2 remain
separate in all tables.

## Paper-facing analyses

The crossover produces two prespecified views:

1. An accuracy/loss versus measured-cost plot containing every method,
   architecture size, timeout, and divergence marker.  A method is on the
   empirical Pareto frontier only if no observed point is at least as accurate
   and no more costly, with one strict inequality.  Wall time and
   hardware-independent work receive separate panels.
2. A within-family scaling plot showing the paired change from smallest to
   largest size for every method.  Raw endpoints, parameter counts, and
   teaching/negative-gradient alignment by depth are shown; family averages
   cannot hide a failed largest model.

The strongest acceptable conclusion is conditional: under the frozen tasks,
budgets, and adapters, SDIL is Pareto-competitive and scales across the tested
families.  The panel does not prove asymptotic scaling, biological
implementation, clean-setting superiority over BP, or universal superiority
over EP/Dual Propagation.

## Execution ledger (2026-07-27)

- Plain-CNN P1 is complete and its selector is frozen.  The 27-cell P2 is
  active on the only authorized physical GPUs, timan107 GPU5 and GPU7.
- ResNet X0, the 19-cell ResNet-20 selector, and its analyzer are frozen.
  Formal R1 waits for an authorized GPU rather than sharing memory with P2.
- Transformer X0 passes in both development PyTorch and the Pascal PyTorch
  2.3 environment.  Its 27-cell T1 registry and analyzer are frozen.  Formal
  T1 likewise waits for an authorized GPU.

GPU serialization changes calendar time, not the cell registry.  No pending
family may borrow an unapproved GPU, and no method or size may be removed to
shorten the queue.