From 70e545c95f32275d55b132ecd1c5ae167db5e258 Mon Sep 17 00:00:00 2001 From: YurenHao0426 Date: Thu, 6 Aug 2026 16:05:49 -0500 Subject: docs: define ICLR 2027 structured-bias roadmap --- ICLR_2027_REPLAN.md | 272 ++++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 272 insertions(+) create mode 100644 ICLR_2027_REPLAN.md diff --git a/ICLR_2027_REPLAN.md b/ICLR_2027_REPLAN.md new file mode 100644 index 0000000..d5444b6 --- /dev/null +++ b/ICLR_2027_REPLAN.md @@ -0,0 +1,272 @@ +# ICLR 2027 replan: local learning under structured bias + +## Decision + +Keep the title **Learning from the Unexpected: Somato-Dendritic Innovations +for Local Credit Assignment**, but replace the old clean-scaling claim with a +more defensible claim: + +> Two-state local learners can average away zero-mean noise, but not a +> differential bias in their local teaching measurements. This bias produces +> error floors, drift and task-switching cycles that can worsen as more local +> populations are biased. A somato-dendritic innovation removes the locally +> predictable component without backpropagation. + +The method remains frozen: + +```text +r = a - P(z) +Delta w = eta * r * eligibility +``` + +`P` is an affine or fixed-basis local adaptive filter trained by LMS on an +instruction-off observation. No backbone-specific recovery stage, copied clean +update, centered target, global loss or backpropagated signal may train it. + +The debiasing mechanism is already de-risked in synthetic experiments. The +paper is not yet de-risked: the main remaining question is whether the same +locally identifiable operation beats strong bias-specific corrections on a +measured physical model and transfers unchanged across local-learning +backbones. + +## What the paper must and must not claim + +The paper may claim that structured differential bias does not disappear with +averaging and that its aggregate estimator norm can grow with the number of +biased local blocks. It may claim task-level scaling only under explicit +curvature, bias-alignment and identifiability assumptions, followed by the +corresponding experiment. It must not claim that EP, Dual Propagation or local +learning in general cannot scale cleanly. + +The comparison target is each backbone's strongest correction for the stated +bias, not BP accuracy in a clean setting. Clean BP and clean native training +are reference points. A result counts as SDIL only after it passes +`HARDWARE_LOCALITY_CONTRACT.md`. + +## Critical path and hard gates + +### Gate 0: freeze the measurement and adapter contract + +Deliverables: + +- one adapter API exposing local neutral state, local teaching measurement and + native local eligibility; +- one shared affine/fixed-basis predictor and neutral schedule across comparable + digital backbones; +- no-grad, local-replay, downstream-independence, instruction-isolation, + common-mode and phase/storage tests; +- a cost ledger counting phases/observations, stored local scalars, local + multiplies, wall time and peak memory; +- exact definitions of common-mode bias, differential fixed bias, + state-dependent bias, slow drift and same-RMS zero-mean noise. + +Do not start expensive scale runs until this gate passes. This prevents a +software-only implementation or a backbone-specific correction from becoming +the reported method. + +### Gate 1: P1 measured-physical mechanism + +Recreate the released small-network coupled-learning dynamics using the +published equations and measured bias traces. Freeze the released task pairs, +bias traces, neutral cadence and predictor basis before comparing: + +1. standard coupled learning; +2. same-RMS zero-mean noise; +3. constant per-edge calibration; +4. overclamping; +5. unchanged SDIL; +6. oracle subtraction. + +Sweep cycle period and bias/drift strength over every compatible released +small-network instance. Report error floor, null-space drift, cycle span, +forgetting and charged cost. Use simulator initializations for uncertainty; +never present them as new independent hardware samples. + +Pass conditions: + +- structured bias remains after temporal averaging while matched zero-mean + noise decreases; +- SDIL closes a substantial fraction of the raw-to-oracle gap on every + identifiable state-dependent-bias setting; +- SDIL beats constant calibration when the bias varies with local state; +- SDIL is nondominated with overclamping on error/cycle span versus charged + observation and circuit cost; +- common-mode bias gives no artificial SDIL advantage; +- task leakage and neutral-to-task shift fail in the direction predicted by + the identifiability theory. + +If SDIL loses to constant calibration or overclamping without a cost or +generality advantage, stop the physical-first paper and reassess before using +more GPU time. + +### Gate 2: frozen-adapter transfer + +Use the identical predictor family, update rule and neutral schedule in all +paper-facing digital experiments. Only the native local teaching measurement +and eligibility change with the backbone. + +**Dual Propagation.** Complete the frozen five-seed miniCNN confirmation, then +run MiniCNN, VGG-style middle scale and author VGG16. The main conditions are +native clean, raw differential bias, constant calibration, SDIL and oracle; +same-RMS noise is the causal control. Use one seed to screen mechanics, three +seeds for the accept matrix and five only for final claims. One existing author +VGG16 run takes about 23,120 seconds (6.4 hours) on a GTX 1080, so VGG16 cells +should move to the A6000 after the one-seed gate. + +**Equilibrium propagation.** Use the author DCHN implementation unchanged for +native dynamics and hyperparameters, with a hand-written paper-facing local +adapter. First test measurement/readout bias, for which constant calibration +and oracle subtraction are matched baselines. Separately test finite-nudge +bias against random-sign beta and centered EP. Do not claim SDIL solves +finite-nudge bias unless the neutral observation actually identifies it. + +**Coupled-learning DCHN.** After EP passes, transfer the same adapter within the +same codebase. Compare with constant calibration, centered coupled learning and +an overclamping analogue. This is the cleanest digital bridge to the measured +physical experiment. + +The accept matrix stops at three seeds and three families: measured physical +coupled learning, Dual Propagation, and EP. Digital coupled learning is added +as soon as the same adapter transfers; it is not allowed to delay the first +complete paper draft. + +Pass conditions: + +- the raw biased endpoint worsens with the declared scale axis in at least two + independent families; +- neutral predictability, rather than corruption RMS, predicts recovery; +- SDIL improves every biased raw endpoint and beats constant calibration on + state-dependent bias; +- SDIL is competitive with each family's strongest correction after cost is + charged; +- the frozen adapter passes the executable BP-free audit in every family; +- clean/common-mode controls show that the gain is not ordinary regularization + or a changed optimizer. + +Failure on one family narrows the scope. It must not trigger a new +backbone-specific SDIL variant. + +## Theory package + +The theory and experiments share one bias model, + +```text +a_l = s_l + b_l(z_l) + epsilon_l. +``` + +The accept version needs five results: + +1. a bias--variance decomposition showing that averaging removes the + conditionally zero-mean term but leaves the predictable differential term; +2. an aggregate-block result, carefully stated as estimator scaling rather + than a universal loss theorem; +3. local quadratic dynamics predicting displaced optima, null-space drift and + alternating-task cycle span; +4. a residual bound decomposing predictor approximation, sample error, drift + and neutral-to-task shift; +5. an instruction-preservation condition plus explicit leakage and + distribution-shift counterexamples. + +Each theoretical quantity must appear on an experimental axis. A theorem that +does not predict a plotted transition, slope or failure control stays in the +appendix or is removed. + +## Main figures + +1. **Problem and real evidence:** Harnett residual motivation, the two-state + measurement model, and the released physical error plateau/cycle drift. +2. **Mechanism on the physical model:** matched noise versus bias; raw, + calibration, overclamping, SDIL and oracle; error/cycle span against charged + cost. +3. **Scaling across backbones:** performance gap and residual bias versus depth, + width, biased-block count and task difficulty for physical learning, DP and + EP/CpL. +4. **What makes recovery possible:** neutral predictability, drift rate, + task leakage and distribution shift, followed by an accuracy--cost Pareto + summary. + +The key scaling plot must show both the failure term and its removal. Plotting +only final accuracy across larger clean architectures is not evidence for the +new claim. + +## Three publication bars + +### Accept bar: target reviewer score 6--7 + +- Gates 0--2 pass for the measured physical model, DP and EP; +- all main comparisons have three seeds or every released physical replicate; +- the five theory results predict the main experiments; +- the method remains two lines and executable without BP; +- strong baselines, cost accounting, negative controls and failure cases are + present; +- the complete manuscript makes the conditional claim, not a universal + anti-EP claim. + +At this point the paper has a new problem, a biologically motivated local +operation, real measured evidence, causal controls and cross-backbone transfer. + +### Oral-B bar: target reviewer score 7--8 + +- extend EP and coupled-learning DCHNs across FashionMNIST, SVHN and CIFAR-10, + multiple depths/biased-block counts and five seeds; +- finish five-seed VGG16 DP and the full cost Pareto comparison; +- cross fixed, state-dependent and slowly drifting bias with matched noise, + predictability and neutral/task shift; +- keep one frozen predictor configuration across comparable backbones; +- show that measured or independently fitted bias parameters predict the + observed scaling curves without post-hoc tuning. + +This bar is breadth plus prediction: the same mechanism must explain when SDIL +works, how much it recovers and when it fails. + +### Oral-A bar: target reviewer score 8--9 + +- demonstrate the local filter and subtraction path in circuit simulation or + on physical hardware, or validate on a second independently measured + substrate; +- complete long large-scale runs such as CIFAR-100/large DCHNs only after the + adapter is frozen; +- compare end-to-end phase, storage, precision, time and energy costs rather + than only optimizer steps; +- obtain a prospective prediction: select bias/predictability conditions from + small systems, then correctly predict recovery and failure on the held-out + larger system. + +These long runs may use spare GPUs in parallel after Gate 0, but their design +must not change in response to their final endpoints. + +## Execution order and rough budget + +1. **Now (hours):** let the existing DP C1 run finish; audit all five seeds and + commit only the completed frozen result. +2. **Next 1--3 working days:** implement and test the released small physical + model, finish Gate 0, and run the cheap P1 CPU sweep. +3. **Following week:** write theory v0 beside the P1 plots; run one-seed DP + scale and EP screens. Reject broken adapters before replication. +4. **Following 1--2 weeks:** run the three-seed accept matrix on the 1080s and + OscarWX's A6000; build the four main figures and write a full paper draft. +5. **After a reviewer-style audit reaches 6:** launch five-seed/multi-dataset + Oral-B jobs. Start Oral-A long runs opportunistically on otherwise idle + cards. + +Expected GPU exposure is roughly 40--100 hours for screened accept-scale +digital runs and several hundred GPU hours for Oral-B breadth. These are +planning ranges, not promises: actual per-cell time must be recorded from the +one-seed screens before the full launch. P1 and theory are the immediate +critical path; adding GPUs cannot compensate for failure there. + +## Reviewer-score ledger + +Update one evidence table after every completed gate: + +| item | evidence | strongest alternative | unresolved failure | score effect | +|---|---|---|---|---| +| real bias exists | released physical P0 | overclamping paper | no SDIL fix yet | necessary, not sufficient | +| physical mechanism | P1 | calibration/overclamping | pending | unlocks 6 | +| general transfer | DP + EP/CpL | native correction | pending | supports 6--7 | +| scaling prediction | theory + held-out scales | descriptive curve | pending | supports 7--8 | +| hardware realization | circuit/second substrate | primitive argument | pending | supports 8--9 | + +Every progress report should state positive results, negative results, compute +and a fresh reviewer score. Accuracy without the matched strongest baseline, +cost coordinate or frozen protocol does not move the score. -- cgit v1.2.3