From a23d761d7330be8d7ac1745be60d4291961087aa Mon Sep 17 00:00:00 2001 From: YurenHao0426 Date: Thu, 6 Aug 2026 15:53:48 -0500 Subject: research: pivot to two-state structured-bias program --- TWO_STATE_BIAS_PROGRAM.md | 280 ++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 280 insertions(+) create mode 100644 TWO_STATE_BIAS_PROGRAM.md (limited to 'TWO_STATE_BIAS_PROGRAM.md') diff --git a/TWO_STATE_BIAS_PROGRAM.md b/TWO_STATE_BIAS_PROGRAM.md new file mode 100644 index 0000000..8933e33 --- /dev/null +++ b/TWO_STATE_BIAS_PROGRAM.md @@ -0,0 +1,280 @@ +# Two-state local learning under structured bias + +## Active paper hypothesis + +The working title remains: + +> **Learning from the Unexpected: Somato-Dendritic Innovations for Local +> Credit Assignment** + +The paper is no longer organized around clean-setting superiority or around +one feedback backbone. Its central hypothesis is: + +> Local learners that update from a difference between two physical or neural +> states are robust to zero-mean noise only after averaging. A deterministic or +> state-dependent differential bias does not average away, eventually dominates +> the shrinking task signal, and becomes more damaging with optimization time, +> flat directions, task count, or the number of biased local populations. A +> neutral, per-cell prediction removes the identifiable component before it is +> consolidated into synapses. + +“Two-state” includes equilibrium propagation (one recurrent network at free +and nudged equilibria), coupled learning, contrastive Hebbian learning, Dual +Propagation, and related positive/negative-state rules. It does not imply that +EP physically contains two networks. Some hardware implementations use two +matched physical copies, which creates an additional differential-mismatch +source. + +## Harnett motivation and unchanged method + +Francioni et al. do not identify raw distal dendritic activity as instruction. +They first remove the component predicted by the normal soma--dendrite +relationship and analyze the per-neuron residual. We apply that operation to a +local teaching channel. + +For local state `z`, observed teaching signal `a`, and eligibility `e`: + +```text +r = a - P(z) +Delta w = eta * r * e +``` + +On an instruction-off observation, the same local predictor uses: + +```text +Delta P = eta_P * (a_neutral - P(z_neutral)) * features(z_neutral)^T +``` + +These equations are frozen. A backbone adapter may expose its local teaching +signal and eligibility, but it may not introduce sparse centered calibration, +oracle targets, copied clean updates, backpropagated gradients, or a new +backbone-specific controller and still call the result SDIL. Neutral +observations, state storage and arithmetic are charged. Baselines receive the +same observation budget when applicable. + +## Unified bias model + +For local population or edge `l`, let the clean two-state teaching estimate be +`s_l`, locally observable neutral state be `z_l`, and measured estimate be + +```text +a_l = s_l + b_l(z_l) + epsilon_l. +``` + +`epsilon_l` is conditionally zero-mean noise. `b_l` is a deterministic or +slowly drifting differential bias. It can arise from unequal physical copies, +sensor offsets, multiplier offsets, asymmetric state readout, or ordinary +state-dependent activity in a mixed dendritic compartment. A common offset +added identically to both states must cancel and is a negative control. + +The claim is conditional. SDIL can remove only bias predictable from its local +neutral information class under a neutral-to-task invariance assumption. It +does not remove arbitrary noise, latent bias, or finite-nudge truncation error +unless that error is observable in the stated neutral channel. + +## Theory program + +The theory must predict experiments, not stop at the standard fact that a +conditional mean minimizes squared prediction error. + +### Non-averaging + +For repeated estimates `g_t + b + epsilon_t`, averaging `T` samples gives + +```text +MSE(mean estimate, mean g) = ||b||^2 + Tr(Sigma) / T. +``` + +Noise vanishes with `T`; fixed bias does not. For distinct parameter blocks, +the squared norm of concatenated local bias is the sum of per-block squared +norms. Under a non-vanishing mean per-block bias, aggregate estimator bias +therefore grows linearly in squared norm with the number of biased populations. +This is an estimator statement, not yet a universal task-loss scaling theorem. + +### Learning dynamics + +Near a solution, use + +```text +dot(theta) = -H(theta - theta_star) + b. +``` + +Positive-curvature modes converge to a displaced solution with displacement +`H^+ b`. Null-space bias produces linear drift `t * Pi_null b`. Alternating +tasks switch `H` and their solution sets, generating bias-dependent periodic +orbits. These quantities give predeclared predictions for error floors, cycle +span and forgetting. + +With residual bias `delta=b-P(z)`, the same expressions replace `b` by +`delta`. Bounds must explicitly include predictor approximation error, +finite-sample error, drift rate and neutral-to-task distribution shift. + +### Instruction preservation + +State and prove the identification condition under which neutral regression +does not subtract instruction. Construct counterexamples for task leakage and +neutral/task shift. Empirically, those counterexamples are mandatory failure +controls. + +### Prior-work boundary + +Ghosh et al. (ICLR 2023) already study how gradient-estimator bias and variance +interact with network size, sparsity, gradient norm and curvature. Laborieux et +al. analyze finite-nudge EP bias; random beta sign cancels its first-order term +in expectation, centered EP removes it with an extra state, holomorphic EP +computes exact finite-amplitude derivatives, and Jacobian homeostasis targets +weight-asymmetry bias. Dillavou et al. already derive physical-bias-driven +null-space drift and introduce overclamping. Our theory must unify local +neutral identifiability with these learning dynamics; it cannot claim that +bias persistence, flat-direction drift, or generic conditional projection was +previously unknown. + +## Evidence ladder + +### P0: published physical evidence, no new learning claim + +Reanalyze the raw experimental data released with Dillavou et al.: + +- demonstrate deterministic, sign-consistent gate-voltage drift and measured + nonzero differential offsets; +- reproduce error/cycle-span trends and standard-versus-overclamping figures; +- show that increasing the averaging window reduces measurement uncertainty + around a nonzero drift rather than driving the drift estimate to zero; +- bind every number to Zenodo record `15692914`, release `v1.0.1`, source tree + `maguzj-imperfect-learning-physical-systems-71b8d72`. + +This stage establishes that the target failure exists in hardware. It does not +show that SDIL fixes it. + +### P1: physical-model mechanism gate + +Use the released small-network equations and measured bias vectors. Compare: + +1. standard coupled learning; +2. same-RMS zero-mean noise; +3. constant per-edge calibration; +4. the authors' overclamping; +5. unchanged two-line SDIL; +6. oracle subtraction. + +Before endpoints, freeze task pairs, bias measurements, predictor features, +neutral cadence and costs. Required plots are bias estimate versus averaging +window, null-space drift, cycle span, error floor and task forgetting. SDIL +must outperform constant calibration when bias is state-dependent and must be +competitive with overclamping after charging observations. Otherwise the +physical branch stops. + +### B1: Dual Propagation adapter + +Finish the already-frozen five-seed differential state-bias confirmation as a +supporting result. Then add a scale axis only after the same-path clean/raw +mechanics pass: miniCNN to VGG-style depth/width using the author code. Common +bias, same-RMS noise, constant calibration and oracle subtraction are required. +Artificial DP bias cannot be the only paper-facing evidence. + +### B2: equilibrium propagation / DCHN adapter + +Use author implementations and unchanged model/training hyperparameters. The +task ladder is FashionMNIST, SVHN, CIFAR-10 and, only after a passed gate, +CIFAR-100. The adapter acts at the local two-state teaching measurement, not +on a backpropagated gradient. + +Required EP baselines depend on the bias source: + +- random-sign beta is the same-cost baseline for finite-nudge first-order + bias; +- centered EP is the extra-state low-variance reference; +- constant offset calibration and oracle subtraction target measurement bias; +- Jacobian homeostasis targets asymmetric dynamics; +- holomorphic EP is discussed or run only where its complex dynamics are a + valid matched alternative. + +Finite-nudge EP is a boundary test unless the unchanged SDIL observation model +actually identifies its bias. Do not design a new staged SDIL variant to beat +random-sign beta. + +### B3: coupled learning and further backbones + +The modern DCHN framework already contains positive, negative and centered EP +and coupled-learning variants on the same five datasets. After P1/B1/B2, use it +to test whether one frozen SDIL adapter transfers between EP and coupled +learning. Contrastive Hebbian learning or Forward--Forward is added only if the +same bias variable and neutral observation have an exact, auditable meaning. + +## Cross-backbone matrix + +Every backbone reports its native clean endpoint, biased raw endpoint, best +bias-specific baseline, SDIL and oracle. “Beat” means a paired advantage under +bias at matched architecture/data/training, together with no meaningful clean +regression and an explicit cost coordinate. It does not mean that SDIL must +beat BP or centered EP on clean accuracy. + +| family | natural/independent bias | strongest required baseline | scale axis | +|---|---|---|---| +| physical coupled learning | measured device/update offset | overclamping, constant calibration | edges, task count, cycle period | +| Dual Propagation | differential state/readout mismatch | clean/raw DP, calibration | CNN depth/width | +| EP/DCHN | measurement offset; separately finite nudge | random-sign beta, centered EP | task difficulty, depth, relaxation | +| coupled-learning DCHN | state/readout mismatch | centered CpL, overclamping analogue | task difficulty, depth | + +The same predictor family, optimizer rule and neutral schedule are frozen +across comparable digital backbones. Backbone-specific retuning of the +debiaser invalidates the “general operation” claim. + +## Main causal matrix + +At fixed corruption RMS cross: + +- zero-mean noise versus fixed, state-dependent and slowly drifting bias; +- common-mode versus differential bias; +- low to high neutral predictability; +- matched versus shifted neutral/task distributions; +- zero versus controlled instruction leakage; +- raw, norm-matched raw, constant calibration, SDIL and oracle. + +The central empirical claim passes only if performance is explained by neutral +predictability rather than by corruption magnitude, update clipping, stronger +nudging, extra observations, or a backbone-specific hyperparameter change. + +## Paper order + +1. Harnett motivation: biological analysis isolates an unexpected per-neuron + dendritic component rather than treating raw dendritic activity as error. +2. Problem: two-state local learning subtracts large nearby quantities; + differential structured bias survives averaging and dominates a shrinking + task contrast. +3. Theory: non-averaging, displaced/flat-direction dynamics, neutral + projection guarantee and failure conditions. +4. Real evidence: released physical-network measurements and predicted drift, + cycles and forgetting. +5. Method: the unchanged two-line local innovation rule. +6. Cross-backbone tests: physical coupled learning, DP, EP and coupled-learning + DCHNs, each against its own strongest correction. +7. Scope: no claim that all local learning fails under all bias, that Harnett + proves plasticity, or that SDIL removes unobservable bias. + +## Advancement and stop rules + +Do not launch the full backbone matrix until P0 is reproduced and P1 shows that +the unchanged rule adds value beyond constant calibration and overclamping. +Stop this paper direction if: + +- physical bias cannot be predicted from a genuinely local neutral variable; +- averaging, constant calibration or overclamping matches SDIL at lower cost; +- gains require a different predictor/controller for each backbone; +- the only positive results use bias generated directly from the predictor + features; +- no predeclared scale variable amplifies raw bias damage; +- realistic neutral/task shift causes instruction removal without a detectable + local certificate. + +## Current status + +- Literature and claim boundary: complete. +- Dillavou artifact: downloaded outside the NFS workspace to + `/scratch/yurenh2/imperfect-learning-physical-systems-v1.0.1`; raw drift, + bow-tie, big-network classification and overclamping data are present. +- Physical P0 reproduction: not yet complete. +- Dual Prop same-path confirmation: active/supporting, not a passed result. +- EP/CpL adapters: not implemented under this bias model. +- Current score for this new paper framing: 5/10 until P1 passes. + -- cgit v1.2.3