summaryrefslogtreecommitdiff
path: root/CONTRASTIVE_BIAS_CONFIRMATION.md
diff options
context:
space:
mode:
Diffstat (limited to 'CONTRASTIVE_BIAS_CONFIRMATION.md')
-rw-r--r--CONTRASTIVE_BIAS_CONFIRMATION.md60
1 files changed, 60 insertions, 0 deletions
diff --git a/CONTRASTIVE_BIAS_CONFIRMATION.md b/CONTRASTIVE_BIAS_CONFIRMATION.md
new file mode 100644
index 0000000..d20b3f4
--- /dev/null
+++ b/CONTRASTIVE_BIAS_CONFIRMATION.md
@@ -0,0 +1,60 @@
+# Same-path contrastive-bias confirmation
+
+## Status and purpose
+
+This is a post-screen confirmation frozen after the B1 results were known. It
+is not the untouched B2 in `CONTRASTIVE_BIAS.md`. B1 showed the intended raw
+failure and innovation recovery at every activity-bias ratio, but its complete
+gate failed because ordinary clean DP ended at 69.66% instead of 70% and a
+common-bias run differed from it by 2.28 accuracy points despite exactly zero
+teaching-difference error.
+
+This confirmation removes that control mismatch. All four conditions use the
+same explicit-difference training path. All four conditions for one seed run
+sequentially on the same physical GPU.
+
+## Frozen panel
+
+The implementation, upstream revision, model, data split, batch order, and
+local bias equations are unchanged from B1. The five new model/minibatch seeds
+are `1989, 1990, 1991, 1992, 1993`. Every run uses author Dual Prop miniCNN on
+CIFAR-10, `alpha=0`, `beta=0.1`, `fwK`, 16 inference passes, batch size 100,
+and the complete 130-epoch author schedule: peak learning rate 0.025, warm-up
+learning rate 0.001, 10 warm-up epochs, decay epoch 120, and final learning
+rate `2e-6`. Test is not evaluated.
+
+For every seed, the four ratio-4 conditions are:
+
+1. `same_path_clean`: common activity bias with the raw rule. Its compartment
+ difference is exactly zero, so this is clean DP through the same code path.
+2. `raw`: differential activity bias with the raw rule.
+3. `innovation`: differential activity bias with the local neutral affine
+ subtraction.
+4. `oracle`: differential activity bias with exact generated-bias subtraction.
+
+The ratio is calibrated once at initialization and frozen. A raw run that
+becomes nonfinite is retained. Clean, innovation, and oracle must complete all
+130 epochs. Test remains untouched until the full validation panel and audit
+exist.
+
+## Confirmation gate
+
+The panel passes only if all source, registry, hardware, split, and test
+isolation checks pass and:
+
+- same-path clean, innovation, and oracle are finite for 130 epochs in all five
+ seeds, and mean same-path-clean final validation accuracy is at least 80%;
+- raw is nonfinite or at least 20 accuracy points below same-path clean in
+ every seed;
+- innovation is above raw in every seed and the one-sided 95% lower bound on
+ the paired gain is above 20 points;
+- the one-sided 95% upper bound on the paired clean-minus-innovation deficit is
+ below 2 points;
+- the two-sided 95% bound on the absolute mean innovation-minus-oracle gap is
+ below 2 points;
+- innovation's post-subtraction bias RMS is at most `1e-3` of raw bias in every
+ run, and every predictor reports zero task-instruction observations;
+- every seed's four conditions record the same physical GPU UUID.
+
+Failure is retained and narrows the claim. Passing permits a separately
+frozen, once-only test evaluation; it does not retroactively make B1 pass.