From 35728eff3e359804c7c74adba8ea06d7f71b3739 Mon Sep 17 00:00:00 2001 From: YurenHao0426 Date: Sat, 29 Aug 2026 19:58:00 -0500 Subject: protocol: freeze high-budget calibration stress test --- CLLN_SCALING.md | 15 +++++++++++++++ 1 file changed, 15 insertions(+) (limited to 'CLLN_SCALING.md') diff --git a/CLLN_SCALING.md b/CLLN_SCALING.md index be98543..c3adb44 100644 --- a/CLLN_SCALING.md +++ b/CLLN_SCALING.md @@ -97,3 +97,18 @@ as independent confirmations. 5. The scaling conclusion is computed from all frozen sizes and task clusters, including failed runs. 6. Hardware-realistic claims use Part 3 only. + +## Frozen calibration-budget stress test + +After the six-size core confirmation, run one supplementary control at side +length 32. Reuse all 40 tasks, component seeds `20260830,20260831,20260832`, +600 epochs, and learning exposure `0.01 s`. Train only the static-calibration +method, increasing its instruction-off calibration observations from 16 to +256. No learning hyperparameter or component draw is reselected. + +The output is `results/coupled_ladder/p4_constant256_side32.json`. Compare its +final error, stable-zero fraction, and error AUC with the frozen 16-observation +static calibration in `p2_confirm_side32.json`. This control tests whether the +static baseline was limited by sampling error. It is supplementary and does +not replace any point in the frozen scaling fit. The 256 calibration reads per +edge are included in its observation cost. -- cgit v1.2.3