diff options
| author | Oscar Wan <oscarwan@stanford.edu> | 2026-07-21 10:48:52 -0700 |
|---|---|---|
| committer | Oscar Wan <oscarwan@stanford.edu> | 2026-07-21 10:48:52 -0700 |
| commit | c4a6b459cdd5b183220cf62996f61efc8b56620d (patch) | |
| tree | 78da709a17b24f274ff5f89ad7faf34bacefce4c /docs/campaign/FW135M_BP_BASELINE.md | |
| parent | ca209ee5df16174d2137b5d5922581fc2f6d7e88 (diff) | |
add 135M baseline
Diffstat (limited to 'docs/campaign/FW135M_BP_BASELINE.md')
| -rw-r--r-- | docs/campaign/FW135M_BP_BASELINE.md | 89 |
1 files changed, 89 insertions, 0 deletions
diff --git a/docs/campaign/FW135M_BP_BASELINE.md b/docs/campaign/FW135M_BP_BASELINE.md new file mode 100644 index 0000000..3f81ba8 --- /dev/null +++ b/docs/campaign/FW135M_BP_BASELINE.md @@ -0,0 +1,89 @@ +# FW135M BP Twin Run Sheet + +Source of truth: `docs/BASELINE_SPEC.md`. + +Status: commands configured on branch `xiang`; no training has started. + +## First width-only rung + +- Shape: `L12/C768/H12/T256` +- Head dimension: `64` +- Exact local parameter count: `135,303,936` +- Data: existing FineWeb-Edu 32k bins +- Batch: `B24` +- Target: `20N = 2,706,078,720` tokens +- Complete-batch exposure: `2,706,081,792` tokens +- Trainer argument: `--steps 440442` (440,443 inclusive updates) + +## What changes from 72M + +- Width: `512 -> 768` +- Heads: `8 -> 12` +- Parameters: `72.11M -> 135.30M` +- Run length: increases to preserve approximately 20 tokens/parameter + +## What stays identical + +- L12 depth and 64-dimensional heads +- T256 and B24 +- FineWeb-Edu data and 32k tokenizer +- OLMo2-style model implementation +- Muon hybrid optimizer and cosine schedule +- Weight decay 0.1 and BF16 autocast +- Seed list and validation cadence +- Existing `casc_bp_train.py` code path + +## LR sweep + +At this smallest new rung only: + +- `7e-4` +- `1e-3` +- `1.4e-3` + +Only Adam-side LR changes. Carry the winner to larger rungs. + +No result may be quoted with fewer than two seeds. Record best validation CE, final-step CE, and tail median. Save every 5,000 steps and log validation every 100 steps. + +## Dry-run launcher + +From `ep_run/`: + +```bash +python -m py_compile fw135m_baseline.py +python fw135m_baseline.py --mode smoke +python fw135m_baseline.py --mode sweep +python fw135m_baseline.py --mode full +``` + +Commands print by default. Add `--execute` only when ready. + +Smoke: + +```bash +python fw135m_baseline.py --mode smoke --execute +``` + +LR sweep: + +```bash +python fw135m_baseline.py --mode sweep --execute +``` + +Selected full run: + +```bash +python fw135m_baseline.py --mode full --lr <selected-lr> --seed 1 --execute +python fw135m_baseline.py --mode full --lr <selected-lr> --seed 2 --execute +``` + +W&B remains off until configured. Later: + +```bash +python fw135m_baseline.py --mode full --lr <selected-lr> \ + --wandb_project <project> --execute +``` + +## Final EP/BP comparison + +The EP arm must use the same shape, data, B/T, steps, optimizer family, selected shared settings, seeds, and evaluation cadence. Only the training rule and EP-specific `β` controls differ. |
