diff options
| author | yurenh <blackhao0426@gmail.com> | 2026-08-31 18:20:12 -0500 |
|---|---|---|
| committer | yurenh <blackhao0426@gmail.com> | 2026-08-31 18:20:12 -0500 |
| commit | 8acf3f94c630c0ef2cc89c56d39464c42a0d5e3a (patch) | |
| tree | c3cec35990ce885b4d6000b2c13240cca7c419c3 /README.md | |
| parent | 17a81b9c86cfedd70812a0e83f33798b64c1678e (diff) | |
gitignore data/runs; data prep documented as on-node
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GkgLsACEF6CCP7EUfA5fZe
Diffstat (limited to 'README.md')
| -rw-r--r-- | README.md | 3 |
1 files changed, 2 insertions, 1 deletions
@@ -16,8 +16,9 @@ product are digital. See the main zobp repo for the method, theory (NSR ≈ c·d - `configs/` — model sizes (60m/124m/350m/1b) x training arms (bp / zbp_n16 / zbp_n4) ## Run +Data prep runs **on the training node** (H200), not on a dev machine; `data/` and `runs/` are gitignored. ``` -python scripts/prepare_data.py --dataset fineweb-edu --tokens 3e9 --out data/fineweb +python scripts/prepare_data.py --tokens 3e9 --out data/fineweb # on the H200 node torchrun --nproc_per_node=8 scripts/train.py --model configs/model/m124.yaml --train configs/train/zbp_n16.yaml ``` Global batch is fixed in the train config; per-rank micro-batch and accumulation adapt to world size. |
