# zbp-scaling Scaling study for **ZBP (zeroth-order backpropagation)** vs BP on decoder-only language models (part 2 of the ZBP paper). Every nonlinear block is treated as a physical black box trained from forward queries only (score-space attention core + per-token SwiGLU FFN); linear maps and the score product are digital. See the main zobp repo for the method, theory (NSR ≈ c·d/n) and part-1/3 results. ## Install ``` pip install -e . # or: pip install -e .[dev] && pytest tests/ ``` ## Layout - `src/zbp_scaling/zbp/` — vendored ZBP package (probes, estimators, ZBPBlock autograd) - `src/zbp_scaling/model.py` — OLMo2-style transformer (RMSNorm, SwiGLU, untied embeddings; learned positions in v1) with the ZBP physical/digital partition - `src/zbp_scaling/data.py` — FineWeb-Edu → GPT-2-BPE uint16 shards (WikiText-103 fallback for smoke) - `src/zbp_scaling/trainer.py` — DDP (torchrun, 4 or 8 GPUs), bf16 autocast with fp32 measurement accumulation, gradient accumulation to a fixed global batch, cosine + warmup, resume, JSONL logs - `src/zbp_scaling/diagnostics.py` — per-block branch-gain rho_k profile and the NSR constant c(scale) - `configs/` — model sizes (60m/124m/350m/1b) x training arms (bp / zbp_n16 / zbp_n4) ## Run (one click, on the H200 node) ``` ./scripts/run_ladder.sh # env check+install -> data prep -> all sizes x arms (resume-safe) VENV=1 ./scripts/run_ladder.sh # fresh node: bootstrap ./.venv first SIZES="s60" ARMS="bp zbp_n16" TOKENS=3e8 ./scripts/run_ladder.sh # any subset / budget DRY=1 ./scripts/run_ladder.sh # print the plan ``` Data prep runs **on the training node** (auto-invoked if `data/` is empty); `data/` and `runs/` are gitignored. ``` python scripts/prepare_data.py --tokens 3e9 --out data/fineweb # manual form torchrun --nproc_per_node=8 scripts/train.py --model configs/model/m124.yaml --train configs/train/zbp_n16.yaml ``` Global batch is fixed in the train config; per-rank micro-batch and accumulation adapt to world size. ## Measured cost (48 GB Ampere-class, d=512 L=8 seq 1024, shared GPU) BP 312 ms/step; ZBP n=16 **8.7x**, n=64 **29x** (FLOPs-bound: probe batching is already saturated at probe_chunk 8; the score core uses its own chunk <= 4 since its memory goes as chunk*B*H*T^2). Ladder arms: **bp / zbp_n16 / zbp_n64** per size (n=4 optional). Headroom if needed: torch.compile on the query path; forward differences (n+1 instead of 2n queries) as a cheaper biased arm. ## Citation Paper in preparation ("Backpropagation Without Jacobians"); see the main zobp research repo for the method, theory and part-1/part-3 results. License: Apache-2.0. ## After the runs (results flow) Checkpoints and logs stay on the node (`runs/` is gitignored). Bundle the small JSONL results: ``` python scripts/collect.py --runs runs --out results/h200node1 ``` then send `results/h200node1/` (and, when asked, specific `runs//ckpt.pt` checkpoints) back over any manual channel — scp / rsync / cloud drive. No git or HF credentials are needed on the node. Analysis (anywhere): `python scripts/plot_ladder.py --results results/h200node1` -> per-run table, loss curves, and the gap-vs-scale figure (the paper's part-2 headline). ## HF upload & security (shared nodes) Results (and optionally checkpoints) can go to a **private** HF repo: `HF_UPLOAD=1 [HF_CKPT=1] ./scripts/run_ladder.sh` or manually `python scripts/upload_hf.py --results results/ [--with-ckpt runs]` (default repo `/zbp-scaling-runs`, created private if missing). Uploads authenticate ONLY via the `HF_TOKEN` environment variable or a standard `hf auth login`; tokens are never CLI arguments (argv is world-readable via /proc on shared machines), never written by our scripts, and `.gitignore` excludes token-like files. On a shared node, mint a **fine-grained HF token scoped to the single private repo** (write permission only), `export HF_TOKEN=...` per session, and revoke it after the campaign. The default flow needs no tokens at all: results come back manually (section above) and any HF upload happens from a trusted machine.