diff options
| -rw-r--r-- | README.md | 25 | ||||
| -rwxr-xr-x | scripts/run_ladder.sh | 15 |
2 files changed, 7 insertions, 33 deletions
@@ -45,27 +45,14 @@ Paper in preparation ("Backpropagation Without Jacobians"); see the main zobp re method, theory and part-1/part-3 results. License: Apache-2.0. ## After the runs (results flow) -Checkpoints stay on the node (`runs/` is gitignored). Collect the small JSONL logs and push them back: +Checkpoints and logs stay on the node (`runs/` is gitignored). Bundle the small JSONL results: ``` python scripts/collect.py --runs runs --out results/h200node1 -git add results && git commit -m "ladder results" && git push ``` +then send `results/h200node1/` (and, when asked, specific `runs/<name>/ckpt.pt` checkpoints) back over any +manual channel — scp / rsync / cloud drive. No git or HF credentials are needed on the node. Analysis (anywhere): `python scripts/plot_ladder.py --results results/h200node1` -> per-run table, loss -curves, and the gap-vs-scale figure (the paper's part-2 headline). If a specific checkpoint is needed for -the estimator audits, scp just that `runs/<name>/ckpt.pt`. - -## Collaborator quickstart (zero tokens on the node) -You receive ONE file: the deploy key `zbp_scaling_deploy` (scoped to this repo only, revocable). Then: -``` -install -m 600 zbp_scaling_deploy ~/.ssh/zbp_scaling_deploy -git clone -c core.sshCommand="ssh -i ~/.ssh/zbp_scaling_deploy -o IdentitiesOnly=yes" \ - git@github.com:YurenHao0426/zbp-scaling.git -cd zbp-scaling && ./scripts/run_ladder.sh # env check -> data prep -> ladder -> results auto-pushed back -``` -The `-c` persists `core.sshCommand` inside the clone, so the auto-push at the end works with no env setup -(nohup-safe; set `PUSH_RESULTS=0` to disable). No GitHub account, no HF token on the node: results JSONL -flow back through the deploy key; checkpoints stay on the node (scp on request) and HF uploads happen on -the maintainer's machine. +curves, and the gap-vs-scale figure (the paper's part-2 headline). ## HF upload & security (shared nodes) Results (and optionally checkpoints) can go to a **private** HF repo: `HF_UPLOAD=1 [HF_CKPT=1] ./scripts/run_ladder.sh` @@ -76,5 +63,5 @@ Uploads authenticate ONLY via the `HF_TOKEN` environment variable or a standard never CLI arguments (argv is world-readable via /proc on shared machines), never written by our scripts, and `.gitignore` excludes token-like files. On a shared node, mint a **fine-grained HF token scoped to the single private repo** (write permission only), `export HF_TOKEN=...` per session, and revoke it after the campaign. -Zero-token alternative: push only the small JSONL results to GitHub (a repo-scoped deploy key suffices) and -upload checkpoints from a trusted machine. +The default flow needs no tokens at all: results come back manually (section above) and any HF upload +happens from a trusted machine. diff --git a/scripts/run_ladder.sh b/scripts/run_ladder.sh index 7402111..1d412a0 100755 --- a/scripts/run_ladder.sh +++ b/scripts/run_ladder.sh @@ -12,8 +12,6 @@ # MICRO_BS=8 per-GPU micro batch [8] # SET="k=v k2=v2" extra --set overrides for every run (e.g. vocab=8192 seq_len=256) # DRY=1 print the plan and exit -# PUSH_RESULTS=0 skip the auto collect + git-push of results JSONL [on] -# HF_UPLOAD=1 [HF_CKPT=1] also upload to a private HF repo (needs HF_TOKEN; see README) # # Runs are sequential (each takes the whole node), resume-safe: a finished run leaves OUT/<name>/DONE # and is skipped on re-invocation, so the script can be re-run after interruptions. @@ -68,19 +66,8 @@ for size in $SIZES; do touch "$dir/DONE" done done -TAG=${HF_TAG:-$(hostname)-$(date +%Y%m%d)} -if [ "${PUSH_RESULTS:-1}" = 1 ] && [ "${DRY:-0}" != 1 ]; then - # zero-token default: collect the small JSONL results and push them back over this clone's git auth - if python scripts/collect.py --runs "$OUT" --out "results/$TAG"; then - git add results - git -c user.email=ladder@zbp -c user.name=ladder commit -m "results: $TAG" || true # nothing new is fine - git push || echo "!! git push failed — push manually later or send results/$TAG" - else - echo "!! collect found no finished runs — skipping push" - fi -fi if [ "${HF_UPLOAD:-0}" = 1 ]; then - # optional direct-to-HF (needs HF_TOKEN in the environment; see README Security) + TAG=${HF_TAG:-$(hostname)-$(date +%Y%m%d)} python scripts/collect.py --runs "$OUT" --out "results/$TAG" python scripts/upload_hf.py --results "results/$TAG" ${HF_REPO:+--repo "$HF_REPO"} ${HF_CKPT:+--with-ckpt "$OUT"} fi |
