<feed xmlns='http://www.w3.org/2005/Atom'>
<title>ept.git/ep_run/lt_ep_train.py, branch xiang</title>
<subtitle>Unnamed repository; edit this file 'description' to name the repository.
</subtitle>
<link rel='alternate' type='text/html' href='https://git.blackhao.com/ept.git/'/>
<entry>
<title>lt_ep_train: --gov_keepjr now keeps the jr floor through reg_delay too (fastfull arm: jr always-on, resreg post-warmup)</title>
<updated>2026-07-09T08:43:54+00:00</updated>
<author>
<name>Yuren Hao</name>
<email>yurenh2@illinois.edu</email>
</author>
<published>2026-07-09T08:43:54+00:00</published>
<link rel='alternate' type='text/html' href='https://git.blackhao.com/ept.git/commit/?id=a015da4a613f71f23dcfe2168c7d783faa0a6c55'/>
<id>a015da4a613f71f23dcfe2168c7d783faa0a6c55</id>
<content type='text'>
Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</pre>
</div>
</content>
</entry>
<entry>
<title>lt_ep_train: --rr_floor (resreg floor during free/delay phase) for govfloor arms</title>
<updated>2026-07-09T08:32:17+00:00</updated>
<author>
<name>Yuren Hao</name>
<email>yurenh2@illinois.edu</email>
</author>
<published>2026-07-09T08:32:17+00:00</published>
<link rel='alternate' type='text/html' href='https://git.blackhao.com/ept.git/commit/?id=e6018f3cbe07d993f9f21a282cb9b7c22ef1544b'/>
<id>e6018f3cbe07d993f9f21a282cb9b7c22ef1544b</id>
<content type='text'>
Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</pre>
</div>
</content>
</entry>
<entry>
<title>lt_ep_train: optional W&amp;B mirroring (--wandb/--wandb_run), failure-proof logging of val/ema/best/res/jr/lr/rho/gov state</title>
<updated>2026-07-09T08:08:42+00:00</updated>
<author>
<name>Yuren Hao</name>
<email>yurenh2@illinois.edu</email>
</author>
<published>2026-07-09T08:08:42+00:00</published>
<link rel='alternate' type='text/html' href='https://git.blackhao.com/ept.git/commit/?id=183863661c35336001bf0226af3900cefe6fd9aa'/>
<id>183863661c35336001bf0226af3900cefe6fd9aa</id>
<content type='text'>
Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</pre>
</div>
</content>
</entry>
<entry>
<title>SPECTRAL GOVERNOR (--governor): the from-scratch recipe candidate, dip-map-calibrated</title>
<updated>2026-07-07T10:48:58+00:00</updated>
<author>
<name>Yuren Hao</name>
<email>yurenh2@illinois.edu</email>
</author>
<published>2026-07-07T10:48:58+00:00</published>
<link rel='alternate' type='text/html' href='https://git.blackhao.com/ept.git/commit/?id=ed367116af71e4688cb3154fa85b20cdca564b32'/>
<id>ed367116af71e4688cb3154fa85b20cdca564b32</id>
<content type='text'>
Dip screening (4/4 seeds have dips; first-dip cluster 1200-1700, width
100-300 steps, depth |lam|~0.998 = s2000-class; s2 shows excursion-&gt;dip in
sequence): reg-free until a scout (warm lead_rho, every gov_k=100, deep-400
subbatch) flags rho&lt;0.985 after gov_min=1000, then ARPACK-certify all top-3
|lam|&lt;1 -&gt; leash ON (pair regs). Post-engagement: sustained rho_scout&gt;1.02
x2 -&gt; rescue boost (resreg x3 for 500 steps; the proven hr2 maneuver).
abl_delay's fixed-2000 death explained: it engaged AFTER the dip cluster,
mid-excursion. Two governor seeds launched on 107 (trained-state engagement
is Pascal-safe per hr2 precedent).

Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Dip screening (4/4 seeds have dips; first-dip cluster 1200-1700, width
100-300 steps, depth |lam|~0.998 = s2000-class; s2 shows excursion-&gt;dip in
sequence): reg-free until a scout (warm lead_rho, every gov_k=100, deep-400
subbatch) flags rho&lt;0.985 after gov_min=1000, then ARPACK-certify all top-3
|lam|&lt;1 -&gt; leash ON (pair regs). Post-engagement: sustained rho_scout&gt;1.02
x2 -&gt; rescue boost (resreg x3 for 500 steps; the proven hr2 maneuver).
abl_delay's fixed-2000 death explained: it engaged AFTER the dip cluster,
mid-excursion. Two governor seeds launched on 107 (trained-state engagement
is Pascal-safe per hr2 precedent).

Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</pre>
</div>
</content>
</entry>
<entry>
<title>Tier-3 gates: Anderson math-yes/impl-no (parked for v2); bf16polish UNSAFE near-edge (eval-only); dp_ep.py ready</title>
<updated>2026-07-06T14:42:52+00:00</updated>
<author>
<name>Yuren Hao</name>
<email>yurenh2@illinois.edu</email>
</author>
<published>2026-07-06T14:42:52+00:00</published>
<link rel='alternate' type='text/html' href='https://git.blackhao.com/ept.git/commit/?id=40be67d4f5b5a6b46c662c70b759e585e136d70e'/>
<id>40be67d4f5b5a6b46c662c70b759e585e136d70e</id>
<content type='text'>
Anderson: res 25-35x deeper per budget but 5.8x slower (naive history stacks
+ per-iter safeguard eval) — v2 = ring buffers + periodic safeguard, est +1.3x
on the speed tier. bf16+20polish: 1.41x free phase, res parity, BUT z-diff
1.2e-3 — near-marginal operators contract too slowly for a 20-step polish
(0.998^20≈0.96), same magnitude as the TF32 kill verdict and the estimator's
50%-sensitivity input. Predicted by our own depth/noise theory. Flags kept
with warnings; neither ships for training. dp_ep.py: manual-allreduce EP DP
(controller in lockstep, aligned collectives), smoke pending freed 1080s.

Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Anderson: res 25-35x deeper per budget but 5.8x slower (naive history stacks
+ per-iter safeguard eval) — v2 = ring buffers + periodic safeguard, est +1.3x
on the speed tier. bf16+20polish: 1.41x free phase, res parity, BUT z-diff
1.2e-3 — near-marginal operators contract too slowly for a 20-step polish
(0.998^20≈0.96), same magnitude as the TF32 kill verdict and the estimator's
50%-sensitivity input. Predicted by our own depth/noise theory. Flags kept
with warnings; neither ships for training. dp_ep.py: manual-allreduce EP DP
(controller in lockstep, aligned collectives), smoke pending freed 1080s.

Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</pre>
</div>
</content>
</entry>
<entry>
<title>aggregate speed bench: speed tier hf+sd 1.47x; accuracy tier hf+sd+t80+avg (0.91x, cos 0.94, avg free); compile demoted</title>
<updated>2026-07-06T14:17:36+00:00</updated>
<author>
<name>Yuren Hao</name>
<email>yurenh2@illinois.edu</email>
</author>
<published>2026-07-06T14:17:36+00:00</published>
<link rel='alternate' type='text/html' href='https://git.blackhao.com/ept.git/commit/?id=9a8b2796ca12e4d4c24717485a635a301aa6d07f'/>
<id>9a8b2796ca12e4d4c24717485a635a301aa6d07f</id>
<content type='text'>
Full-ep_step wall times on quiet A6000 (warm s2000, B24), res parity across
all 8 configs. compile only 1.12x at this shape (historical 1.46x was a
different workload split); FULL cmp_sdpa saves 4% over eager at t80 — not
worth the guard complexity. tforce_sdpa added (flash baked into compiled
graph, flag-free so grad paths never see SDPA). bp_lm --tie probe in flight.

Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Full-ep_step wall times on quiet A6000 (warm s2000, B24), res parity across
all 8 configs. compile only 1.12x at this shape (historical 1.46x was a
different workload split); FULL cmp_sdpa saves 4% over eager at t80 — not
worth the guard complexity. tforce_sdpa added (flash baked into compiled
graph, flag-free so grad paths never see SDPA). bp_lm --tie probe in flight.

Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</pre>
</div>
</content>
</entry>
<entry>
<title>--holoavg shipped (gate 0.913-&gt;0.936 @t2sel160, batch2 +0.051) + farm wave-2 tooling</title>
<updated>2026-07-06T12:15:30+00:00</updated>
<author>
<name>Yuren Hao</name>
<email>yurenh2@illinois.edu</email>
</author>
<published>2026-07-06T12:15:30+00:00</published>
<link rel='alternate' type='text/html' href='https://git.blackhao.com/ept.git/commit/?id=488c50e1bdbf8f420b2ad5b4021a7f950d121835'/>
<id>488c50e1bdbf8f420b2ad5b4021a7f950d121835</id>
<content type='text'>
trend-aware stop + plateau averaging recovers the semi-convergence victims
exactly as predicted. Estimator pack now: holofast+sdpa+t2sel80(+holoavg).
Also: bp_lm stdinit/beta2/sched flags (anchor archaeology), dipfarm_freezer,
staged tol_sweep.sh (gated on hr2 verdict), bp_sweep.sh.

Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
trend-aware stop + plateau averaging recovers the semi-convergence victims
exactly as predicted. Estimator pack now: holofast+sdpa+t2sel80(+holoavg).
Also: bp_lm stdinit/beta2/sched flags (anchor archaeology), dipfarm_freezer,
staged tol_sweep.sh (gated on hr2 verdict), bp_sweep.sh.

Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</pre>
</div>
</content>
</entry>
<entry>
<title>--sdpa: fused flash attention in the no_grad relax loop — 1.45x free phase, z* parity 4e-7</title>
<updated>2026-07-05T10:09:36+00:00</updated>
<author>
<name>Yuren Hao</name>
<email>yurenh2@illinois.edu</email>
</author>
<published>2026-07-05T10:09:36+00:00</published>
<link rel='alternate' type='text/html' href='https://git.blackhao.com/ept.git/commit/?id=75ff326dcb40cd960abd56e9c9c18a45d9e5e2c2'/>
<id>75ff326dcb40cd960abd56e9c9c18a45d9e5e2c2</id>
<content type='text'>
Scoped via blk._sdpa set only inside relax()'s loop (grad paths jvp/vjp/resreg
keep the manual attention: no forward-mode-through-flash risk). Combined with
--holofast: ~1.51x full-step exact-math tier.

Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Scoped via blk._sdpa set only inside relax()'s loop (grad paths jvp/vjp/resreg
keep the manual attention: no forward-mode-through-flash risk). Combined with
--holofast: ~1.51x full-step exact-math tier.

Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</pre>
</div>
</content>
</entry>
<entry>
<title>holofast: exact halved-jvp AEP track — 1.55x nudged phase, gradient-gate passed</title>
<updated>2026-07-05T08:53:57+00:00</updated>
<author>
<name>Yuren Hao</name>
<email>yurenh2@illinois.edu</email>
</author>
<published>2026-07-05T08:53:57+00:00</published>
<link rel='alternate' type='text/html' href='https://git.blackhao.com/ept.git/commit/?id=cbecb171b1af77fe6510fd60f1dfa4c7938a20d6'/>
<id>cbecb171b1af77fe6510fd60f1dfa4c7938a20d6</id>
<content type='text'>
holo_a_track computed the doubled-batch jvp/vjp on [v0; -v0] at a shared
anchor zbar — exact antisymmetric redundancy (phase deviations from the
common mode are exact negatives). holo_a_track_fast computes at batch B and
mirrors: single-eval parity 6e-7 (exact); trajectory-level 45% divergence
SHARED with the original's own FD noise floor (1e-6 state noise -&gt; 49%
self-divergence — the 2r=0.04 finite difference amplifies fp noise; training
averages it via pema/momentum). Ship gate: cos(EP,BPTT) orig vs fast
indistinguishable (0.907/0.912, 0.853/0.853, 0.918/0.918). Timing 6.43-&gt;4.16s
on the T2=40 nudged phase (contended GPU, relative). --holofast flag,
default off; queued ablation arms deliberately stay on orig for fidelity.

Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
holo_a_track computed the doubled-batch jvp/vjp on [v0; -v0] at a shared
anchor zbar — exact antisymmetric redundancy (phase deviations from the
common mode are exact negatives). holo_a_track_fast computes at batch B and
mirrors: single-eval parity 6e-7 (exact); trajectory-level 45% divergence
SHARED with the original's own FD noise floor (1e-6 state noise -&gt; 49%
self-divergence — the 2r=0.04 finite difference amplifies fp noise; training
averages it via pema/momentum). Ship gate: cos(EP,BPTT) orig vs fast
indistinguishable (0.907/0.912, 0.853/0.853, 0.918/0.918). Timing 6.43-&gt;4.16s
on the T2=40 nudged phase (contended GPU, relative). --holofast flag,
default off; queued ablation arms deliberately stay on orig for fidelity.

Co-Authored-By: Claude Fable 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</pre>
</div>
</content>
</entry>
<entry>
<title>magic-s2000 study: reg_delay/noadaptc flags, 4-arm queue v2, redx trajectory audit</title>
<updated>2026-07-05T03:35:16+00:00</updated>
<author>
<name>Yuren Hao</name>
<email>yurenh2@illinois.edu</email>
</author>
<published>2026-07-05T03:35:16+00:00</published>
<link rel='alternate' type='text/html' href='https://git.blackhao.com/ept.git/commit/?id=b11d9c6da6ce32471e1c25a6f1b5e7a0a568774d'/>
<id>b11d9c6da6ce32471e1c25a6f1b5e7a0a568774d</id>
<content type='text'>
- lt_ep_train: --reg_delay N (reg-free early phase: resreg/jr/floss/adaptc off
  for first N steps) + --noadaptc (kill hidden jacreg==0 damping feedback that
  would pollute single-reg ablation arms)
- queue v2: 4 arms delay-first (abl_delay = reg-free 2k -&gt; proven pair)
- eig_traj/2/3: ARPACK audit of redx_traj — the run crossed the edge EARLY and
  oscillated (s1000 rotating-unstable, s1400 excursion mu=+2.1 self-recovered,
  s2000 the ONLY stable snapshot mu=-0.02, s2100/s2200 already back out) =&gt;
  s2000 is a post-excursion STABILITY-DIP capture, dip width &lt;100 steps;
  learning survives mild instability (val fell through unstable stretches).
  lead_rho cold-40 under-reads clusters — NOT a classifier; ARPACK for audits.

Co-Authored-By: Claude Opus 4.8 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
- lt_ep_train: --reg_delay N (reg-free early phase: resreg/jr/floss/adaptc off
  for first N steps) + --noadaptc (kill hidden jacreg==0 damping feedback that
  would pollute single-reg ablation arms)
- queue v2: 4 arms delay-first (abl_delay = reg-free 2k -&gt; proven pair)
- eig_traj/2/3: ARPACK audit of redx_traj — the run crossed the edge EARLY and
  oscillated (s1000 rotating-unstable, s1400 excursion mu=+2.1 self-recovered,
  s2000 the ONLY stable snapshot mu=-0.02, s2100/s2200 already back out) =&gt;
  s2000 is a post-excursion STABILITY-DIP capture, dip width &lt;100 steps;
  learning survives mild instability (val fell through unstable stretches).
  lead_rho cold-40 under-reads clusters — NOT a classifier; ARPACK for audits.

Co-Authored-By: Claude Opus 4.8 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
</pre>
</div>
</content>
</entry>
</feed>
