diff options
| author | Yuren Hao <yurenh2@illinois.edu> | 2026-07-12 08:33:55 -0500 |
|---|---|---|
| committer | Yuren Hao <yurenh2@illinois.edu> | 2026-07-12 08:33:55 -0500 |
| commit | 2484a7ef7dbf0e0996424a5797aea9002a7c3a52 (patch) | |
| tree | 44ccf3cafb85e984dfbd594e670aef6f1b87e16f /docs/outreach | |
| parent | 68325e9fc389498b6e1c935c34b70fe37adb3b8e (diff) | |
Hardware outreach v2: gate lifted, clockless-MVP story, Dillavou wave-1, brief rewritten to cascade-era
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
Diffstat (limited to 'docs/outreach')
| -rw-r--r-- | docs/outreach/OUTREACH_TARGETS.md | 143 |
1 files changed, 140 insertions, 3 deletions
diff --git a/docs/outreach/OUTREACH_TARGETS.md b/docs/outreach/OUTREACH_TARGETS.md index ffbd177..e4cfdad 100644 --- a/docs/outreach/OUTREACH_TARGETS.md +++ b/docs/outreach/OUTREACH_TARGETS.md @@ -1,4 +1,140 @@ -# EP analog-hardware collaboration — outreach targets (2026-06-21) +# EP analog-hardware collaboration — outreach targets (2026-06-21; REVISED 2026-07-12) + +## ⚡ 2026-07-12 REVISION — gate LIFTED, story changed by the clockless MVP +**User instruction 2026-07-12: outreach begins now.** Gate artifacts in hand: 42.75M full-epoch +"能看" demo (first BP-free transformer LM, gap 0.05 nats), E-tier tolerance ledger, cost model, +and **CLOCKLESS_ANALOG_MVP_PLAN.md** (user-authored) which replaces the $5–20k CIM Demo-0 with a +**$170–300 clockless twin-network tile** (Dillavou-lineage; EP-current vs CL-voltage on one board; +exact vs sign local update; no processor/converter/clock in the learning loop). + +**What the MVP changes about outreach:** +1. The ask shrinks from "help us engineer a CIM demonstrator" to "host/advise a $300, six-week, + scope-and-DMM bench build" — any analog lab can say yes. +2. The scientific lineage points at the **physical-learning community (Penn/Dillavou)**, not only + CIM-VLSI. Dillavou becomes a wave-1 target (his PNAS 2024 board is the design's ancestor; our + deltas: true EP current nudge, exact-vs-sign matrix, the LM program + β-SNR law transfer). + Affiliation note (checked 2026-07-12): LinkedIn = "Independent Researcher, ARIA R&D Creator"; + Penn pages still list postdoc (Durian/Liu). Email the Penn address; keep title-neutral wording. +3. Substrate groups (Shanbhag CIM, Zhu FeFET, THU, Stanford) are **rung-5 partners** (CIM block) — + still first-mover whitespace, pitched as the rung AFTER the tile, which makes us look staged + rather than speculative. Hanumolu's converter relevance drops (MVP deletes converters) → wave-2. +4. Shared quantitative hook everywhere: the board's Factor-4 (bias-vs-nudge-magnitude) = the wall-1 + β-SNR law we measured in fp32 + the E-tier error-channel result (10% relative noise free) — + "the same law, measured in simulation and in physics." + +**Revised sequencing:** wave-1a **Dillavou** (design review + natural collaborator; fastest +credible yes/no) → wave-1b **Shanbhag trio** (local bench + rung-5 CIM; E-tier speaks compute-SNR) +→ wave-2 Zhu (nonvolatile-weight rung: film cap → FeFET conductance), Hanumolu (rung-5 mixed-signal +glue), Mingu Kang, Stanford → unicorns (Grollier/Querlioz) once the 32-edge board exists. +**Ben/Rain thread stays separate** — the MVP plan flows there after the current Overleaf beat. + +**Attachments per send:** COLLABORATOR_BRIEF (rev. 2026-07-12, rewritten to cascade-era) + +CLOCKLESS_ANALOG_MVP_PLAN.md (Dillavou/Shanbhag) — render to PDF and VISUALLY VERIFY before send. +Sender-title TODO still open. Current drafts: §"Email drafts v2" below; the 2026-06-21 drafts at +the bottom are SUPERSEDED (looped-era framing, CIM-first ask). + +--- + +## Email drafts v2 (2026-07-12) — copy-paste after title/attachment check + +### Draft B (wave-1a) — Sam Dillavou · To: dillavou@sas.upenn.edu +Subject: A true-EP current nudge on a twin-network learning circuit — building on your PNAS design + +Hi Dr. Dillavou, + +I'm Yuren Hao (UIUC). Two results may interest you. On the algorithm side, we recently trained a +standard 12-layer transformer language model entirely without backpropagation — equilibrium +propagation on a layered energy, all updates local — to within 0.05 nats of a tuned backprop +control over a full epoch; it generates coherent text. On the hardware side, we are starting a +physical-learning build whose design descends directly from your processor-free network: two +continuously-running replicas, shared weight capacitors, local contrast updates, no clock or +processor in the learning loop. + +The planned departures from your architecture are the reason I'm writing. First, an OTA current +nudge alongside the voltage clamp, so EP's force nudge and Coupled Learning's constraint can be +compared on the same board — the distinction McGinnis, Li and Mori recently formalized. Second, one +exact difference-of-squares contrast channel running in parallel with sign-only update cells, for a +continuous exact-versus-sign comparison. Third, a measured bias-versus-nudge-magnitude curve: in +simulation we find the EP error channel tolerates 10% relative noise, while an additive precision +floor sets a hard threshold on the nudge amplitude — the board should exhibit the same law in +physics, and your imperfection-characterization paper is the closest existing treatment. + +Would you have 20–30 minutes to talk? We would value your judgment on the design before we commit +the board, and there may be a natural collaboration — we bring the transformer/LM program and the +simulation tolerance data; the physical-learning lineage is yours. A one-page brief and the build +plan are attached. + +Best, Yuren + +### Draft A (wave-1b) — Shanbhag group · To: Soonha Hwang (soonhah2@), Mihir Kavishwar (mihirvk2@) · cc: Shanbhag +Subject: Backprop-free transformer training — GPU-scale results and a staged path to CIM + +Hi Soonha and Mihir, + +I'm Yuren Hao, working on backprop-free training in ChengXiang Zhai's group at UIUC. The project +recently crossed a threshold worth reporting: we trained a standard 12-layer transformer language +model with equilibrium propagation — no backpropagation anywhere in training, every update local — +to within 0.05 nats of a tuned backprop control over a full epoch, and it generates coherent text. +Inference is an ordinary forward pass. Every operation in the recipe was chosen to have a known +analog implementation, and we have measured the fault tolerances the learning rule actually needs: +8-bit effective weights are lossless and 6-bit marginal; the error channel tolerates 10% relative +noise; 1% forward state noise costs nothing. + +We are deliberately starting the hardware small: a ~$300 clockless twin-network tile (descended +from the Penn processor-free learning circuits) that validates the physical learning rule with no +processor, converter, or clock in the loop. The reason to write to your group is the rung after +that: a CIM transformer block with in-situ EP updates — analog MVM plus a local two-phase weight +update. Your DiT accelerator and the compute-SNR ADC line are the closest existing substrate for +that rung, and the tolerance table above is, in effect, its SNR budget. + +Could I grab 20 minutes to show the results and the staged plan? A one-page brief is attached. +(cc'ing Prof. Shanbhag.) + +Thanks, Yuren + +### Draft C (wave-2) — Wenjuan Zhu · To: wjzhu@illinois.edu +Subject: Nonvolatile analog weights for a physical equilibrium-propagation learner — FeFET fit? + +Dear Prof. Zhu, + +I'm Yuren Hao, working on backprop-free training in ChengXiang Zhai's group at UIUC. We train +transformers with equilibrium propagation — no backpropagation; each weight updates from a local +contrast between two settled states — and recently demonstrated this at language-model scale in +simulation (a 12-layer model within 0.05 nats of its backprop control). We are now building a small +clockless analog learning network in which each weight is a capacitor charged by its own local +update circuit. + +The capacitor is the honest weakness: it is volatile. The natural upgrade is exactly your group's +territory — a nonvolatile, electrically-programmable, multilevel conductance, and your vdW / +CuInP2S6 FeFETs are the closest devices I know of. I realize that work has centered on memory and +logic rather than training; the question is whether a FeFET conductance could replace the weight +capacitor in a continuously-learning analog network, with the update current driving the gate. + +Would you have 20 minutes to discuss feasibility? A one-page brief and the build plan are attached. + +Best, Yuren + +### Draft D (wave-2) — Hanumolu · To: hanumolu@illinois.edu +Subject: Mixed-signal partner for the CIM phase of an analog learning program — student pointer? + +Dear Prof. Hanumolu, + +I'm Yuren Hao, working on backprop-free training in ChengXiang Zhai's group at UIUC. We train +transformers with equilibrium propagation (no backpropagation; local two-phase updates), recently +at language-model scale in simulation, and are starting the hardware side with a deliberately +minimal clockless analog tile — no converters at all in the learning loop. + +The phase where your group's expertise becomes central is the one after: an in-memory-compute +transformer block, where settled-state readout, nudge injection, and loop stability are +mixed-signal problems. Nearer-term, the tile itself has one control-loop question — enforcing a +100–1000× time-scale separation between state settling and weight motion — that a student who +enjoys discrete analog and feedback loops might find fun as a side project. + +Could you point me to a student for either, or spare 15 minutes? One-page brief attached. + +Best, Yuren + +--- Per-group PhD/PI profiles from 5 research agents. Accuracy discipline: emails only where published or netid on an official directory; "—" = not public, route via PI (no invented addresses). Verify "current" status before sending — students graduate. Companion: COLLABORATOR_BRIEF.md (the one-pager), HW_RESEARCH_FINDINGS.md (citations). @@ -133,7 +269,8 @@ world — EP-rich, mostly device-light. Pair one of each. --- -## ⏸ STATUS (2026-06-21): HOLD — DO NOT SEND until the 33M demo + scaling dossier +## ~~⏸ STATUS (2026-06-21): HOLD~~ → **GATE LIFTED 2026-07-12 (user instruction; artifacts delivered). Use "Email drafts v2" above; everything below is the superseded 06-21 record.** +## (superseded) ⏸ STATUS (2026-06-21): HOLD — DO NOT SEND until the 33M demo + scaling dossier **User decision (CONFIRMED 2026-06-21): outreach is gated on the ~33M "能看" demo + scaling-law dossier (task #15) — NOT the C512/2.09 milestone.** Send nothing until there's a readable-generation ("能看") demo + a scaling-law dossier to lead with. (C512 EP descending past the 2.09 wall toward ~1.8 is a prerequisite step that validates the recipe, NOT the outreach gate — @@ -141,7 +278,7 @@ the gate is the bigger, showable 33M artifact.) Until then: no contact with anyo When the bar is met: set sender title, render COLLABORATOR_BRIEF.pdf, attach + ept_method_intro.pdf, optionally ask Prof. Zhai for a warm intro to Shanbhag/Hanumolu first. All profiles/contacts/pairing/drafts above are durable and ready. -## Email drafts (READY, gated — copy-paste when the bar is met) +## Email drafts v1 (2026-06-21) — SUPERSEDED by v2 above (looped-era framing, CIM-first ask; kept for the record) ### Draft 1 — Shanbhag group · To: Soonha Hwang (soonhah2@), Mihir Kavishwar (mihirvk2@) · cc: Shanbhag Subject: Backprop-free (Equilibrium-Propagation) transformer training — a fit for your in-memory CIM work? |
