summaryrefslogtreecommitdiff
path: root/docs/hardware/TINYEP_HOSTILE_AUDIT.md
blob: 96f896c0ae2fae3a7af08b98b97695674ebedb9e (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
# HOSTILE AUDIT — TinyEP-256-R32 (TINYSTORIES_EQPROP_ANALOG_TRAINER_MVP.md), 2026-07-13

Reviewer stance: mixed-signal tape-out + bench + referee. Baselines: MSCALE_SUPPLY_CHAIN_SEARCH.md (M-study), HW_FIRST_PRINCIPLES_SEARCH.md ([D]/[X]/[H] audits).

## (a) PLACEMENT — evades the residency theorem, walks into the retention wall

- **Residency theorem: genuinely evaded.** The theorem's antecedent (model > physical array → sub-block residency → digitized inter-block activations → B/B− ceiling) does not hold: all 37k weights are resident, R=1, no streaming, no tile edge to digitize. Reuse-forced volatility is also evaded — cells absorb no reload writes, only analog update currents (no endurance arithmetic).
- **But the theorem has a corollary the proposal never confronts: full residency on volatile cells trades digitization for droop.** 2N7002-class access FETs leak 1–100 nA typ at datasheet spec (~10 pA–10 nA derated at ~1 V cell bias); on 10–100 nF that is 1 LSB (8 mV @ 8-bit, ±1 V span) per **~0.1–100 s per cell**, consistent with [D]'s measured CD4066 line (1 LSB in 3–6 min on ≥4.7 µF ⇒ 4–8 s scaled to 100 nF). Effective per-weight decay τ ≈ **minutes–hours**, with 10–100× per-cell spread. Training updates shrink toward zero near convergence; droop doesn't. The machine is the M-study's "refresh machine" with training as the refresh — a droop-limited loss floor, the LF398 arithmetic transposed from S/H to the weight bank itself. "Leak = weight decay" is a euphemism: mismatched per-cell λ is a random decay field, not a regularizer.
- **≤$3k flagship-enabler? No.** It escapes the M-band's walls by shrinking the object, but it does not land in the band (see (c): honest parts $5–12k), and the object it shrinks to is no longer the thing the band's flagship sentence needs (see (d)/(e)). The M-study's ruling — ≤$3k buys metrology, not a trainer — stands.

## (b) PHYSICS — the energy is real; the scanner spends EP's one advantage

- **Conservative construction: PASS on paper, and it is the proposal's best idea.** Tied-K/V softmax attention over clamped context embeddings is exactly the modern-Hopfield/DenseAM energy (∂E/∂z of −logsumexp = softmax-weighted sum of the same vectors); the rank-32 factors as physical bilinear layers, the 512-unit rectified memory, tied embed/decode, and leak give a legitimate layered energy. Classic EP legal, no transpose circuitry: same-cell forward/transpose access is the analog analogue of T64's word-transpose — **mode (i) exact by construction**, the strongest reciprocity evidence class in either prior study. The p−y nudge current is standard and clean.
- **Simultaneous settle: broken and replaced, semi-defensibly.** 16–32 shared lanes over 37k weights means the state never relaxes under the full physical energy; it is Gauss-Seidel coordinate descent by scanned analog MACs. For a contractive energy (their leak gate) async relaxation reaches the same fixed point (Bertsekas-Tsitsiklis conditions), and EP only needs the fixed points — so the *gradient theorem* survives, contingent on gate 4. "Both replicas settled before update integration" is enforceable (two residual envelopes AND-ed into a C-element; ± replicas in matched parallel lanes). But EP's selling point — physics does the relaxation in constant time — is forfeited; the machine is a self-timed serial processor whose ALU is analog.
- **Wall-clock: the kill.** Per-sweep MACs ≈ 131k (attention: 8×256×32 scores + value sums) + ~70k (bilinear banks); context-embedding instantiation is 256×8,192 ≈ **2.1M MACs/example** — which requires either an uncosted ~65k-cell analog hold bank (1.8× the weight array) or recomputation every sweep (5–10× time). Per example: 2 phases × 15–30 sweeps × 200k MACs ÷ 32 lanes × 1–5 µs/slot (LM13700/AD633 GBW floor; T64's own kill line was 7 µs) ≈ **0.1–0.5 s/token floor** (4-way batch broadcast included), 1–2 s realistic. 100M char-tokens = **115–580 days; 500M = 1.6–8 years. DEAD** by the 30-day rule. The only survivable envelope: ≤10–15M tokens at ≤0.25 s/token = 12–30 days, MARGINAL. (T64 does 100M tokens in 9–14 d with 15× the model — digital logistics pipeline the settles; all-analog scanning is slower, not faster.) Add read-disturb: ~80 cell accesses/example × charge injection 1–30 pC ⇒ 0.01–0.3 mV/event on 100 nF; even dummy-switch-cancelled ×10, an accumulating per-example bias competing with a 0.1–1 mV contrast signal [D].

## (c) COSTS — line-by-line correction

| Line | Proposed | Corrected (parts) | Basis |
|---|---:|---:|---|
| 37k cells + cards | $650–1,250 | **$1.5–6k** | SC-1M audited c_cell $0.033–0.05 all-in ⇒ $1.2–1.9k floor at cheap caps; but the proposal's own leakage/DA gate forces low-DA dielectric: 100 nF C0G exists only in 1206-class (TDK C3216NP0/Yageo) at ~$0.2–0.5 ⇒ $6–19k in caps alone, vs 10 nF C0G cheap but 10× droop/injection. The fork is unpriced. |
| 16–32 dual-sign MAC/update lanes | $180–420 | **$1–2.5k + trim labor** | 64 matched four-quadrant multipliers: AD633JRZ $10.97 audited = $700 alone; LM13700M $1.52 route reopens the [D] offset war (±5–50 mV vs 0.1–1 mV contrast) ⇒ per-lane auto-zero front-ends, the [D] unpriced critical path ×37k cells through 32 lanes. SC-1M's functional equivalent (analog product engine) audited at **$8–15k**. |
| 256-ch softmax bank | $80–220 | $300–800 + cal | V_T mismatch 1 mV = 4% exp error; 256-way matching = trim/selection labor, not parts. |
| State banks/integrators/detectors | $180–400 | $800–2k **+ $2–5k embedding hold bank (or 5–10× wall-clock)** | 8 replicas × ~800 nodes; the 65k-cell context-memory store is absent from the BOM entirely. |
| Row/col RMS optimizer + controller | $100–250 | $500–1.2k | ~1,200 envelope-detector + variable-gm cells. |
| Power/backplane/spares | $200–450 | $300–700 + one respin $200–600 [D] | 4–8 large guarded cards. |
| Missing lines | — | async address chains $300–800; corpus reader $150–400; **null/cal bench labor (the [D] finding: the money is the null, not the BOM)** | |

**Corrected: parts $5–12k; loaded $25–50k** (M-study accounting). The proposal's $1.4–3.0k is a 3–5× parts and ~15× loaded understatement; the missing money is exactly where SC-1M's budget went — the product engine, the hold/readout periphery, and calibration. Sanity anchor: SC-1M, 49k cells, $40–78k all-in.

## (d) QUALITY RISK — 6× below the published coherence floor, on harder tokens

Eldan-Li's demonstrated range is **1M–35M params on a reduced-vocab word-level tokenizer**; coherent multi-paragraph stories start at ~1M. TinyEP is 164k *effective* (37k physical), char-level — below the paper's floor by ~6× on capacity and on a strictly harder tokenization, before the EP-estimator tax (cos(EP,BPTT) ≈ 0.9 ceiling in this program's own software) and the analog fault stack. Honest forecast: held-out bits/char plausibly beats a 5-gram (n-grams plateau ~1.9–2.2 b/c on TinyStories-simple text; a 164k char model can reach ~1.7–1.9), so **gate 2 as literally written is passable**; completions will show correct spelling and local syntax and zero story coherence. "Readable completions" honestly = readable *words*, not stories. To reach the paper's coherence floor: ~1–3M effective ⇒ rank 128+ and wider memory ⇒ 150–300k physical cells ⇒ the M-band machine ($40–80k) this proposal was built to escape. Rank 64 (328k eff) does not get there either.

## (e) CLAIM — mechanism exceeds the bar, object fails it

- Per the M-study ruling, the minimum object licensing "an LM trained in analog hardware" is D2/B−: both phases physically settled, edge digitization, digitally composed updates — *on a real LM* (T64's pre-registered sentence names a 2.5M transformer). TinyEP is **stronger than D2 on mechanism** (no ADC/DAC/processor anywhere: updates, optimizer state, control all analog — an A-shaped loop) and **weaker on object** (164k-effective char model, below the field's own coherence floor).
- The M-study's referee kill ("digital training with an analog gradient estimator") does **not** fire — nothing is digital. It is replaced by two new kills: (1) **"not really an LM"** — the referee samples text, reads word salad, and the headline noun dies; (2) **"your self-timed one-hot scanner over 37k weights is a serial processor with an analog ALU — the physics performs no parallel relaxation; what does EP buy?"** Kill (1) is near-certain at flagship if the claim sentence in §9 is kept verbatim.
- Licensed sentence (defensible, NeurIPS/Nat-Comms hardware track): *"End-to-end in-situ training of an energy-based attention **sequence model** by Equilibrium Propagation in a self-timed switched-analog circuit, with weights, nudges, updates, and optimizer statistics as physical charges; X bits/char vs n-gram baseline, N=37k physical weights."* — "language model" demoted from headline to caveat. Also mandatory citation exposure: Kerjan-Høier-Scellier (2606.03584) already holds EP-transformer territory in software; Yi 2022/Momeni 2023 hold hardware-training precedent; the only unclaimed cell is "attention + fully-analog loop," and that is what the sentence must sit on.

## (f) VERDICT

- **Corrected cost:** parts $5–12k, loaded $25–50k (claimed $1.4–3.0k: rejected).
- **Corrected wall-clock:** 0.1–0.5 s/token floor ⇒ 100M tokens 115–580 d **DEAD**; feasible envelope ≤10–15M char-tokens in 12–30 d, which further degrades (d).
- **Grade:** attempted A(37k); expected **B(37k)-contested** (scanner ruling + droop floor); C if the LM noun is retained against word-salad samples.
- **Survival per its own claim class** ("end-to-end analog-EP TinyStories LM, $1.4–3k"): **DEAD** — cost ×4–15, wall-clock ×10–40, quality 6× under the published floor. **SURVIVES-DEGRADED** as re-scoped: 10M-char budget, mechanism-first sequence-model claim, corrected budget class.
- **Ladder position:** does not open the ≤$3k band (that band's correct spend remains the M-study's metrology package). Slots *between* the $150–260 grade-A tile (still the correct first physical spend) and T64 ($52–82k, B−, real 2.5M LM) — at roughly SC-1M's price point with a purer mechanism and a weaker object. Its genuinely novel, program-portable assets: the tied-K/V DenseAM energy (conservative attention, no J^T — relevant to the ept software line), and same-cell bidirectional access as analog mode-(i)-by-construction.
- **Load-bearing gates (of its own 1–7):** **Gate 5** (fault twin — but extended: per-cell random decay τ=0.5–10 h with 10× spread, read charge-injection accumulation, lane offsets vs 0.1–1 mV contrast; this gate decides existence), **Gate 4** (async coordinate-order stability — decides both EP-legality and the sweep count that sets wall-clock), **Gate 2 run at the feasible token budget** (≤10–15M chars, not unlimited — otherwise it gates nothing). Add a missing **Gate 0**: pre-registered wall-clock ledger (slots × slot-time × sweeps × tokens ≤ 30 d) plus a costed design for the 65k-cell context-embedding store. All gates are $0 GPU work and worth running before any board spend; the board is not.

Sources: 2N7002 leakage — Nexperia/Infineon/onsemi datasheets; 100nF C0G 1206 — TDK C3216NP01H104J, Yageo CC1206JKNPO9BN104 (DigiKey); TinyStories arXiv:2305.07759.

---
*Hostile audit by a single English-only agent against the two prior workflow studies, 2026-07-13.*