diff options
| author | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-29 18:57:24 -0500 |
|---|---|---|
| committer | YurenHao0426 <Blackhao0426@gmail.com> | 2026-07-29 18:57:24 -0500 |
| commit | 0c70a6c80f50dc48241a9860b2364affd3a09042 (patch) | |
| tree | 0e097dfd3a781ebd823bc6104b71a0bd52b0bf4f /index.html | |
EPT project page — Academic Project Page Template port (grant-review noindex mode)
Template: github.com/eliahuhorwitz/Academic-project-page-template (CC BY-SA 4.0, attribution retained)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FAPDWQ49M5Ye3NpTndTpn
Diffstat (limited to 'index.html')
| -rw-r--r-- | index.html | 357 |
1 files changed, 357 insertions, 0 deletions
diff --git a/index.html b/index.html new file mode 100644 index 0000000..f5e863b --- /dev/null +++ b/index.html @@ -0,0 +1,357 @@ +<!DOCTYPE html> +<html lang="en"> +<head> + <meta charset="utf-8"> + <meta name="viewport" content="width=device-width, initial-scale=1"> + + <!-- Primary Meta Tags --> + <meta name="title" content="PAPER_TITLE - AUTHOR_NAMES"> + <meta name="description" content="Standard transformer LMs trained end-to-end by Equilibrium Propagation, matching backprop within 4-5% perplexity at 72M parameters, with matched controls and hardware-grade ablations."> + + <meta name="author" content="FIRST_AUTHOR_NAME, SECOND_AUTHOR_NAME"> + <meta name="robots" content="index, follow"> + <meta name="language" content="English"> + + <!-- Open Graph / Facebook --> + <meta property="article:published_time" content="2024-01-01T00:00:00.000Z"> + <meta property="article:author" content="FIRST_AUTHOR_NAME"> + <meta property="article:section" content="Research"> + <meta property="article:tag" content="KEYWORD1"> + <meta property="article:tag" content="KEYWORD2"> + + <!-- Twitter --> + + <!-- Academic/Research Specific --> + + <!-- Additional SEO --> + <meta name="theme-color" content="#2563eb"> + <meta name="msapplication-TileColor" content="#2563eb"> + <meta name="apple-mobile-web-app-capable" content="yes"> + <meta name="apple-mobile-web-app-status-bar-style" content="default"> + + <!-- Preconnect for performance --> + <link rel="preconnect" href="https://fonts.googleapis.com"> + <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin> + <link rel="preconnect" href="https://ajax.googleapis.com"> + <link rel="preconnect" href="https://documentcloud.adobe.com"> + <link rel="preconnect" href="https://cdn.jsdelivr.net"> + + + <title>Training Transformer Language Models Without Backpropagation</title> + <meta name="robots" content="noindex, nofollow, noarchive"> + + <!-- Favicon and App Icons --> + <link rel="icon" type="image/x-icon" href="static/images/favicon.ico"> + <link rel="apple-touch-icon" href="static/images/favicon.ico"> + + <!-- Critical CSS - Load synchronously --> + <link rel="stylesheet" href="static/css/bulma.min.css"> + <link rel="stylesheet" href="static/css/index.css"> + + <!-- Non-critical CSS - Load asynchronously --> + <link rel="preload" href="static/css/bulma-carousel.min.css" as="style" onload="this.onload=null;this.rel='stylesheet'"> + <link rel="preload" href="static/css/bulma-slider.min.css" as="style" onload="this.onload=null;this.rel='stylesheet'"> + <link rel="preload" href="static/css/fontawesome.all.min.css" as="style" onload="this.onload=null;this.rel='stylesheet'"> + <link rel="preload" href="https://cdn.jsdelivr.net/gh/jpswalsh/academicons@1/css/academicons.min.css" as="style" onload="this.onload=null;this.rel='stylesheet'"> + + <!-- Fallback for browsers that don't support preload --> + <noscript> + <link rel="stylesheet" href="static/css/bulma-carousel.min.css"> + <link rel="stylesheet" href="static/css/bulma-slider.min.css"> + <link rel="stylesheet" href="static/css/fontawesome.all.min.css"> + <link rel="stylesheet" href="https://cdn.jsdelivr.net/gh/jpswalsh/academicons@1/css/academicons.min.css"> + </noscript> + + <!-- Fonts - Optimized loading --> + <link href="https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600;700;800&display=swap" rel="stylesheet"> + + <!-- Defer non-critical JavaScript --> + <script defer src="https://ajax.googleapis.com/ajax/libs/jquery/3.5.1/jquery.min.js"></script> + <script defer src="https://documentcloud.adobe.com/view-sdk/main.js"></script> + <script defer src="static/js/fontawesome.all.min.js"></script> + <script defer src="static/js/bulma-carousel.min.js"></script> + <script defer src="static/js/bulma-slider.min.js"></script> + <script defer src="static/js/index.js"></script> + + <!-- Structured Data for Academic Papers --> + + + <!-- Website/Organization Structured Data --> + +</head> +<body> + + + <!-- Scroll to Top Button --> + <button class="scroll-to-top" onclick="scrollToTop()" title="Scroll to top" aria-label="Scroll to top"> + <i class="fas fa-chevron-up"></i> + </button> + + <!-- More Works Dropdown --> + <div class="more-works-container"> + <button class="more-works-btn" onclick="toggleMoreWorks()" title="View More Works from Our Lab"> + <i class="fas fa-flask"></i> + More Works + <i class="fas fa-chevron-down dropdown-arrow"></i> + </button> + <div class="more-works-dropdown" id="moreWorksDropdown"> + <div class="dropdown-header"> + <h4>More Works from Our Lab</h4> + <button class="close-btn" onclick="toggleMoreWorks()"> + <i class="fas fa-times"></i> + </button> + </div> + <div class="works-list"> + <a href="https://arxiv.org/abs/PAPER_ID_1" class="work-item" target="_blank"> + <div class="work-info"> + <h5>Paper Title 1</h5> + <p>Brief description of the work and its main contribution.</p> + <span class="work-venue">Conference/Journal 2024</span> + </div> + <i class="fas fa-external-link-alt"></i> + </a> + <a href="https://arxiv.org/abs/PAPER_ID_2" class="work-item" target="_blank"> + <div class="work-info"> + <h5>Paper Title 2</h5> + <p>Brief description of the work and its main contribution.</p> + <span class="work-venue">Conference/Journal 2023</span> + </div> + <i class="fas fa-external-link-alt"></i> + </a> + <a href="https://arxiv.org/abs/PAPER_ID_3" class="work-item" target="_blank"> + <div class="work-info"> + <h5>Paper Title 3</h5> + <p>Brief description of the work and its main contribution.</p> + <span class="work-venue">Conference/Journal 2023</span> + </div> + <i class="fas fa-external-link-alt"></i> + </a> + </div> + </div> + </div> + + <main id="main-content"> + <section class="hero"> + <div class="hero-body"> + <div class="container is-max-desktop"> + <div class="columns is-centered"> + <div class="column has-text-centered"> + <h1 class="title is-1 publication-title">Training Transformer Language Models Without Backpropagation</h1> + <div class="is-size-5 publication-authors"> + <span class="author-block">Author list forthcoming</span> + </div> + + <div class="is-size-5 publication-authors"> + <span class="author-block">University of Illinois Urbana-Champaign<br>2026</span> + </div> + + <div class="column has-text-centered"> + <div class="publication-links"> + <span class="link-block"> + <a class="external-link button is-normal is-rounded is-dark" style="pointer-events:none;opacity:.55"> + <span class="icon"><i class="fas fa-file-pdf"></i></span> + <span>Paper — in preparation</span></a> + </span> + <span class="link-block"> + <a class="external-link button is-normal is-rounded is-dark" style="pointer-events:none;opacity:.55"> + <span class="icon"><i class="fab fa-github"></i></span> + <span>Code — on request</span></a> + </span> + <span class="link-block"> + <a href="mailto:yurenh2@illinois.edu" class="external-link button is-normal is-rounded is-dark"> + <span class="icon"><i class="fas fa-envelope"></i></span> + <span>Contact</span></a> + </span> + </div> + </div> + </div> + </div> + </div> + </div> +</section> + + +<!-- TLDR --> +<section class="section"> + <div class="container is-max-desktop"> + <div class="columns is-centered"> + <div class="column is-four-fifths"> + <h2 class="title is-3">TL;DR</h2> + <div class="content has-text-justified"> + <ul> + <li><b>The largest models trained from scratch without backpropagation at any level.</b> + Every parameter update follows the Equilibrium Propagation (EP) rule — no layer, block, or + output head is trained with a backprop rule. The models are standard transformer LMs + (OLMo2-style blocks, 32k vocabulary, FineWeb-Edu); the trained network runs ordinary forward inference.</li> + <li><b>Matches backprop within a 4–5% perplexity band at 72M parameters</b>, against a backprop + twin trained on identical data, steps, and optimizer — with multi-seed controls on both sides.</li> + <li><b>First transformer language model in the EP family, at 2.15× the size of the largest prior + EP-family result</b> (VGG10 ImageNet classifier, ~63M params): models trained up to <b>135M</b> + parameters on 2.7B tokens.</li> + <li><b>2× wall-clock overhead versus backprop — where the closest prior EP work pays + 6.7–12×</b>: one nudged phase × 3 relaxation sweeps per step versus two phases × 10 iterations.</li> + <li><b>Verified gradient fidelity:</b> the EP update maintains cosine ≈0.99 to the true backprop + gradient throughout training — measured, at scale. Larger non-backprop transformers in the + literature train most parameters with <i>local backprop inside blocks</i>; ours use none.</li> + <li><b>Hardware-ready by measurement, not assumption:</b> 8-bit quantization shows no EP-specific + penalty against an equally-quantized backprop twin; under injected analog faults (1% forward noise, + 10% error-channel noise, device-tolerance mismatch) the EP estimator tracks the faulted network at + cosine ≈0.97 — learning co-adapts to the hardware.</li> + </ul> + </div> + </div> + </div> + </div> +</section> +<!-- End TLDR --> + +<!-- Paper abstract --> +<section class="section hero is-light"> + <div class="container is-max-desktop"> + <div class="columns is-centered has-text-centered"> + <div class="column is-four-fifths"> + <h2 class="title is-3">Abstract</h2> + <div class="content has-text-justified"> + <p> + Analog and physical accelerators promise order-of-magnitude energy savings for training, but they + cannot run backpropagation natively: exact gradients on a physical substrate require per-step + digitization or an adjoint copy of the hardware. Equilibrium Propagation extracts gradients from the + physics itself — two relaxations and local reads — and is the only member of its family with a + gradient-equivalence guarantee that we verify directly at scale. What the field has lacked is scale and + rigor: EP results stopped at mid-size vision models, without matched controls. We train standard + transformer language models end-to-end with EP — 72M parameters within 4–5% perplexity of a + matched backprop twin, models up to 135M — at 2× backprop wall-clock, with backprop twins, + multi-seed discipline, and hardware-relevant ablations (quantization, analog faults, nudge operating + windows, energy accounting) at every stage. Scaling a physical learning rule also surfaces new science: + we identified a width-scaling loss in the EP gradient invisible to per-step alignment metrics, built an + instrument that measures it in 90 minutes per candidate recipe, mapped its dose–response law, and + demonstrated an estimator-side treatment that recovers 97% of it without touching the model or the + cost budget. + </p> + </div> + </div> + </div> + </div> +</section> +<!-- End paper abstract --> + + +<!-- Results --> +<section class="section"> + <div class="container is-max-desktop"> + <div class="columns is-centered"> + <div class="column is-four-fifths"> + <h2 class="title is-3">Headline results</h2> + <div class="content"> + <table class="table is-fullwidth"> + <thead><tr><th>Model</th><th>Data</th><th>EP (val CE)</th><th>BP twin</th><th>Gap</th></tr></thead> + <tbody> + <tr><td>72M transformer LM</td><td>FineWeb-Edu, 1.4B tok</td><td>3.33</td><td>3.29</td><td>+4–5% ppl</td></tr> + <tr><td>135M transformer LM</td><td>FineWeb-Edu, 2.7B tok</td><td colspan="3">trained end-to-end, zero instability + events; scaling analysis below</td></tr> + </tbody> + </table> + <p class="is-size-6 has-text-grey">Twin discipline: identical architecture, tokenizer, data order, + optimizer, steps, and evaluation; multi-seed on both sides (BP n=3, band ±0.006; EP n=2).</p> + <table class="table is-fullwidth"> + <thead><tr><th>Cost vs backprop</th><th>This work</th><th>Closest EP work (VGG10, ImageNet)</th></tr></thead> + <tbody> + <tr><td>Wall-clock overhead</td><td><b>2.0×</b></td><td>6.7× (single-sided) / 12× (centered)</td></tr> + <tr><td>Relaxation iterations / step</td><td><b>3</b> (one phase)</td><td>20 (two phases × 10)</td></tr> + </tbody> + </table> + </div> + </div> + </div> + </div> +</section> + +<section class="section hero is-light"> + <div class="container is-max-desktop"> + <div class="columns is-centered"> + <div class="column is-four-fifths"> + <h2 class="title is-3">The scaling science</h2> + <div class="content has-text-justified"> + <p>Scaling a physical learning rule surfaces phenomena backprop never meets. Between widths 512 and 768 + we identified a width-scaling loss in the EP gradient — localized to the top half of the network, + invisible to every per-step alignment metric, and traced to response components that finite nudge + displacement under-reaches. We built a screening instrument that measures this leak in 90 minutes per + candidate recipe, mapped its dose–response law (logarithmic across two decades of displacement + amplification), and demonstrated a pure estimator-side treatment that closes 97% of it — no change to + the model, its inference path, or the cost budget. The same instruments provide the go/no-go protocol + for each next rung of the ladder.</p> + </div> + <h2 class="title is-3">Hardware line</h2> + <div class="content"> + <ul> + <li>Measured energy projection for an integrated weight-stationary realization: + <b>0.21–0.63 pJ/MAC</b> (SPICE-measured analog core + datasheet periphery), against a + 0.3–1 pJ/MAC digital INT8 system envelope.</li> + <li>Single-column analog prototype: SPICE-modeled, discrete multiplying-DAC parts list — kept at the + “hardware someone can actually build” level.</li> + <li>Nudge-amplitude operating windows and their evolution over training are mapped — the + dynamic-range spec an analog implementation must meet.</li> + </ul> + </div> + <h2 class="title is-3">Roadmap</h2> + <div class="content has-text-justified"> + <p>Staged scaling with matched BP controls and hardware-relevant ablations at every rung: a 150M–600M + ladder (does the gap grow or shrink with scale — measured, not assumed), then 1B–3B; each stage + gated on the previous stage’s loss, alignment, and throughput numbers. In parallel: the + algorithm→regime map across the activity-difference family (contrastive / coupled-learning arms on + the same harness), and a bounded single-column hardware feasibility study.</p> + </div> + </div> + </div> + </div> +</section> + + + +<!--BibTex citation --> + <section class="section" id="BibTeX"> + <div class="container is-max-desktop content"> + <div class="bibtex-header"> + <h2 class="title">BibTeX</h2> + <button class="copy-bibtex-btn" onclick="copyBibTeX()" title="Copy BibTeX to clipboard"> + <i class="fas fa-copy"></i> + <span class="copy-text">Copy</span> + </button> + </div> + <pre id="bibtex-code"><code>@misc{ept2026, + title={Training Transformer Language Models Without Backpropagation}, + author={(author list forthcoming)}, + year={2026}, + note={Project page} +}</code></pre> + </div> +</section> +<!--End BibTex citation --> + + + <footer class="footer"> + <div class="container"> + <div class="columns is-centered"> + <div class="column is-8"> + <div class="content"> + + <p> + This page was built using the <a href="https://github.com/eliahuhorwitz/Academic-project-page-template" target="_blank">Academic Project Page Template</a> which was adopted from the <a href="https://nerfies.github.io" target="_blank">Nerfies</a> project page. + You are free to borrow the source code of this website, we just ask that you link back to this page in the footer. <br> This website is licensed under a <a rel="license" href="http://creativecommons.org/licenses/by-sa/4.0/" target="_blank">Creative + Commons Attribution-ShareAlike 4.0 International License</a>. + </p> + + </div> + </div> + </div> + </div> +</footer> + +<!-- Statcounter tracking code --> + +<!-- You can add a tracker to track page visits by creating an account at statcounter.com --> + + <!-- End of Statcounter Code --> + + </body> + </html> |
