openTPU, an open-source AI accelerator built by AI, runs ten models on an FPGA card

openTPU is an open-source AI accelerator whose README describes it as "developed by AI". It builds on the project's earlier auto-arch-tournament and asks two questions: how far can AI agents go at hardware design, and can they build the chip that runs their own inference? The README also pitches it as a learning project. Everything sits in one small monorepo that can be read end to end: the SystemVerilog hardware design, the instruction set, a bit-exact simulator, a kernel language with its compiler, and the host software that drives a real PCIe card. The license is Apache 2.0.
The design runs ten modern models with their real weights on an Inspur YPCB-00338 FPGA card (Xilinx Kintex-7 xc7k480t, two DDR3 channels). The README says the card produces the same tokens as the simulator, bit for bit, and that every configuration matches the simulator token for token. Measurements were taken between 2026-09-29 and 2026-10-01. The first three models were measured on 2026-09-29 with production image deploy_champ_e698dcd7. LFM2-2.6B, SmolLM3-3B and Phi-4-mini followed on 2026-09-30, and Qwen3.5-2B, Qwen3.5-4B and Gemma 4 on 2026-10-01, with build B (deploy_fused133c_79c5707a), which has been in production since.
The image runs at 133.33 MHz with one bitstream for all models. It has LiteDRAM controllers calibrated by a small CPU inside the memory core, a four-column systolic matrix unit and a stream engine. Memory is DDR3-1066 with a 17.1 GB/s peak. The host is an Intel Core i7-4790 machine called opentpu. The benchmark method (tools/qual/perf.py) is 64 greedy tokens after a 512-token prompt, with the host's argmax in the loop. "Device" counts only the cycles the accelerator runs; "wall" adds the host. Prefill is the 512-token prompt, on the device.
Build B decodes LFM2-2.6B, SmolLM3 and Phi-4-mini 8-9% faster than e698dcd7 (Gemma 4 E2B 10%), reaching 91-94% of the DRAM peak instead of 82-87%. With logits streamed back while the card runs (96 tokens), 4-bit decode in device / wall tokens per second is: LFM2 89.5 / 84.5, Qwen3 33.7 / 33.3, Qwen3.5 24.6 / 24.2, LFM2-2.6B 11.07 / 11.02, SmolLM3-3B 8.92 / 8.89 and Phi-4-mini 6.69 / 6.67 (the last three on build B). In the card's own decode loop, where the card picks every token, Gemma 4 E2B decodes at 11.01 tok/s with an int8 head and 12.73 tok/s with a 4-bit head, and E4B at 3.83 tok/s, all on the device. E2B keeps its per-layer embedding tables on the card (images of 3.5 to 3.6 GiB) and matches Hugging Face's greedy tokens on three prompts. E4B's 2.95 GB table stays on the host, which copies one 11 KB row into the card per token; its image is 3.96 GiB.
Against the previous production image, se-cand3 (Xilinx MIG, two-column matrix unit, 120.755 MHz), the new image decodes within 2.3% in every configuration, because decode is bound by DRAM and LiteDRAM reads at 82-85% of the DDR3 peak, as the MIG did. Prefill is 1.3x (Qwen3.5) to 2.0x (LFM2 4-bit) faster. At start, the core's CPU calibrates both DDR3 channels in 12 s with no host involvement.
Mixture-of-experts models larger than the card's 4 GiB run with experts streamed from host storage. The card routes each token and computes every expert, keeping experts in per-layer slots in its DRAM; the host only copies missing experts into those slots. Measured on 2026-10-01 with build B, with 4-bit experts and an int8 head: LFM2.5-8B-A1B (8.5B parameters, 1.7B active) reaches 10.6 tok/s over 160 tokens, with 98.5% of expert uses hitting the slots and 5.2 MB streamed per token. Qwen3.5-35B-A3B (34.7B parameters, 3.0B active) reaches 3.95 tok/s, with 62% of expert uses hitting and 153 MB streamed per token at 1.41 GB/s over PCIe. Both match the simulator bit for bit.
The 4-bit format uses FP4 values with two-level block scales, at 4.25 bits per weight, and keeps the LM head in int8. It cuts bytes per token by about a third and raises decode speed by 40% (Qwen3.5) to 45% (Qwen3, LFM2), at a measurable cost in perplexity that docs/quant.md reports per model. The host is described as nearly out of the way: for LFM2 and Qwen3 the card runs one decode program compiled once, and the host adds 0.17 to 0.30 ms per token on omarchy (0.45 to 1.3 ms on opentpu).
The machine is deliberately simple. A sequencer issues one instruction per cycle to a few units: DMA, a matrix unit that multiplies int8 weights streamed from DRAM, a vector unit doing fp32 math, and a quantizer that turns results back into int8. There is no cache and no hidden scheduling, so every data movement is an instruction and a trace shows where the cycles go. Kernels are written in Python with an @ol.jit decorator in a language with a compiler that handles layouts, affine loop addressing and fusion. Lens, the project's profiler, records a run from the RTL, the simulator or the card and opens it in the browser with a roofline, a timeline and per-instruction tables.
The README lists three roadmap items: the last few percent of DRAM efficiency (LiteDRAM path work is under way), timing margin and area (the design closes 133.33 MHz with worst negative slack of only +0.032 ns), and faster prefill, which is still limited by the matrix unit's multiply rate. Everything except the card runs on a laptop, and most contributions need only Python and Verilator.
Key facts
- The openTPU README describes an open-source AI accelerator "developed by AI": SystemVerilog hardware, ISA, bit-exact simulator, kernel language and compiler, and host software in one monorepo under Apache 2.0.
- It runs ten modern models with real weights on an Inspur YPCB-00338 FPGA card (Kintex-7 xc7k480t, two DDR3 channels) at 133.33 MHz, and the README says the card matches the simulator token for token.
- Build B decodes LFM2-2.6B, SmolLM3 and Phi-4-mini 8-9% faster than the earlier production image, at 91-94% of the 17.1 GB/s DRAM peak instead of 82-87%.
- Mixture-of-experts models larger than the card's 4 GiB work by streaming experts from the host: LFM2.5-8B-A1B reaches 10.6 tok/s and Qwen3.5-35B-A3B reaches 3.95 tok/s.
- Decode is bound by DRAM bandwidth; the timing margin at 133.33 MHz is only +0.032 ns worst negative slack.
Why it matters
The project puts a concrete test behind a question many people ask: can AI agents design the chip that runs their own inference? The README states that aim directly. What it publishes is a complete, working stack, from a matmul in Python down to RTL and a physical PCIe card, with models running on real hardware and checked against a bit-exact simulator. It also builds on the team's earlier auto-arch-tournament. The numbers are modest, with tokens per second ranging from single digits to under 90 depending on model, but the point is that the whole chain exists and can be inspected.
Who it affects
Hardware and ML-systems engineers and students are the clear audience: the README calls openTPU a learning project and points readers to docs on the ISA, compiler, simulator and board bring-up. Anyone with a Kintex-7 card of this type can try to build the bitstream. Everyone else can run the simulator on a laptop, since everything except the card runs there. Contributors need mostly Python and Verilator, not an FPGA.
How to use it
Install with pip install -e ., add pytest, torch and transformers, and run python3 -m pytest -q (the RTL tests also need Verilator 5). Download a model such as LiquidAI/LFM2.5-230M with hf download, then chat on the simulator with otpu-chat --model lfm2 --backend isa. With the Inspur YPCB-00338 card, build the bitstream with make bit in boards/ypcb-00338, load it over JTAG, then run sudo otpu-setup and otpu-chat --backend board. docs/board.md walks through bring-up, and docs/isa.md is the suggested starting point for reading. The license is Apache 2.0.
How solid is it
The performance figures come from the project's own README, with the method stated: 64 greedy tokens after a 512-token prompt, device versus wall time separated, DRAM traffic read from the card's own counters, and measurement dates and image builds named. The README says every configuration matches the simulator token for token, and Gemma 4 E2B matches Hugging Face's greedy tokens on three prompts. The text does not say which AI agents or models designed the hardware, nor how much human involvement there was, so the "developed by AI" framing rests on the README's own description.
Risks and caveats
Decode is bound by DRAM, and LiteDRAM reads at 82-85% of the DDR3 peak, so the design gains little from a faster clock. The 133.33 MHz timing closes only just, with worst negative slack of +0.032 ns. 4-bit weights speed up decode by 40% to 45% at a measurable cost in perplexity, which the README defers to docs/quant.md. Large mixture-of-experts models are limited by PCIe streaming: Qwen3.5-35B-A3B hits the on-card slots for 62% of expert uses and streams 153 MB per token at 1.41 GB/s. The text gives no comparison with GPUs, Google TPUs or other commercial accelerators, and no power figures, so how it ranks against them is not stated. Prefill is still limited by the matrix unit's multiply rate.
“how far can AI agents go at hardware design, and can they build the chip that runs their own inference?”
— openTPU README