OpenTPU: AI-Built Open-Source AI Accelerator

Original: OpenTPU – An open-source AI accelerator, developed by AI

Why This Matters

Demonstrates AI agents producing functional, end-to-end hardware accelerator designs.

Developer FeSens released openTPU, an open-source AI accelerator built by AI agents, on GitHub. It includes RTL, ISA, simulator, compiler, and profiler—and runs Qwen3, LFM2.5, and Qwen3.5 on a Kintex-7 PCIe FPGA card.

OpenTPU is a single-repo AI accelerator project where AI agents handled the hardware design work—asking, per the README, 'how far can AI agents go at hardware design, and can they build the chip that runs their own inference?' The repo bundles everything: SystemVerilog RTL, a custom ISA, a bit-exact simulator, a kernel compiler, and host PCIe software. It already runs three real models—Qwen3, LFM2.5-230M, and Qwen3.5—on a Kintex-7 FPGA card, with a companion tool called otpu-smi displaying live card utilization and DRAM bandwidth. The project bills itself as a learning resource, letting anyone trace a matrix multiplication from Python all the way down to hardware. It has 204 GitHub stars and 8 forks since launch, licensed under Apache-2.0.

Source

github.com — Read original →