openTPU is a monorepo containing every layer of a working AI accelerator: the SystemVerilog hardware design, the instruction set architecture, a bit-exact Python simulator, a kernel language with compiler, a browser-based profiler, and host software that drives a real PCIe card. The project was built using AI agents as hardware designers, extending the methodology of auto-arch-tournament to silicon-level design. It runs on an Inspur YPCB-00338 card with a Xilinx Kintex-7 xc7k480t FPGA and two DDR3 channels, a board you can buy used for a couple hundred dollars. The results are concrete and verifiable. Ten models — including Qwen3-0.6B, Qwen3.5-0.8B, LFM2.5-230M, Gemma 4 E2B/E4B, SmolLM3-3B, Phi-4-mini (3.8B), and Qwen3.5-4B — run on the card and produce tokens matching the simulator bit for bit. LFM2.5-230M decodes at 85.8 tok/s in 4-bit mode; Qwen3-0.6B at 31.3 tok/s; the 3.8B Phi-4-mini at 6.56 tok/s. DRAM utilization consistently hits 82-94% of the DDR3-1066 peak of 17.1 GB/s. Mixture-of-experts models that exceed the card's 4 GiB memory run via expert streaming from the host — Qwen3.5-35B-A3B (34.7B parameters, 3.0B active) decodes at 3.95 tok/s with experts streamed over PCIe at 1.41 GB/s. The architecture is deliberately simple and fully transparent. A sequencer issues one instruction per cycle to a DMA unit, a four-column systolic matrix unit (int8), a vector unit (fp32), and a quantizer. There is no cache and no hidden scheduling — every data movement is an explicit instruction, which means the included profiler (Lens) can show exactly where every cycle goes. The ISA uses 8×32-bit instruction words. Kernels are written in a Python DSL that compiles down through an affine loop addressing system with fusion. The LiteDRAM memory controllers calibrate autonomously via a small on-chip CPU in 12 seconds, with no host involvement. The 4-bit quantization scheme uses FP4 values with two-level block scales at 4.25 bits per weight, keeping the LM head in int8 for accuracy. This cuts bytes per token by roughly a third and boosts decode speed by 40-45%, with a documented perplexity cost reported per model. The host overhead is nearly eliminated for smaller models: LFM2 and Qwen3 decode adds only 0.17-0.30 ms per token on the faster host machine. Two production bitstream builds are documented with dates (September 29 to October 1, 2026) and commit hashes. Build B (deployfused133c79c5707a) improved decode by 8-10% and raised DRAM efficiency from 82-87% to 91-94% for larger models. The design closes timing at 133.33 MHz with a worst negative slack of just +0.032 ns — functional but tight. The generative potential here is significant. This is a complete, readable, open-source reference implementation of an AI accelerator stack — the kind of artifact that previously existed only inside Google (TPU), Groq, or Cerebras. It lowers the barrier to understanding and experimenting with custom AI hardware from "hire a team of 50 ASIC engineers" to "clone a repo and read the docs." The fact that AI agents did much of the design work is itself a data point about the capability frontier of AI-assisted hardware development. The constraints are real: DDR3-1066 bandwidth caps throughput, the timing margin is razor-thin, and the largest models require host-side expert streaming. But the project isn't claiming to compete with an H100 — it's demonstrating that a legible, AI-designed accelerator can run real models with real weights on real hardware, bit-exact against simulation, with the entire stack open for inspection. That's a different and arguably more important achievement.