Hardware accelerator speeds up diffusion transformer machine learning models
The World Model Hardware Accelerator
Hardware Architecture
Summary
Many machine learning models generate data step by step, which can be slow. The authors worked on a type of model called diffusion transformers that process all steps at once instead of one by one. They designed a special computer chip that speeds up these models by making the process more parallel and efficient. The chip was tested carefully and showed great accuracy and performance improvements for various model tasks.
What this means in practice
- •For machine learning engineers: Run diffusion transformer models faster and with lower latency using a specialized hardware accelerator.
- •For chip designers: Design and verify high-performance inference accelerators for emerging AI models using the detailed micro-architecture and verification approaches.
- •For image generation developers: Accelerate real-time latent diffusion workflows by integrating a tailored hardware accelerator to reduce inference time.$Commercial implications: Enables production of faster AI image generation hardware products by improving inference speed and efficiency.
Authors
Shashank Chaurasia
Abstract
Diffusion transformers invert the arithmetic that autoregressive decoding made familiar. There is no token-by-token recurrence: every denoising step is a full-sequence forward pass over static shapes, so the entire schedule is known at compile time and the only serial dimension is the step count itself. We exploit that structure in WMHA, a latency-first diffusion-transformer inference accelerator: a very-long-instruction-word sequencer issues four engines from one instruction word, a weight-stationary 16x16 dual-dot array streams FP8 and BF16 contractions, and a single-pass online-softmax attention pipeline keeps keys and values resident through a skewed software pipeline. The design is specified in a frozen micro-architecture document, implemented in synthesizable SystemVerilog, and verified against a double-precision reference model by a UVM environment whose acceptance criterion is semantic: the device must run a real denoising trajectory and reduce mean squared error against a clean latent by at least a factor of ten. It does so by a factor of 23, at both synthesized configurations, with zero element failures across 237 million checked values. Eleven application benchmarks built from published model shapes, including the original diffusion-transformer configuration, run on the device and report measured occupancy beside separately labelled projections. Five engines are taken to routed layout in sky130 with parasitic-annotated timing and measured-activity power; the full chip is synthesized, and the host limit that stopped its place-and-route is quantified together with the machine that would remove it.