FSGen: Agile Fused and Sparse Accelerator Generator with Accurate Power Model for LLM Applications
2026-08-10 • Hardware Architecture
Hardware Architecture
AI summaryⓘ
The authors introduce FSGen, a new tool to help design better computer chips that run large language AI models more efficiently. FSGen can explore many design options faster and finds chip designs that use less power or run much faster than previous methods. Their approach includes better estimators for key chip features, making the design process quicker and more accurate. Overall, FSGen helps create chip designs that work better across various AI language tasks.
large language modelsAI chip acceleratordesign space explorationpower efficiencyperformance estimationfused operator dataflowssparsityPareto-optimal designsfigures of meritearly-stage estimator
Authors
Jay Zhe-An Mok, Qijun Zhang, Zhiyao Xie
Abstract
With the growing demand of artificial intelligence (AI) applications, large language models (LLMs) have become important workloads in many domains. The question of how to efficiently generate optimal AI chip accelerator designs remains unresolved and challenging. Currently, there is a lack of end-to-end design methodologies for efficient design space exploration (DSE). We propose FSGen, an agile framework for attention-based LLM accelerator generation with an early-stage PPA estimator. FSGen supports fused operator dataflows and sparsity with a diverse design space and finds designs with 1.4x better power efficiency or 10x speedup with similar PPA metrics compared to prior work. Pareto-optimal designs have much better performance over a wide range of LLM benchmarks and have 58x better figures of merit (FoM). Design exploration is also faster due to our PPA estimators, which have better accuracy than prior art and reduce DSE runtime drastically.