Machine learning performance tools built from design documents not code

Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool

Programming LanguagesArtificial Intelligence

Summary

Machine learning systems change quickly, making it hard to keep their performance analysis tools up to date. The authors show a new approach where the driving documents for the tool are written in clear, detailed natural language instead of code. Then AI agents automatically rewrite the tool’s code from these documents whenever updates are needed. This method reproduces complex models very accurately and keeps the tool easy to maintain as machine learning evolves.

machine learningperformance modelingAI code generationdesign documentationsymbolic computationintermediate representationschedulingtech debtsoftware maintenanceanalytical modeling

Authors

Samuel Kushnir, Kimia Noorbakhsh, Kavya Sreedhar, Liqun Cheng, Ming Liu, Parthasarathy Ranganathan, Mohammad Alizadeh, Fred Kjolstad, Suvinay Subramanian

Abstract

Machine-learning performance modeling is a uniquely hostile terrain for long-lived software: the assumptions baked into today's abstractions are invalidated by tomorrow's models and systems, forcing perpetual refactoring of performance-modeling frameworks. Meanwhile, AI coding agents have become fast and capable enough that regenerating an entire library is cheaper than paying down the tech debt of incrementally patching it. We describe SMART, a rigorous symbolic performance-modeling library for ML systems whose main branch contains almost no code: the repository is a DAG of self-contained natural-language design docs, coding sub-agents regenerate the implementation from only the docs on new version updates, and every human change is a natural-language edit to a doc--self-documenting by construction. Two ingredients make regeneration reliable: (i) a design-doc style built around step-by-step worked examples that act as in-context demonstrations for the generating agents, and (ii) a minimal, recursively defined operator IR with symbolic (SymPy) cost expressions, a fast analytical roll-up mode for large sweeps, and a slow modulo-scheduling mode for fine-grained schedule studies. Regenerated implementations reproduce hand-audited reference models--including DeepSeek-V3 serving on a TPU pod slice--to round-off precision, suggesting that design docs--not code--can be the durable artifact for ML-systems co-design tools.