Seven ways physical AI capabilities can develop in diverse systems
Seven Sources of Physical AI Capability Formation
Artificial Intelligence
Summary
Capabilities in physical AI systems arise in many ways, but existing classifications don’t clearly explain how these abilities form. The authors identified seven key sources that contribute to forming capabilities, such as learning from experience, building predictive models, and evolutionary processes. They analyzed many studies and found all examples fit into these seven categories without needing new ones. This framework helps understand how similar abilities can come from different origins, which is useful for explaining, copying, or governing AI systems.
What this means in practice
- •For robotics developers: Identify different foundational sources of robot capabilities to tailor training and design strategies for robotic systems effectively.
- •For ai governance teams: Use the framework to analyze how physical AI capabilities formed, supporting better evidence gathering and regulatory decisions.
A survey. It maps existing work.
Authors
Gang Chen
Abstract
Capabilities relevant to Physical AI can arise from materially different formation histories, yet existing taxonomies organized by morphology, architecture, learning algorithm, task, or domain do not directly answer what gives rise to a capability. We define a capability-formation source as a factor materially contributing to capability formation, distinct from components or construction steps. We identify seven non-exclusive sources: Recorded-Experience (RE), Predictive-Modeling (PM), Evaluative-Interaction (EI), Surrogate-Environment (SE), Mechanism-Grounded (MG), Embodied-Coupling (EC), and Evolution-Driven (ED) Formation. Using reconstructive induction with theoretical saturation, we traced a research matrix to primary studies, deduplicated the literature, set coding rules, and conducted three rounds of maximum-difference and negative-case sampling. Challenges included curriculum and self-supervised learning, active inference, open-ended and developmental learning, planning and search, neuro-symbolic architectures, digital twins, generative physical world models, and morphology-control co-design. Within the scope and criteria fixed as of September 4, 2026, all 49 evidence records were explainable by the seven sources individually or in combination. No R1-R3 challenge produced an irreducible eighth source, and R3 required no new core definition or substantive boundary rule. We therefore claim theoretical saturation within the stated scope, not logical completeness or exhaustive future coverage. The framework distinguishes similarity in observed capability from similarity in how it was formed, supporting analysis of explanation, transfer, replication, dependencies, governance evidence, and geoeconomic foundations.