Agents in the Wild: Where Research Meets Deployment

2026-07-21Artificial Intelligence

Artificial IntelligenceComputation and Language
AI summary

The authors discuss advanced AI systems that use large language models to think, plan, act, and work with tools or other AIs. These systems are moving from research to real-world uses in areas like science and finance. The tutorial covers challenges in making these AI systems safe, reliable, and robust during deployment. By examining real examples in drug discovery and finance, the authors identify design patterns and strategies to handle failures, including checks, backups, and human oversight. Attendees learn practical methods to deploy these systems responsibly across different industries.

large language modelsagentic systemsreasoning and planningmulti-agent coordinationdeployment challengesrobustnesssafetyverification pipelineshuman-in-the-loopfallback mechanisms
Authors
Grace Hui Yang, Pranav N. Venkit, Hooman Sedghamiz, Enrico Santus, Victor Dibia, Ioana Baldini
Abstract
Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and finance. While academic work has emphasized benchmarks and algorithmic innovation, deployment raises new challenges around robustness, safety, and reliability. This tutorial brings together researchers and practitioners to explore advances in reasoning and planning, multi agent coordination, and evaluation, highlighting open challenges arising from deployment experience. Through applied case studies in pharmaceutical discovery and financial systems, we analyze common design patterns that make agentic systems successful, and discuss practical mitigation strategies for failure modes, such as verification pipelines, fallback mechanisms, and human in the loop supervision. Attendees will gain a comprehensive view of the field along with concrete design patterns, evaluation checklists, and templates for safe and reliable deployment across industries.