Large language models generate executable multi-agent system specifications

LLMs for Executable Multi-Agent System Specification Generation

Artificial Intelligence

Summary

Specifying how multiple agents act and interact over time usually requires formal expertise and labeled examples, which are hard to obtain. The authors propose a method called genRTEC that uses large language models to turn simple natural language descriptions into executable specifications. These specifications can help monitor agent behavior during runtime and handle complex dependencies. Their tests show that genRTEC produces accurate and efficient specifications without needing extensive expert input.

What this means in practice

  • For software developers: Automatically generate executable behavior rules from simple descriptions for multi-agent applications to enable easier runtime monitoring.
  • For automation engineers: Build and verify complex agent systems by creating formal specifications from plain language without deep formal methods expertise.

Authors

Andreas Kouvaras, Periklis Mantenoglou, Alexander Artikis

Abstract

MAS specifications express the effects of the actions of the agents and their environment, as well as other temporal phenomena, such as the intervals during which an agent may perform an action. The specification of a MAS should also be executable in order to allow for run-time monitoring. Constructing the specification of a MAS requires formal language expertise, while machine learning techniques depend on labelled data which are rarely available. To address these issues, we propose `genRTEC', a method that leverages pre-trained Large Language Models (LLMs) to generate executable MAS specifications, in the language of the `Run-Time Event Calculus' (RTEC), from natural language descriptions. genRTEC constructs MAS specifications with complex hierarchical and cyclic dependencies based only on short natural language descriptions of the concepts involved. We present an extensive empirical evaluation of genRTEC, spanning various MAS specifications, including both a qualitative and a quantitative assessment. Our results demonstrate that genRTEC constructs executable MAS specifications of high predictive accuracy without compromising reasoning efficiency.