LLMs struggle to accurately reason network protocol state machines
RFCLLM: Evaluating LLMs' Reasoning Ability of Network Protocol State Machines
Computation and Language
Summary
When designing network protocols, people write rules in plain language that need to be carefully translated into formal state machines to ensure they work correctly. This paper studies how well large language models (LLMs) can understand those plain language rules and represent them as state machines. The researchers found that LLMs often do not fully capture the precise behavior described in the specifications. They tested many tasks across different network protocols to better understand where and why LLMs make mistakes in reasoning about these state machines.
What this means in practice
- •For network protocol developers: Assess automated tools that use LLMs for generating and verifying protocol state machines to identify correctness gaps.
- •For security testing teams: Evaluate the reliability of LLM-based models in reproducing protocol behavior for security audit and fuzz testing.
Authors
Anqi Chen, Dan Goldwasser, Cristina Nita-Rotaru
Abstract
Mapping textual specifications into formal representations is essential for ensuring the correctness of protocol designs and implementations. LLM-generated mappings, used for networking security or testing, are assumed to capture a perfect understanding of the specification, which may not hold in practice. The goal of this paper is to assess the extent to which LLMs can interpret the specification correctly. We examine the degree to which an LLM's implicit representation of a finite-state transition system-defined via natural language descriptions-aligns with a manually generated ground-truth model. We designed 4 tasks and 1482 task queries for 16 protocols. We evaluated different judge biases, observed the inherent difficulty gaps between tasks, looked into the effect of 4 context types, and the influence of protocol characteristics. Our work contributes to a step toward verifying whether LLMs can really be trusted in FSM (Finite State Machine) reasoning of protocol specifications.