Local checks improve safety and accuracy in network automation
Can You Check That? The Checkability Boundary for Local LLM Network Automation
Networking and Internet ArchitectureArtificial Intelligence
Summary
Sending sensitive network data to big online language models can risk privacy, so running smaller models locally is safer but less reliable for some tasks. The authors introduce the idea of 'checkability,' meaning a task is suitable for local models if there is a fast and clear way to check if the answers are correct. They build a system called Touchstone that first tries local small models and checks outputs before asking a big model for help only when needed. Their approach works well on tasks like detecting conflicts or understanding commands while keeping most data local, but it doesn't perform as well on tasks without clear checks.
What this means in practice
- •For network operations teams: Use local small language models combined with task-specific checks to automate configuration and intent tasks with high accuracy and minimal data exposure.
- •For cybersecurity teams: Implement local-only network data processing pipelines that automatically escalate uncertain cases to trusted online models, reducing leakage of sensitive configurations.
Authors
Maleeha Masood, Momina Nofal
Abstract
Sending every network-automation input to a third-party frontier LLM exports sensitive artifacts such as production configurations, topologies, and logs. Querying small language models (SLMs) locally avoids this egress, but SLM outputs can be error-prone for direct use. This work introduces checkability as a criterion for determining which tasks are suitable for local inference. A task is checkable when it exposes a cheap, deterministic test - an intrinsic check - that rejects outputs violating a necessary correctness condition. We instantiate this idea in Touchstone, a local-first pipeline that uses seven off-the-shelf SLMs (1-8B parameters) to generate candidates, uses task-specific intrinsic checks to reject responses, and escalates unresolved inputs to a frontier LLM. On conflict detection and intent translation tasks, Touchstone reaches 98.6% and 93.8% end-to-end accuracy while escalating only 16% and 17% of inputs, respectively. On TeleQnA, a knowledge-only control that has no task-specific intrinsic checks, Touchstone is unable to match the accuracy of the frontier baseline. Our results support a simple deployment rule: keep inference local when task semantics support precise, low-cost checks; escalate the rest.