Human Grounded Evaluation of Large Language Models for Optical Network Automation
2026-07-20 • Networking and Internet Architecture
Networking and Internet ArchitectureArtificial Intelligence
AI summaryⓘ
The authors created HuGLEN, a method to compare different large language models (LLMs) by using another LLM as a judge along with some expert ratings. They used it to turn technical outputs from an AI model that estimates network quality into easy-to-understand explanations for operators. Their tests showed that a medium-sized LLM gave the best balance between explanation quality and speed. HuGLEN helps reduce the need for lots of human labeling and makes choosing the right model easier for network automation tasks.
Large Language ModelsNetwork AutomationExplainable AIQuality of TransmissionModel EvaluationInference CostExpert RatingsHuman LabelingModel SelectionQuality Efficiency Score
Authors
Kiarash Rezaei, Omran Ayoub, Paolo Monti, Carlos Natalino
Abstract
Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substantially across LLM families. We present HuGLEN, a stepwise evaluation pipeline that uses an LLM-as-a-judge together with a small set of expert ratings to enable scalable and reproducible comparison of candidate LLMs, and to rank them using a quality efficiency score (QES). We demonstrate HuGLEN for translating outputs from an explainable artificial intelligence (XAI) model for the optical network quality of transmission (QoT) estimation task into operator-friendly explanations. Our results show that a medium-sized LLM (12B parameters) achieves the highest QES, indicating the best trade-off between explanation quality and efficiency. Overall, HuGLEN reduces the human-labeling burden while supporting consistent model selection for operator-facing automation tasks.