Papers for

network operations teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Modeling uncertainty and dependency improves multivariate time series anomaly detection

GT-PSSM: Unified Probabilistic Framework for Stochastic Dynamics Modeling and Dependency Learning in Multivariate Time Series Anomaly Detection

Abstract: Multivariate time series anomaly detection (MTAD) is crucial for ensuring the safe and reliable operation of complex systems. Many existing methods learn normal patterns by training reconstruction or forecasting models on predominantly normal data. However, a large portion of these approaches rely on deterministic models and their associated point-wise output errors for anomaly scoring. Since real-world multivariate time series are inherently stochastic due to measurement noise and intrinsic system randomness, purely error-based scores can be unreliable, as large errors may arise from benign fluctuations rather than true anomalies. Probabilistic approaches address this limitation by quantifying uncertainty in model outputs. In particular, probabilistic state-space models (PSSMs) provide a principled framework by modeling stochastic system dynamics through latent state transitions and measurement noise via emission models. Despite this advantage, existing PSSM-based MTAD methods often struggle to capture long-range temporal dependencies and inter-variable dependencies, as they typically rely on noise-sensitive recurrent architectures and lack explicit cross-variable structure modeling. To address these limitations, we propose Graph-Transformer-Enhanced Probabilistic State-Space Model (GT-PSSM), a novel PSSM-based MTAD method that tightly integrates PSSM-based probabilistic modeling of stochastic dynamics with Graph Transformer-based learning of temporal and inter-variable dependencies. By jointly modeling stochasticity, long-range temporal dependence, and variable interactions within a unified probabilistic framework, GT-PSSM enables more robust anomaly detection.

Mon 28 SeptMachine Learning
The gist
Finding unusual events in data that changes over time and across many parts is important for safety. The authors point out that usual methods often treat these data as predictable and may mistake normal random changes for problems. They created a new approach that combines a way to handle randomness with a method to understand both how things change over time and relate to each other. This combined approach helps detect true problems more reliably.
Open → 2609.34161v1

Local checks improve safety and accuracy in network automation

Can You Check That? The Checkability Boundary for Local LLM Network Automation

Abstract: Sending every network-automation input to a third-party frontier LLM exports sensitive artifacts such as production configurations, topologies, and logs. Querying small language models (SLMs) locally avoids this egress, but SLM outputs can be error-prone for direct use. This work introduces checkability as a criterion for determining which tasks are suitable for local inference. A task is checkable when it exposes a cheap, deterministic test - an intrinsic check - that rejects outputs violating a necessary correctness condition. We instantiate this idea in Touchstone, a local-first pipeline that uses seven off-the-shelf SLMs (1-8B parameters) to generate candidates, uses task-specific intrinsic checks to reject responses, and escalates unresolved inputs to a frontier LLM. On conflict detection and intent translation tasks, Touchstone reaches 98.6% and 93.8% end-to-end accuracy while escalating only 16% and 17% of inputs, respectively. On TeleQnA, a knowledge-only control that has no task-specific intrinsic checks, Touchstone is unable to match the accuracy of the frontier baseline. Our results support a simple deployment rule: keep inference local when task semantics support precise, low-cost checks; escalate the rest.

Fri 25 SeptNetworking and Internet ArchitectureArtificial Intelligence
The gist
Sending sensitive network data to big online language models can risk privacy, so running smaller models locally is safer but less reliable for some tasks. The authors introduce the idea of 'checkability,' meaning a task is suitable for local models if there is a fast and clear way to check if the answers are correct. They build a system called Touchstone that first tries local small models and checks outputs before asking a big model for help only when needed. Their approach works well on tasks like detecting conflicts or understanding commands while keeping most data local, but it doesn't perform as well on tasks without clear checks.
Open → 2609.31540v1

Graph dynamics model predicts changing networks in uncertain settings

Beyond Static Graph World Models: Learning Stochastic Latent Dynamics over Evolving Topologies

Abstract: Graph-based world models have recently emerged as a means of learning transitions over relational state representations. However, existing approaches are largely limited to fixed-topology graphs or deterministic, fully observable environments. We propose the Graph Dynamics Model (GDM), a world model for graph-structured observations that is designed to handle the more general setting of evolving topologies in stochastic and partially observable environments. The GDM uses a sparse recurrent adjacency matrix to model topology updates and perform message passing, together with a recurrent state-space architecture for modelling stochastic transitions. Furthermore, we identify a gap in the evaluation of graph-based world models, as existing methods do not provide a means of comparing predicted and true distributions over the joint graph state comprising the interdependent topology, node features, and graph features. We therefore introduce the Graph Distribution Distance (GDD) metric, which uses maximum mean discrepancy with a graph kernel to comprehensively compare joint next-state distributions. We evaluate the GDM across several environments, including stochastic and partially observable settings. We demonstrate that GDM outperforms baseline models and displays zero-shot generalisation on large graphs.

Wed 23 SeptMachine LearningArtificial IntelligenceSocial and Information Networks
The gist
Many models assume networks (graphs) don’t change over time, but real-world situations often involve networks that evolve unpredictably. This paper presents the Graph Dynamics Model (GDM), which can learn and predict how both the structure and attributes of a network change in environments that have randomness and only partial observations. The authors also introduce a new way to measure how well these models capture the true behavior of evolving networks. Their experiments show that the GDM produces better predictions and can handle larger networks than previous methods.
Open → 2609.28670v1

CacheDyG speeds up learning on dynamic graphs with less memory

CacheDyG: Decoupling Temporal Propagation for Efficient Dynamic Graph Learning

Abstract: Dynamic graphs are widely used to model time-evolving relational systems in real-world applications. Dynamic graph neural networks provide an effective framework for capturing both structural dependencies and temporal dynamics in such data. However, they typically intertwine temporal graph propagation with every optimization epoch and often maintain large trainable representations for each node-time pair. This design repeatedly recomputes largely unchanged historical structures, leading to substantial training and parameter overhead. To address this critical issue, we propose CacheDyG, a Cache-refine framework for efficient Dynamic Graph learning. Specifically, it decouples temporal propagation from routine parameter updates by constructing a time-ordered temporal dependency cache that stores graph-aware node-time representations in non-trainable buffers. During standard training epochs, CacheDyG reads from the cache and updates only a lightweight cache refiner, an adaptive residual gate, and the link predictor. Selective cache refresh further keeps cached representations aligned with the supervised objective while avoiding epoch-wise sparse propagation. Experiments on five dynamic graph benchmarks show that CacheDyG adopts substantially fewer trainable parameters and lower runtime to obtain more competitive predictive performance than baselines. These results demonstrate that cache-based decoupling provides an effective principle for scalable dynamic graph learning.

Tue 22 SeptMachine Learning
The gist
Dynamic graphs represent things that change over time, like social networks or traffic patterns, and analyzing them takes a lot of computing power. The authors found that current methods waste time recalculating parts of the graph that don't change much between training steps. Their solution, CacheDyG, keeps a special memory bank of past computations and only updates small parts needed for learning. This makes training faster and uses fewer parameters while still making good predictions.
Open → 2609.25814v1

Scalable algorithm finds tight communities around many nodes fast

Improved Methods for k-core Community Search

Abstract: Community search based on user-specified query nodes is complementary to community finding or graph clustering. Prior work in community search is divided into optimizing for external separate- ness or internal cohesiveness, which does not scale well networks of over a billion edges. We present SteinerKCore, a new scalable k-core based community search algorithm for multi-vertex queries. We also present Par-ShellStruct, a parallel algorithm for building the ShellStruct data structure used for k-core community search. We show that our implemen- tations in Icebug, an open-source toolkit for large-scale network analysis, are both more efficient and more scalable than comparative tools, being able to perform on a benchmark network of 273M and 5.1B edges using just 64GB RAM and under 4 hours runtime with 16 CPUs.

Tue 15 SeptSocial and Information Networks
The gist
Finding groups of connected friends or items around specific people in huge networks is hard. Past methods struggled with very large networks or only worked well for certain group shapes. The authors created SteinerKCore, a new method that quickly finds strong, tightly connected groups including many chosen people. They also made Par-ShellStruct, a fast way to prepare necessary network information using multiple computers. Their tools work well on extremely large networks using reasonable memory and time.
Open → 2609.17822v1

Network traffic classification improved with human guided semantic validation

Beyond Measurement Metrics: A Human-Centered Framework for Semantic Validation of Network Traffic Classification

Abstract: Machine learning (ML) has become the dominant approach for network traffic classification, achieving very high predictive performance. However, a model is only valuable if it learns semantically meaningful and trustworthy patterns rather than exploiting spurious correlations. Conventional evaluation practices predominantly assess predictive performance. Consequently, whether the model relies on semantically meaningful patterns remains unknown. To address these challenges, we adapt the knowledge generation framework for network traffic classification. The adapted framework combines data, ML models, explainability, visualization, and expert reasoning to support the iterative exploration, verification, and refinement of model behavior and data preprocessing. The framework is grounded in findings from the literature, benchmark dataset analyses, practical experience with XAI-based traffic classification, and expert feedback, providing practical guidance for semantic model validation. By complementing predictive performance with semantic validation and human expertise, the proposed framework supports the development of network traffic classification models that are not only accurate but also robust and trustworthy.

Tue 15 SeptNetworking and Internet ArchitectureHuman-Computer InteractionMachine Learning
The gist
High accuracy in identifying types of network traffic using machine learning is common, but it is unclear if the models learn meaningful patterns or just tricks. The authors propose a new framework that brings together data, machine learning models, explanations for those models, visual tools, and expert judgment to better understand and improve model behavior. This approach helps ensure that the models work in a trustworthy way and are robust, not just accurate. It encourages combining human feedback with machine performance to make smarter network traffic classification.
Open → 2609.17014v1

Signals improve safety checks for network operation agents

Safety Signals to Verify NetOps Agents with Action-Level Granularity

Abstract: Agentic Network Operations (NetOps) are an emerging paradigm promising to enable workload-aware, self-adjustable, and reliable autonomous networks. While agents have proven their value in incident summarization and telemetry signal extraction, their effectiveness as autonomous control-loop engines heavily relies on their long-horizon reliability. One such setting is the datacenter fabric, where an agent must respond to alarms and operator intents while abstaining from high-risk actions that may cause or extend downtime. Abstention, however, presupposes that an action's impact is known pre-execution, which necessitates a per-action ground truth that NetOps agent benchmarks do not provide. We construct such a ground truth for the network repair task of NetArena. A symbolic replay of the emulated network, validated against the environment at every turn, yields the exact value of every action. From the action-level value, we derive two pre-execution targets, namely whether an action reduces the repair distance (progress) and whether it increases it (harm). We show across 10 agent models, that agent verifiers leveraging internal signals predict both harm and progress more reliably than a baseline using observable signals only. Perspectively, we aim to use these signals as safety feedback to an agent harness to abstain from risky actions and protect the target system.

Sun 13 SeptArtificial Intelligence
The gist
Managing computer networks automatically is tricky because some agent actions can cause problems or downtime. The authors created a way to precisely judge every action an agent might take before it happens, using a detailed replay of network conditions. They showed that by analyzing the agent's internal signals, they can better predict if an action will help fix the network or cause harm. This approach could help agents avoid risky moves and keep networks running smoothly.
Open → 2609.14422v1

Benchmark compares cloud-edge scaling and placement strategies under deadlines

ContinuumBench: Benchmarking Joint Autoscaling and Placement Across Evaluation Regimes in the Cloud-Edge Continuum

Abstract: Cloud-edge controllers coordinate service placement, replica scaling, and resource pre-warming to keep end-to-end latency within application deadlines. But evaluations often obscure the source of a reported gain: placement and scaling are studied separately; workload, connectivity, and calibration assumptions remain implicit; and metrics over completed tasks hide unfinished work. We present ContinuumBench, a benchmark that controls these factors. Its completion-aware accounting treats late, unfinished, and discarded tasks as deadline misses. A common protocol compares placement-only and scale-capable controllers under declared regimes and stressors. Built on the ECLYPSE simulator, ContinuumBench adds arrivals, worker elasticity, intermittent transport, buffering, and failures to close the control loop. We evaluate nine controllers across four scenarios and two regimes. The studied regimes are capacity-bound: elastic capacity, not placement sophistication, drives completion, and once capacity suffices, the choice of autoscaling policy decides how much of that work arrives on time. Placement re-planning has no measurable effect without relocation, while cost-free migration defines the observed exception. Consequently, scale-capable controllers approach an over-provisioned reference while placement-only controllers degrade with load; and placement quality separates controllers only once capacity is exhausted. Finally, the accounting choice itself changes the reported result: completion-only and completion-aware scoring can rank controllers differently.

Tue 8 SeptDistributed, Parallel, and Cluster Computing
The gist
Coordinating where and how much computing power to use in cloud and edge systems is tricky, especially to keep tasks finishing on time. The authors created ContinuumBench, a tool that fairly compares different controllers that decide service placement and scaling, making clear when tasks miss deadlines or fail. They found that scaling capacity often matters more than placement choices until resources are very tight, and different ways of counting task completion can change which controller looks best. This helps clarify how to better manage computing across clouds and edges for time-sensitive applications.
Open → 2609.08946v1

Network fault effects predicted with uncertainty using new model

Information-Entropy-Driven Fault Propagation Modeling for Probabilistic Network Performance Prediction

Abstract: Network faults can trigger cascading effects that cause abrupt and nonstationary performance degradation. Existing learning-based performance predictors mainly focus on normal operation or treat fault-induced topology and routing changes as static inputs, and typically produce deterministic point estimates. They overlook fault-propagation dynamics and uncertainty in performance evolution. The predefined-rule and purely data-driven propagation models lack a unified representation of fault definition, propagation mechanism, and impact quantification. Additionally, generic denoisers in conditional diffusion models fail to incorporate fault propagation into uncertainty modeling. To address these limitations, we propose an information-entropy-driven fault propagation paradigm (IEFP) that characterizes fault propagation via relative entropy, mutual information and transfer entropy. We then design a fault-aware graph message-passing mechanism that propagation contexts modulate network representation learning. We further develop FEMNet, which employs this mechanism as a tailored denoiser within a conditional diffusion model to enable probabilistic network performance prediction under complex fault scenarios. Compared with the strongest baselines, IEFP improves fault-prediction performance, while FEMNet reduces errors in both point and probabilistic KPI prediction.

Tue 8 SeptNetworking and Internet Architecture
The gist
Networks can experience faults that cause complicated and unexpected problems as issues spread through them. Existing prediction methods often treat faults as fixed changes and give only one guess about how the network will perform. The authors present a new way to understand and model how faults spread by measuring the information shared and transferred between parts of the network. They use this idea to build a tool called FEMNet that can predict not only the likely network performance during faults but also the uncertainty around those predictions.
Open → 2609.08143v1