Network traffic classification improved with human guided semantic validation
Beyond Measurement Metrics: A Human-Centered Framework for Semantic Validation of Network Traffic Classification
Networking and Internet ArchitectureHuman-Computer InteractionMachine Learning
Summary
High accuracy in identifying types of network traffic using machine learning is common, but it is unclear if the models learn meaningful patterns or just tricks. The authors propose a new framework that brings together data, machine learning models, explanations for those models, visual tools, and expert judgment to better understand and improve model behavior. This approach helps ensure that the models work in a trustworthy way and are robust, not just accurate. It encourages combining human feedback with machine performance to make smarter network traffic classification.
What this means in practice
- •For network operations teams: Improve network traffic monitoring by verifying that classification models base decisions on meaningful patterns rather than chance correlations.
- •For cybersecurity analysts: Enhance detection of network threats by using semantically validated models that provide trustworthy traffic classifications.
Authors
Igor Cherepanov, David Sessler, Alex Ulmer, Thorsten May, Jörn Kohlhammer
Abstract
Machine learning (ML) has become the dominant approach for network traffic classification, achieving very high predictive performance. However, a model is only valuable if it learns semantically meaningful and trustworthy patterns rather than exploiting spurious correlations. Conventional evaluation practices predominantly assess predictive performance. Consequently, whether the model relies on semantically meaningful patterns remains unknown. To address these challenges, we adapt the knowledge generation framework for network traffic classification. The adapted framework combines data, ML models, explainability, visualization, and expert reasoning to support the iterative exploration, verification, and refinement of model behavior and data preprocessing. The framework is grounded in findings from the literature, benchmark dataset analyses, practical experience with XAI-based traffic classification, and expert feedback, providing practical guidance for semantic model validation. By complementing predictive performance with semantic validation and human expertise, the proposed framework supports the development of network traffic classification models that are not only accurate but also robust and trustworthy.