Knowledge distillation shapes detection habits in encrypted traffic classifiers

Unknown-Traffic Detection, Calibration and Shortcut Reliance in Distilled Encrypted-Traffic Classifiers over One Year

Networking and Internet ArchitectureMachine Learning

Summary

Classifiers compress complex models to run on smaller devices using knowledge distillation, but usually just check accuracy. This paper finds that students (simpler models) inherit some traits from their teachers, like confidence or how they detect unknown traffic, but not consistently or fully. The study showed that model size impacts shortcut reliance more than distillation does, and some detection skills can be obtained without a teacher. The authors tested these effects over a year of real encrypted network traffic and found mixed results.

What this means in practice

  • For network security teams: Evaluate and calibrate compressed encrypted-traffic classifiers considering inherited detection traits and confidence shifts over time.
  • For edge device engineers: Design smaller encrypted traffic classifiers informed by how shortcut reliance and unknown-traffic detection persist after distillation.

Authors

Mahmoud Abbasi

Abstract

Knowledge distillation is the standard way to compress encrypted-traffic classifiers for the edge, and almost all such work judges students by accuracy alone. We ask what else a student inherits: unknown-traffic detection, calibration, shortcut reliance, and whether any survives a year of drift. Resemblance proves little on its own, since soft targets also regularise. We therefore distil one 101k-parameter student from two teachers of equal accuracy but different construction, a five-member ensemble and a single wider model, so that following one rather than the other is attributable to it. The design was pre-registered before any test result was seen. We tested ten hypotheses on CESNET-TLS-Year22, a year of real TLS traffic, across 18 test windows over 35 weeks. Two are supported: a student's per-flow unknown-scores shift toward its own teacher, but only at a conventional temperature, not the accuracy-optimal one; and a shortcut-reliant teacher passes its over-confidence to a student that never sees the feature. The drift prediction is reversed under both scores, the gap narrowing rather than widening and the student overtaking under the energy score in two of three replicates, as is the prediction that such a teacher harms its student's detection, which improves slightly. Shortcut reliance is set by model size, not distillation. Under the logit-based scores nothing else transfers: distillation beats neither a temperature-scaled direct student nor label smoothing. Exploratory analysis shows this turns on the scoring rule: with a feature-space detector the teacher detects unknown traffic 0.073 AUROC better than the direct student, where the energy score sees 0.000, and the conventional-temperature student inherits most of it. Label smoothing, with no teacher, recovers more. Distillation transfers the teacher's habits; what looks like an inherited ability is available without one.