Poisoning malware detectors by exploiting antivirus label weaknesses

Weaponizing Ground Truth: Data Poisoning Attacks by Exploiting Boundary Misalignment Between Antivirus Software and Learning-Based Detectors

Cryptography and Security

Summary

Malware detectors that learn from antivirus labels can be tricked by tiny changes to software that confuse antivirus engines but look similar to the detectors. The authors created a tool called Bi-Iocane that changes such bytes to flip antivirus labels, making malware appear safe or clean software look dangerous. These poisoned examples, when used to train detectors, cause them to misclassify most original software without hurting overall performance. Their tests show that current defenses struggle against this attack and that it poses real risks for systems relying on antivirus-generated labels.

What this means in practice

  • For security software developers: Protect machine-learning malware detectors from training data poisoning that mislabels malware and benign files using antivirus label weaknesses.
  • For enterprise cybersecurity teams: Assess and mitigate risks in malware classification workflows that rely heavily on antivirus-generated labels from services like VirusTotal.

Authors

Jieshuai Yang, Zhi Wang, Yan Jia, Zhenhua Wu, Jianfei Tang, Chenbin Su, Jingwei Ye, Jianwen Tian, Wanpeng Li

Abstract

Machine-learning (ML)-based malware detectors are commonly trained using labels obtained from antivirus (AV) engines and aggregation services (e.g., VirusTotal). This practice assumes AV-generated labels provide reliable supervision. However, small byte-level modifications can substantially alter AV verdicts while leaving the representations perceived by downstream ML detectors largely unchanged, producing label-feature inconsistencies that can contaminate training datasets and create poisoning opportunities for ML-based malware detection. We present Bi-Iocane, a black-box poisoning framework that exploits the reliance of malware-labeling pipelines on AV-generated labels. Bi-Iocane identifies AV-sensitive bytes and modifies them to induce label changes. It rewrites such bytes in malware to obtain benign labels (evasion-oriented poisoning) and injects malware-associated byte patterns into benign software to obtain malicious labels (defamation-oriented poisoning). These poisoned samples and their lightly modified variants corrupt training data and cause selected targets to be misclassified. We evaluate Bi-Iocane with 13 AV engines simulating AV aggregation services and eight ML detectors. For 30 malware and 30 benign clean targets, Bi-Iocane combines AV-specific manipulations to generate malware-to-benign and benign-to-malware poisoned samples whose all tested AV-based labels are flipped. After these poisoned samples and variants are used for downstream training, the resulting ML models misclassify 92.08% of the original clean targets on average with only a 0.06\% poisoning budget per target. Meanwhile, the poisoned models largely preserve clean-set performance, and six evaluated poisoning defenses show only limited mitigation. VirusTotal evaluation further confirms practical defamation risk and reveals potential evasion risk in real-world AV-to-ML labeling supply chains.