Cybersecurity AI model scores vary greatly with testing methods

Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks

Cryptography and SecurityArtificial IntelligenceComputation and Language

Summary

Cybersecurity AI models are often judged using set tests, but this study shows that the way these tests are run can change a model's score a lot. The researchers looked at different evaluation methods and found that changing how tests are done can move a model’s ranking by many positions. Even tasks that seem the same can give very different results if tested differently. The study suggests that to fairly compare AI models, the testing process itself should be checked and standardized.

Large language modelsBenchmarkingEvaluation pipelineCybersecurityModel rankingMeasurement reliabilityTask semanticsEvaluation harnessSystematic failure modes

Authors

Aymene Berriche, Cathrine Shalby, Mohannad Alhanahnah, Yazan Boshmaf

Abstract

Large language model (LLM) benchmarks are often treated as fixed datasets with stable scores, yet their outcomes depend on configurable evaluation pipelines. We audit eight cybersecurity benchmarks across 10 proprietary, open-weight, and cybersecurity-specialized LLMs. By modeling benchmarks as measurement pipelines, we identify 15 systematic failure modes and show that a single pipeline choice can change a model's score by more than 80 percentage points and substantially alter model rankings. At the cross-benchmark level, two semantically similar task pairs rank the same models differently because of incompatible evaluation conventions. Under an evaluation harness that standardizes pipeline choices while preserving task semantics, nine of 10 models shift by at least three ranks on at least one benchmark. These results show that cybersecurity LLM benchmark scores are pipeline-dependent and motivate pipeline-aware auditing as a core requirement for reliable model evaluation.