JevVibe improves secure code generation with fast weak spot detection
JevVibe: Efficient Classification-Guided Secure Code Generation
Cryptography and SecurityArtificial Intelligence
Summary
Code-writing AI sometimes creates programs that work but have security flaws. The authors study Jev, a tool that quickly and accurately finds known types of security problems in code. They combine Jev with a fixer tool to create JevVibe, which uses the identified problem types to make the code safer. Their tests show JevVibe raises the security pass rate notably, doing better and cheaper diagnosis than other AI methods.
What this means in practice
- •For software development teams: Improve the security of AI-generated code by diagnosing vulnerability types quickly for targeted fixing.
- •For cybersecurity operations teams: Use fast vulnerability classification to prioritize and guide swift code security repairs in development pipelines.
Authors
Arshak Rezvani, Sasha Behrouzi, Ahmad-Reza Sadeghi
Abstract
Large language models can generate functionally correct code that still contains security weaknesses, motivating repair pipelines that first diagnose a weakness type before deciding how to fix it. The Common Weakness Enumeration (CWE) provides a standardized vocabulary for such diagnoses, but asking an autoregressive language model to generate a CWE label and extracting it from the response raises questions about output validity, speed, and cost, as well as accuracy. We evaluate Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, GPT-5.6-Sol, on a controlled 50-way CWE classification task over 1,916 CyberSecEval benchmark examples. Jev outperforms all six open-weight baselines on every classification and ranking metric, while its comparison with GPT-5.6-Sol depends on the metric: GPT-5.6-Sol achieves higher Top-1 accuracy and Macro-F1, whereas Jev achieves higher Top-3 and Top-5 accuracy and a nearly identical MRR, at $6.27\times$ lower median API latency and $55.9\times$ lower estimated API cost. We further build JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code generated by Qwen2.5-Coder-32B-Instruct. With Jev providing the diagnosis, the agent increases the detector-measured security pass rate from 63.5% before repair to 70.7%, compared with 66.1% for LLM-guided repair. These results show that JevVibe is effective at improving the security of generated code, with Jev providing reliable and efficient CWE classification.