GraftyVul: Synthesising Insecure Programs Through Real-World Vulnerability Grafting
Cryptography and SecuritySoftware Engineering
Summary
The authors created GraftyVul, a new dataset that embeds real-world security bugs into open-source programs to make realistic and testable vulnerable code. This dataset covers five programming languages and many types of vulnerabilities, and each vulnerability is verified to actually affect the program. They also developed a new way to compare vulnerabilities that works across languages and setups, showing their dataset closely matches real bugs. Compared to other datasets, GraftyVul offers a good balance of variety, realism, and reproducibility. They demonstrated its usefulness by testing a real-world security fix system.
Authors
Omri Ram, Mitchell Horner, Ron Van der Meyden, Alsharif Abuadbba, Hammond Pearce
Abstract
Vulnerability datasets underpin a wide range of security research, including vulnerability detection, automated remediation, and secure code generation. However, existing datasets sacrifice at least one of three desirable properties: diversity (of language or vulnerability type), reproducibility/executability, or realism. We therefore present GraftyVul, a system that constructs vulnerable programs by grafting real-world vulnerabilities into open-source projects. This grounds the dataset in vulnerabilities observed in real-world contexts while harnessing known good build and test environments, enabling exploit-verification scripts to guarantee that an introduced vulnerability successfully alters a program's behaviour. Using GraftyVul, we generate 212 verified and exploitable vulnerable programs spanning five programming languages (Python, TypeScript, Java, Go, and C#) across 23 CWE categories. To evaluate fidelity, we introduce a language- and context-agnostic semantic embedding that compares vulnerabilities by sink, mechanism and host-feature rather than surface code. This approach outperforms standard code embeddings on cross-language clone and CWE classification. These embeddings demonstrate that GraftyVul samples retain a strong semantic signature to their source vulnerability. We additionally compare GraftyVul against 13 widely used datasets, where it attains competitive diversity while being the only reproducible-exploit dataset with broad language and CWE coverage. Finally, we illustrate GraftyVul's practical utility through an industrial case study evaluating a production vulnerability remediation system.