Papers for

security teams in ai companies

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

TraceGuard detects poisoned data in image text training datasets

TraceGuard: Adaptive Multimodal Poison Filtering through Cross-Feature Rank Agreement

Abstract: Multimodal training relies on image-text corpora collected from external sources, creating opportunities for attackers to poison the data. Stealthy attacks can preserve plausible image-text pairs while concealing the differences used by detectors, so apparently clean data can still redirect the trained model. We therefore ask which properties a poison set must preserve for the attack to remain effective. A small poison set must still exert enough collective influence during training to induce the attacker's target behavior. We analyze this influence in terms of how often an attack pattern occurs and how strongly the examples carrying it jointly affect the model. This analysis motivates six corpus-level features that examine cross-modal neighborhoods, recurring text, and changes after text-span erasure without training the victim model. We introduce TraceGuard, an adaptive rank-based filtering method that uses agreement among complementary feature rankings to identify suspicious examples. It refines the selected set through shared patterns and adapts the removal threshold to each corpus without knowing the attack or poison rate. Across 19 attack configurations spanning image-text learning, generative vision-language model fine-tuning, and encoder-transfer tests, TraceGuard removes an average of 98.4% of poisoned examples and 5.4% of clean examples. After training on the filtered corpora, the residual attack metric is at most 1% in 13 configurations. Matched-removal controls and ablations support the contributions of sample selection and adaptive removal. Stress tests also identify detection failures under adaptive attacks and unnecessary removal on poison-free corpora.

Thu 24 SeptCryptography and SecurityMachine Learning
The gist
Training AI models with images and text from the internet can be risky because attackers can hide harmful examples that trick the model. The authors study what makes a small set of bad data powerful enough to affect training. They develop TraceGuard, a tool that spots suspicious examples by looking for agreement among different patterns in the data without needing to train the model first. Their tests show TraceGuard removes most bad examples while keeping most good ones, making AI models safer to train.
Open → 2609.29099v1

Language models balance fast output and traceable text with new sampling method

Watermarkable Multi-Draft Speculative Sampling via Poisson Processes

Abstract: Large language models (LLMs) have achieved state-of-the-art performance across a wide range of tasks, motivating two important aspects of deployment: inference efficiency and output provenance, which can be tackled by speculative sampling and watermarking, respectively. However, recent works have shown that combining these two goals is highly nontrivial and can be potentially impossible. In this work, we develop a novel multi-draft speculative sampling algorithm based on Poisson processes that improves the frontier of this fundamental trade-off. The proposed algorithm has strong sampling efficiency on its own and, more interestingly, is naturally watermarkable: we can embed an unbiased watermark without degrading speculative acceptance. Moreover, our algorithm is based on an exact list-coupling-without-communication scheme, which yields a drafter invariance property that benefits both sampling and watermarking. It is the first multi-draft, drafter-invariant speculative sampling scheme that maintains both watermark strength and sampling efficiency, and we experimentally verify its strong performance in both aspects.

Fri 18 SeptCryptography and SecurityMachine Learning
The gist
Large language models can create text quickly and let people know that the text came from them. But doing both at the same time is hard. The authors designed a new way to pick words using a process based on random timing, which helps keep the text fast and easy to check. Their method also makes sure the text carries a special signature (a watermark) without slowing things down. They tested their idea and found it works well for both speed and watermark quality.
Open → 2609.21858v1