Papers for
security teams in tech companies
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
LLM based agents get better defense against hidden prompt attacks
CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents
Abstract: Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based defenses are typically optimized on a static distribution of explicit injections. They learn surface-form cues rather than the boundary between serving the user and obeying an injected objective, and therefore fail when malicious intent is folded into a plausible workflow and deferred for several turns. We present CoDeL, a defense that hardens agent against an attack distribution it reshapes as it trains. The defender is updated each round via LoRA-based GDPO under a decoupled reward over safety, task progress, and format compliance, so refusing injections and completing the user's task jointly define fitness. To keep supplying it with the failures worth learning from, a co-evolving prober searches over injection rounds, attack methods, and payloads for injections that still penetrate the current defender, guided jointly by attack success and attack latency so that it preferentially mines breaches the defender notices too late. Each defender update invalidates part of the attack population and forces the next round onto a new frontier, turning the defender's own failures into a moving curriculum. Extensive experiments on three IPI benchmarks, nine baselines, and two base models show that CoDeL reduces attack success rate (ASR) by 88.5% and outperforms other baselines largely (+38.0%). Codes are available.
PerceptFence controls sensitive content in screen-share AI assistants
PerceptFence: Content-Mediation Architecture and Deterministic Coverage for Screen-Share AI Assistants
Abstract: Live screen-share AI assistants observe raw screen and speech streams, but users have little runtime control over what an assistant may observe, retain, or disclose. Prompt-level privacy settings are insufficient because sensitive content enters through the capture stream. We present PerceptFence, a content-layer mediation architecture between capture, memory, and model responses, with a deterministic synthetic-fixture scaffold; the artifact omits live capture, category inference, authenticated re-consent, cross-session state, and an external model adapter. On 9,600 protocol-documented adversarial strings scored by a separately implemented exposure oracle, PerceptFence neutralises 0.828 of digit-PII payloads on the 5 seeds both systems run, versus 0.183 for Microsoft Presidio; outside that family Presidio leads 0.238 to 0.154, so the overall 0.398 to 0.260 comparison is only indicative. We then evaluate the path a deployed assistant uses: 480 synthetic developer-support screens rendered by Chrome, degraded, and read by OCR, with rules frozen before testing and three screen types held out. PerceptFence neutralises 889 of 968 OCR-surviving secrets and PII values (0.918; Wilson 95% 0.899-0.934) against 0.581 for Presidio and 0.179 for gitleaks, and 0.974 on the held-out screen types, at a measured cost of 0.763 task-token retention on those types. The contribution is a documented mediation architecture and an evaluation method with explicit coverage boundaries, not a claim of live deployment, formal privacy, novel redaction primitives, or general model robustness.