Papers for

security teams in tech companies

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

LLM based agents get better defense against hidden prompt attacks

CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents

Abstract: Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based defenses are typically optimized on a static distribution of explicit injections. They learn surface-form cues rather than the boundary between serving the user and obeying an injected objective, and therefore fail when malicious intent is folded into a plausible workflow and deferred for several turns. We present CoDeL, a defense that hardens agent against an attack distribution it reshapes as it trains. The defender is updated each round via LoRA-based GDPO under a decoupled reward over safety, task progress, and format compliance, so refusing injections and completing the user's task jointly define fitness. To keep supplying it with the failures worth learning from, a co-evolving prober searches over injection rounds, attack methods, and payloads for injections that still penetrate the current defender, guided jointly by attack success and attack latency so that it preferentially mines breaches the defender notices too late. Each defender update invalidates part of the attack population and forces the next round onto a new frontier, turning the defender's own failures into a moving curriculum. Extensive experiments on three IPI benchmarks, nine baselines, and two base models show that CoDeL reduces attack success rate (ASR) by 88.5% and outperforms other baselines largely (+38.0%). Codes are available.

Mon 28 SeptCryptography and SecurityArtificial Intelligence
The gist
Large language models (LLMs) used in smart agents can get tricked by hidden instructions that make them do harmful things without clear warnings. The authors introduce CoDeL, a method that trains the agent and a set of attackers together in a cycle so the agent learns from the latest tricky attacks. This approach helps the agent spot and refuse hidden harmful prompts even when they are disguised over multiple steps. Their tests show CoDeL reduces how often attacks succeed by a large margin compared to other methods.
Open → 2609.34463v1

PerceptFence controls sensitive content in screen-share AI assistants

PerceptFence: Content-Mediation Architecture and Deterministic Coverage for Screen-Share AI Assistants

Abstract: Live screen-share AI assistants observe raw screen and speech streams, but users have little runtime control over what an assistant may observe, retain, or disclose. Prompt-level privacy settings are insufficient because sensitive content enters through the capture stream. We present PerceptFence, a content-layer mediation architecture between capture, memory, and model responses, with a deterministic synthetic-fixture scaffold; the artifact omits live capture, category inference, authenticated re-consent, cross-session state, and an external model adapter. On 9,600 protocol-documented adversarial strings scored by a separately implemented exposure oracle, PerceptFence neutralises 0.828 of digit-PII payloads on the 5 seeds both systems run, versus 0.183 for Microsoft Presidio; outside that family Presidio leads 0.238 to 0.154, so the overall 0.398 to 0.260 comparison is only indicative. We then evaluate the path a deployed assistant uses: 480 synthetic developer-support screens rendered by Chrome, degraded, and read by OCR, with rules frozen before testing and three screen types held out. PerceptFence neutralises 889 of 968 OCR-surviving secrets and PII values (0.918; Wilson 95% 0.899-0.934) against 0.581 for Presidio and 0.179 for gitleaks, and 0.974 on the held-out screen types, at a measured cost of 0.763 task-token retention on those types. The contribution is a documented mediation architecture and an evaluation method with explicit coverage boundaries, not a claim of live deployment, formal privacy, novel redaction primitives, or general model robustness.

Sun 27 SeptCryptography and Security
The gist
When people share their screens with AI helpers, private information can unintentionally be seen or kept by the AI. The researchers created PerceptFence, a system that carefully filters what the AI can see and remember from the screen to protect privacy better. They tested it using tricky fake data and found it was better at hiding secrets than other tools they compared it to. This system is not meant for live use yet but shows a clear way to limit what private content an AI assistant can access.
Open → 2609.34027v1