Papers for

enterprise risk managers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Language model agents defend enterprise boards against deception attempts

DGF-Bench: A Benchmark for Simulating and Auditing Deception Against Multi-Agent Governance Boards

Abstract: Tool-using language-model agents can review enterprise projects as governance boards do: they read the evidence, apply written rules and decide whether the project may proceed. Part of that evidence comes from suppliers and project members with a stake in the decision. DGF-Bench is a benchmark in which a board of agents (specialist gates and a General gate that consolidates their decisions) reviews synthetic dossiers while an attacker plants deceptive content in evidence the organization does not vouch for. Dossiers are generated from canonical facts under 61 executable rules, with 42 authoritative records and 32 narrative documents; every gate is certified decidable from those records. Attacks never change an authoritative value, so an attacked dossier keeps the reference decisions of its clean copy. A success is attributable only when the agent receives the injection and takes the exact injected action, which it does not take on the paired clean dossier; the DGF score is the share of applicable fixed attacks a model blocks. Reading documents and records themselves, five of six models were outcome-strict (disposition, findings, actions and authorization all correct) on 82 to 85 of 85 gates. Over 2,622 attacked gate runs, seven direct-order, false-data and false-authority attacks obtained one attributable success against these five, whereas task-aligned attacks imitating the organization's own process passed against four of them: a record note citing a fake review procedure lowered GPT-6 Luna Pro from 34 to 6 outcome-strict gates and DeepSeek V4 Pro from 33 to 7. DGF scores ranged from 96.2 to 26.9, and a policy-aware adaptive attacker writing in records succeeded against five of six models. The approval tool executed no forged approval, yet deceived agents submitted approvals that the rules forbid. The open-source package dgf-bench computes the DGF score with one command.

Mon 28 SeptArtificial IntelligenceCryptography and Security
The gist
Sometimes groups using AI agents to check projects get tricked by false information hidden inside project files. The researchers created a test called DGF-Bench where AI agents act like a company's review board and try to spot lying or fake data planted by attackers. They measured how well different AI systems could avoid being fooled by these tricks and found that some attacks were very effective when they copied the company’s own rules falsely. The work helps understand how trustworthy these AI reviewers are in spotting deception in complex decision processes.
Open → 2609.34913v1

Provenance based runtime guard stops cascading attacks on llm agents

AGATE: Provenance-Based Runtime Defense Against Compositional Attacks on LLM Agents

Abstract: LLM agents can produce harmful effects through sequences of ordinary operations. Judging such actions requires establishing both the authority that permits them and the origin of the data they carry. We present AGATE, an authorization and data-provenance gate at instrumented agent-harness boundaries. Operator declarations and host approval events ground authorization; delegated actions are constrained by grants that bind to exact parameters, expire, and permit a limited number of uses. Source registration connects observed inputs to subsequent transfers, while an effect ledger tracks repeated requests. Deterministic checks make decisions without an LLM in the decision path and retain their grounds with execution evidence for forensic replay. Adapters integrate three production harnesses -- DeepSeek Harness, OpenCode, and OpenClaw -- without modifying host code, translating each host's native observation and veto points into a single shared gate interface; the judgment core is identical in all three, and only enforcement depth differs. Our evaluation combines 153 exercised attack-chain records with deployment, utility, and reconstruction experiments. The deployment observations expose how tool declarations and data checks govern business actions, including a bypass through parameter rewriting. Six of eleven benign file-processing scenarios contain denial events, revealing the utility cost of content-based provenance policies. Across 252 runs on 63 sanitized scenarios, replay agrees with live graph projections for all 63 scenarios on each of two platforms. These results establish the feasibility of provenance-based runtime judgment and identify content transformation, legitimate reuse, and observation coverage as concrete limits.

Fri 25 SeptCryptography and SecuritySoftware Engineering
The gist
Large language model (LLM) agents can cause harm by chaining normal actions that look safe on their own. The authors created AGATE, a tool that checks who allowed each action and where the data comes from to decide if the action is okay. It watches agent boundaries and records detailed evidence so decisions can be replayed for review. AGATE works with three existing LLM frameworks without changing their code and was tested on many attack and normal scenarios to verify its accuracy and limits.
Open → 2609.30830v1