Tracekit secures autonomous coding agents by detecting tampering

Tracekit: Tamper-Evident Intent-Reasoning-Action Auditing for Autonomous Coding Agents

Computational Engineering, Finance, and ScienceCryptography and Security

Summary

Autonomous coding agents do complex work but their activity logs can be changed, making it hard to trust what they did. The authors present Tracekit, a system that records what the user wants, what the agent thinks, and what it actually executes, all linked securely to prevent tampering. The system finds edits and fakes in these records and helps spot suspicious alterations. It works with multi-agent setups and shows good detection in tests. Tracekit is open source for others to build on.

What this means in practice

  • For software security teams: Securely audit and verify autonomous code agent actions to ensure accountability and detect tampering in software development workflows.
  • For cloud infrastructure operators: Monitor and gate code-executing autonomous agents in multi-tenant cloud environments to prevent unauthorized or harmful operations.
  • For automated compliance auditors: Use tamper-evident ledgers to build verifiable audit trails of autonomous system actions for regulatory or internal review purposes.$Commercial implications: Enables sales of compliance and auditing software products that prove agent behavior integrity to regulated customers.

Authors

Bravish Ghosh

Abstract

Autonomous coding agents read untrusted files, run shell commands and spawn sub-agents with little supervision, yet their record is usually an editable log. We present Tracekit, an open-source, dependency-free system that captures three channels for every agent session: what the human asked (intent), what the model said of its reasoning (self-report), and what it actually executed (actions). These are written to a hash-chained, externally anchorable ledger and cross-checked. Tracekit hooks into Claude Code's lifecycle events, reconstructs multi-agent hierarchies, gates tool calls with a pre-execution policy, accepts events from other agents via an SDK or HTTP API, and renders a live observer that re-verifies the ledger in the browser. We evaluate Tracekit in five experiments. (1) Across 1,600 random mutations, the chain detects every edit, deletion, reordering, forged insertion and torn write; tail truncation and full re-chaining are caught only by anchors, with detection falling to 0.47 at an anchoring interval of 300 records, matching a closed-form model. (2) A hook costs 23.9 ms median, flat up to 100,000 ledger records, and the chain stays correct under 16 concurrent writers. (3) A regular-expression gate blocks only 18 of 44 harmful tool calls (41%) while wrongly blocking 3 of 40 benign ones; trivial rewrites evade it. (4) In 14 real Claude Code runs, the agent never acted on four planted indirect prompt injections and disclosed each. The provider withheld the text of all 27 thinking blocks, so self-report was limited to visible prose. (5) With seeded-fault splicing, a new method that inserts concealed misaligned steps into real traces, over 63 traces and 126 reviewer calls, rule flags caught 40/49 (82%) of faulted traces and an independent LLM reviewer caught 98/98 (100%), with a false-positive rate of 1/14 on unmodified traces. We release the system, harness and all traces.