From Runtime Records to Legal Findings: An Evidentiary-Adequacy Criterion for Agentic AI Oversight

2026-07-01Computers and Society

Computers and Society
AI summary

The authors explain that just having logs or records from AI systems isn't enough to prove certain legal facts, like whether data crossed a boundary or if a human could intervene. They define a clear rule saying these records need to connect events to legal categories and show relationships like timing or authority to be useful for those decisions. They tested this rule on some European AI laws and found that common tools like tamper-proof logs or generic frameworks aren't enough by themselves. Their work ties into ideas about how good systems need to capture enough detail to be trustworthy.

agentic AI systemsruntime recordsevidentiary adequacybinary findings of factEU AI Actprovenanceaudit logsGood Regulator Theoremruntime verification
Authors
Jeroen Janssen
Abstract
Agentic AI systems generate runtime records, logs, traces, and audit artefacts, but the existence or integrity of such records does not by itself establish that legally operative oversight findings can be recovered from them. This technical report defines an evidentiary-adequacy criterion for a bounded class of determinations: binary findings of fact about specific events and their relations, such as whether protected data crossed a boundary, whether a human could intervene, whether an information barrier held, or whether delegated authority was valid at the moment of use. The criterion states that a runtime record can answer such a determination only if it carries both a typing that maps recorded events to the legally operative category and the relation, such as provenance, authority, derivation, or temporal validity, on which the determination's truth depends. The claim is one of necessity, not sufficiency. The report instantiates the criterion against selected EU AI Act oversight obligations and explains why tamper-proof logs, generic process frameworks, and provenance structures alone cannot establish the relevant findings. It further relates the argument to requisite variety, the Good Regulator Theorem, and the trace-versus-hyperproperty boundary of runtime verification. Companion materials and the experiment protocol are archived on Zenodo.