Human oversight needs anticipation for planning ai systems

Anticipatory Human Oversight of Agentic AI: A Philosophical Account

Computers and Society

Summary

Current ways people watch AI can work well when AI makes quick decisions, but they fall short with AI that plans and acts over time. The authors explain that for smarter AI that carries out multi-step tasks, oversight has to happen before the AI acts by setting clear rules and goals. This anticipatory oversight works alongside watching what AI does reactively, making sure harmful patterns are caught early. They also discuss how this approach helps clarify who is responsible when things go wrong and addresses common concerns about control and fairness.

What this means in practice

  • For ai safety teams: Design oversight strategies that specify AI goals and limits before execution to better prevent long-term risks from multi-step AI actions.
  • For government regulators: Develop regulations that require anticipatory human oversight frameworks to ensure accountability and proactive control over autonomous AI agents.

A position paper. It proposes an approach and reports no results.

Authors

Kevin Baum, Maximilian Kiener, Markus Langer, Johann Laux

Abstract

Human oversight is widely held to mitigate the risks of AI systems. Even for systems that produce discrete outputs at identifiable decision points, the realisation of human oversight as a reactive measure is empirically fragile, yet increasingly well understood. However, for agentic AI -- systems that plan, decompose goals, and execute multi-step actions over extended horizons -- reactive oversight reaches its structural limits: intervention on individual actions defeats the autonomy that motivates the deployment, while intervention on aggregate patterns is too coarse for harms whose cumulative consequences only become legible after the fact. This paper argues that reactive oversight must be complemented by an anticipatory mode: oversight exercised before the agent acts, by specifying the normative agenda that structures the space of permissible action and refining it iteratively through specification, runtime, and inspection. The two are complements -- the agenda's escalation conditions specify when reactive intervention is invoked. Drawing on Meaningful Human Control, we read anticipatory oversight as the operationalisation of distal-reason tracking. In addition, we argue that the proposed framework yields a specific responsibility architecture by design: occupying the anticipatory mode is the discharge of a role-grounded prospective obligation, and backward-looking responsibility takes the form of strict moral answerability -- rationalistic, relational, and holding regardless of fault, in virtue of the principal's prior opportunity for precaution. We develop bridging failure modes, address objections including moral luck and the illusion of control, and close with regulatory, architectural, and empirical implications