AI trust layer helps check flight plans for safety and reliability

Toward a Decision-Assurance Layer for AI-Assisted Flight Planning in Air Traffic Management

Artificial Intelligence

Summary

Using AI to help plan flights can speed things up, but AI sometimes gives unexpected or unsafe answers. The authors propose a system called ATAL that checks AI flight plans for reliability using different safety checks. ATAL gives a clear signal about how safe the AI output is, so human planners can decide whether to trust it. They tested the system with flight planning tasks and showed it can spot problematic AI suggestions before they cause issues. This approach can also work in other areas where humans must oversee AI decisions for safety.

What this means in practice

  • For air traffic controllers: Use ATAL to evaluate AI-generated flight plans for safety before approving them in daily operations.
  • For rail traffic planners: Adapt the ATAL framework to verify AI-assisted route planning for trains ensuring compliance with operational constraints.

Authors

Alexandre Barreto, Shou Matsumoto, Jorge Valverde-Rebaza, Cleiton Ataide, Paulo Costa

Abstract

Generative AI is increasingly being used informally in Air Traffic Management (ATM) for tasks such as flight plan generation, trajectory interpretation, and constraint checking. Although these tools can reduce workload and accelerate planning, their non-deterministic outputs create safety and operational risks in human-in-the-loop settings. This paper proposes the AI Trust and Assurance Layer (ATAL), a model-agnostic decision assurance architecture that evaluates whether AI-generated flight-planning outputs are sufficiently reliable for operational use. ATAL combines semantic stability under prompt variation, operational consistency of structured outputs, and normative constraint validation against domain rules, and maps these signals to a Decision Readiness Level (DRL) for human operators. An ATM-inspired experimental study shows how unsafe, inconsistent, or misleading outputs can be identified before influencing flight-plan validation or execution. Although demonstrated in aviation, the framework is also transferable to other safety-critical decision-support domains that require human oversight under regulatory constraints.