Continuous assurance improves security checks in software delivery

Continuous Assurance of Agentic Security Auditors for Software Delivery Decision Gates

Cryptography and Security

Summary

Security tools that use large language models help decide if software changes are safe to merge, but their results can change over time and depend on policies. The authors found that checking these tools only once is not enough for reliable decisions. They created a new method that keeps policies separate from the tool’s actions, tracks all evidence versions, and quickly updates security decisions as conditions change. Their system works fast enough to fit within software development timelines and handles many decision points simultaneously.

What this means in practice

  • For software development teams: Keep security audit decisions current by continuously updating assurance when policies or audit evidence change during software integration.
  • For security operations teams: Manage multiple software security decision points reliably by recomputing assurance posture quickly across diverse audit models and environments.

Authors

Guy Lupo, Nguyen Hung Nguyen, Viet Vo, M. A. P. Chamikara, Guangdong Bai, Nazatul Haque Sultan, Alsharif Abuadbba

Abstract

Large language model (LLM)-based repository auditors are increasingly deployed as security controls within continuous integration (CI) pipelines, where their findings admit, block, or delay software changes. As Agentic Software Development Life Cycle (SDLC) Security Controls, their non-deterministic behaviour changes the evidence, while organisational risk appetite and jurisdictional or data-sovereignty policy change its interpretation. Point-in-time audits therefore cannot maintain current assurance for merge decisions. We propose the Policy-Evidence-Execution Separation Pattern, implemented by the Trustworthy AI Posture (TAIP) Assurance Engine and operated as Continuous Control Posture Assurance (CCPA). By separating policy from stable execution and binding admitted evidence to a versioned Posture Tree, the same assurance logic operates across models, environments, and policy profiles. We evaluate the approach using unmodified RepoAudit on a fixed Python Null Pointer Dereference benchmark. The retained evidence repository contains 80 RepoAudit executions across two OpenAI model configurations, gpt-4o-mini and gpt-4.1. TAIP recomputes assurance posture after policy, evidence, and model-context changes and is evaluated across increasing numbers of independent Decision Gateway contexts. The maximum observed policy-to-posture latency was 1.1 ms across three policy-class cycles in one execution. At 1,000 independent assurance contexts, full policy-triggered recomputation with one worker recorded a maximum aggregate refresh of 1.62 s, below the predeclared 5 s Decision Gateway budget. These single-host measurements concern assurance over retained evidence and exclude RepoAudit execution and provider inference.