Sandbox security flaws in ai orchestration expose high risk exploits

Weird Machine Compositors: Exploiting AI Orchestration at the Expression Layer

Cryptography and Security

Summary

The paper shows that common security measures used to lock down user expressions in AI orchestration platforms are not effective. The authors demonstrate that these sandboxes behave like 'weird machines' that can be manipulated to bypass restrictions. They found multiple serious vulnerabilities in a popular platform called n8n and explain how attackers can move from untrusted inputs to fully trusted access by tricking these systems. They also show that AI helps attackers find these weaknesses faster. The authors suggest a better defense approach that focuses on allowing only known safe operations instead of blocking certain unsafe ones.

What this means in practice

  • For security engineers: Use the new AST coverage tool and methodology to assess and harden expression sandboxes in orchestration platforms before deployment.
  • For devops teams: Incorporate policy inversion strategies to reduce risks of expression-based sandbox exploits in automated workflow systems.

Authors

Eilon Cohen, Ariel Fogel

Abstract

Orchestration platforms secure user-provided expressions through enumerate and block sandboxing: AST rewriting, runtime property blocklists, template sandbox environments. We demonstrate that these sandboxes are weird machines whose instruction set is the underlying language specification, and that the enumerate and block approach is unfixable, following the same trajectory that led to the deprecation of past sandboxing technologies such as Java's SecurityManager and vm2. We validate this claim through three rounds of escalating bypasses against n8n's expression sandbox (three CVEs, two CVSS 9.4, one unauthenticated), and frame these findings within a broader pattern of sandbox failures across the orchestration products category. We identify a trust laundering pattern where orchestration pipelines and applications move attacker controlled input from untrusted to fully credentialed through transformations that strip taint at each level. AI-assisted enumeration accelerates the discovery of these coverage gaps, compressing the timeline between a sandbox's deployment and its compromise. We provide an AST coverage analysis methodology, an accompanying open-source tool, and a defensive playbook that includes policy inversion (allowlist over blocklist) as a structural mitigation.