How adversarial signals weaken multi-agent trading with large language models

Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems

Artificial Intelligence

Summary

The paper looks at how trading systems that use multiple AI agents and large language models can be tricked by harmful social media posts. The researchers created a system to simulate attacks that only add normal-looking posts, which can still change the trading team’s beliefs and decisions. They measure how bad signals spread inside the system and show that even simple attacks can hurt how well these AIs trade, making them less profitable and more risky. However, the study also finds that thoughtful design choices can make these systems less vulnerable to such attacks.

What this means in practice

  • For quantitative finance teams: Detect potential vulnerabilities in AI-driven trading systems to social media manipulation and improve system resilience.
  • For ai system architects: Design multi-agent AI frameworks with coordination mechanisms that reduce the impact of adversarial information injection.

Authors

Qi Rong Sua, Junhao Dong, Nguyen Duc Thai, Yuqing Wen, Cheston Tan, Yew-Soon Ong

Abstract

Multi-agent trading systems built on large language models (LLMs) are beginning to appear in quantitative finance, yet their robustness to adversarial inputs is largely unknown. We study the vulnerability of LLM trading stacks to black-box, input-only attacks that enter solely via admissible social-media feeds. We introduce the Generic Multi-Agent Trading System (GMATS), a framework that captures modern multiagent trading architectures and instantiate a class of black-box poisoning attackers that treat an LLM as a post generator and inject budget-constrained, plausibly benign social-media content into the analyst's evidence stream. We define contagion metrics that trace how adversarial content propagates through the stack, including belief-shift scores at analyst and coordinator layers and attack-clean deltas on standard backtest metrics. Experiments on a safe offline benchmark with historical market and social data show that even simple input-only attackers can materially degrade risk-return profiles, sharply reducing Sharpe ratios. At the same time, we find that suitably designed multi-agent topologies and coordinator prompts can dampen adversarial shocks and improve average robustness under identical poisoning budgets.