AI trading methods perform worse in real crypto markets than backtests suggest
Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets
Artificial Intelligence
Summary
It’s hard to know if AI programs that trade cryptocurrencies will actually make money when used live because most tests use old data. The authors created a way to test these AI trading methods step-by-step—from looking at past data, to pretending to trade in the future, to actually trading with real money. Their system shows how much the AI’s performance drops when moving from simulations to real market conditions. They also compare different kinds of AI approaches to see which handle real trading better.
What this means in practice
- •For crypto trading teams: Test AI trading strategies in real future markets using a unified system that moves from simulation to live trading while measuring actual performance drops.
- •For financial software developers: Build or improve AI-driven trading platforms for crypto by integrating a standardized benchmark that evaluates models in increasingly realistic conditions.
Authors
Xingtong Yu, Jiarun Zhou, Guanlin Ding, Wenkang Wei, Jiarui Liu, Chang Zhou, Fangzhou Ge, Chenyi Xu, Xikun Zhang, Renqiang Luo, Jie Zhang, Hong Cheng, Xinming Zhang, Hui Zhang, Yuan Fang
Abstract
AI-based trading methods have rapidly evolved from machine learning and reinforcement learning to large language models (LLMs) and trading agents, yet their performance is still predominantly assessed through historical backtesting. Such evaluations provide limited evidence of whether a method can generalize to unseen future markets or whether its backtested performance can be sustained in realistic trading frictions (e.g., latency, slippage, liquidity constraints, and market impact). We present a unified benchmark that evaluates representative machine learning, reinforcement learning, LLM-based, and agent-based trading methods in cryptocurrency markets through three progressively more realistic stages: historical backtesting, prospective exchange-based paper trading, and real-money live trading. These stages jointly increase temporal realism by moving from historical to unseen future markets, and execution realism by moving from offline simulation toward live trading. This protocol enables us to quantify the backtest-to-realization gap, identify when performance begins to deteriorate, and compare how this gap differs across major classes of AI trading methods. We further provide a unified open-source system supporting all three evaluation stages, together with a public platform that continuously updates benchmark results. Code is available at https://github.com/Starlien95/Awesome-TradingAI.