Papers for

quantitative finance teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Deep learning improves uncertainty and explanation in option pricing models

Uncertainty and Explainability in Deep Rough Volatility: A Neural Information-Theoretic Posterior Approach

Abstract: Deep learning has substantially accelerated the calibration of complex stochastic-volatility models, but neural point calibration alone does not capture the uncertainty remaining after an implied-volatility (IV) surface has been observed. We develop a simulation-based inference framework for rough Heston (rHeston) calibration that learns the posterior distribution of the model parameters conditional on an IV surface. Using neural ratio estimation, we obtain calibrated posterior samples that can be propagated through heteroscedastic neural surrogate pricers for path-dependent exotic options. The resulting posterior-predictive distributions combine residual parameter uncertainty with conditional surrogate uncertainty and yield uncertainty-aware price intervals. We further introduce Hellinger-SHAP, an information-theoretic explainability method for posterior inference. Rather than attributing a single parameter point estimate, it applies local-background Kernel SHAP to a posterior-information functional measuring contraction from the prior to the posterior. This identifies maturity--moneyness regions associated with posterior information gain for individual rHeston parameters. In a simulation study, posterior-predictive intervals provide calibrated or conservative coverage across forward-start, barrier, and realized-variance claims, while point plug-in prices can be materially unreliable for selected contract regimes. Together, the UQ and XAI analyses provide a transparent framework for uncertainty-aware neural calibration and downstream exotic pricing under the specified prior-predictive model.

Fri 25 SeptMachine Learning
The gist
Financial markets often use complex mathematical models to price options, but it's hard to know how uncertain these prices are after using observed market data. The authors develop a new deep learning approach that not only estimates possible values of the model parameters but also quantifies the uncertainty of those estimates. They also introduce a method to explain which parts of the market data influence each parameter estimate. Their approach provides more reliable price estimates and clear insights about uncertainty for different types of financial contracts.
Open → 2609.31570v1

Online method adapts portfolio window size to reduce trading costs

Cost-Sensitive Online Window Size Selection for Portfolio Management

Abstract: This paper investigates cost-sensitive online window size selection for portfolio management under changing market conditions. Specifically, we propose a two-level framework that constructs portfolios using candidate window sizes and dynamically aggregates them through online learning. By treating candidate window sizes as ``experts,'' we dynamically update their aggregation weights using turnover-inclusive losses. Moreover, we derive finite-horizon cost-sensitive tracking-regret bounds that account for turnover of the aggregated portfolio, with static regret as a special case. Under bounded losses and cost rates, suitably tuned Fixed Share achieves asymptotically no tracking regret for sublinear switching budgets, with Hedge covering the static case.

Thu 24 SeptMachine Learning
The gist
Choosing how far back in time to look at stock prices matters for investment strategies, especially when markets change. This paper shows a new way to automatically pick and combine different time windows so portfolios adapt better over time without excessive trading costs. The authors treat each window choice as an expert and update their influence based on both performance and trading costs. Their approach comes with mathematical guarantees on how well it tracks the best switching strategy under realistic cost conditions.
Open → 2609.29887v1

AlphaDiverse improves stock factor research with diverse local AI agents

AlphaDiverse: Post-Training Local Quantitative Research Agents for Diverse Exploration in Alpha Factor Mining

Abstract: Large language model (LLM)-based multi-agent systems can automate alpha factor mining, but their reliance on external APIs limits control over cost, availability, and confidentiality. Long research loops also tend to revisit a few successful economic mechanisms that lead to research path collapse. To address these limitations, we propose AlphaDiverse, a framework that integrates a multi-agent alpha research system, diverse research path collection, and post-training for local agents. We let the research system generate complementary plan portfolios and vary research environments across loops to collect diverse research paths. Using these diverse traces, we warm-start local Planner and Realizer agents with supervised fine-tuning. Then, we propose a joint GRPO method to optimize both of them using predictive quality and diversity of contributions. Research feedback is confined to inner period data, while a frozen final model is evaluated on a later outer period data, thereby avoiding test-set tuning. Experiments across four Chinese stock universes show that AlphaDiverse can combine competitive prediction with broader exploration.

Thu 24 SeptArtificial IntelligenceComputational Engineering, Finance, and ScienceMultiagent Systems
The gist
Finding new strategies to predict stock prices often relies on repeated ideas, limiting success. The authors created AlphaDiverse, a system where multiple AI agents generate varied research paths and learn from them locally, rather than relying on external tools. This approach encourages exploring a wider range of possibilities and reduces repeated focus on a few popular strategies. Their experiments with Chinese stocks show that AlphaDiverse balances good predictions with more diverse research outcomes.
Open → 2609.29014v1

How adversarial signals weaken multi-agent trading with large language models

Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems

Abstract: Multi-agent trading systems built on large language models (LLMs) are beginning to appear in quantitative finance, yet their robustness to adversarial inputs is largely unknown. We study the vulnerability of LLM trading stacks to black-box, input-only attacks that enter solely via admissible social-media feeds. We introduce the Generic Multi-Agent Trading System (GMATS), a framework that captures modern multiagent trading architectures and instantiate a class of black-box poisoning attackers that treat an LLM as a post generator and inject budget-constrained, plausibly benign social-media content into the analyst's evidence stream. We define contagion metrics that trace how adversarial content propagates through the stack, including belief-shift scores at analyst and coordinator layers and attack-clean deltas on standard backtest metrics. Experiments on a safe offline benchmark with historical market and social data show that even simple input-only attackers can materially degrade risk-return profiles, sharply reducing Sharpe ratios. At the same time, we find that suitably designed multi-agent topologies and coordinator prompts can dampen adversarial shocks and improve average robustness under identical poisoning budgets.

Thu 17 SeptArtificial Intelligence
The gist
The paper looks at how trading systems that use multiple AI agents and large language models can be tricked by harmful social media posts. The researchers created a system to simulate attacks that only add normal-looking posts, which can still change the trading team’s beliefs and decisions. They measure how bad signals spread inside the system and show that even simple attacks can hurt how well these AIs trade, making them less profitable and more risky. However, the study also finds that thoughtful design choices can make these systems less vulnerable to such attacks.
Open → 2609.19789v1

AlphaRJM improves formulaic alpha discovery using reward-jump memory

AlphaRJM: Reward-Jump Memory for Stochastic Return-Guided Alpha Discovery

Abstract: Formulaic alpha discovery is a pool-dependent symbolic search problem in which informative feedback is observed primarily when a complete expression is evaluated. This delayed feedback creates two coupled difficulties: the retained alpha pool does not preserve the full history of realized evaluation feedback, and the value of an intermediate construction action is uncertain because its consequence depends on the formula eventually completed. We introduce AlphaRJM, which addresses these difficulties through Reward-Jump Memory, an event-driven latent state that remains fixed during token construction and updates only at terminal evaluation events using the realized pool reward and evaluation outcome, and an action-conditioned SDE return critic that represents future discounted discovery returns with stochastic particles. The particles guide action selection through their mean and uncertainty and are learned using a distributional Bellman objective combining energy-distance matching, mean calibration, and jump regularization. Empirically, AlphaRJM delivers strong and stable gains across multiple equity universes, forecasting horizons, and random seeds, while ablations confirm the complementary roles of persistent evaluation history, stochastic return modeling, and distributional supervision.

Tue 8 SeptMachine Learning
The gist
Finding formulas that predict stock returns is tricky because feedback only comes after a complete formula is evaluated, making it hard to judge partial progress. The authors introduce AlphaRJM, which keeps track of rewards from complete formulas in a special memory that only updates when a formula finishes. It also models future payoffs with a technique that accounts for uncertainty. Together, these approaches help select better actions when building formulas and lead to more stable and better predictions in stock markets.
Open → 2609.08581v1

Nystrom attention matches full attention in stock prediction models

Nyström Attention Matches Full Attention for Cross-Sectional Stock Prediction

Abstract: MASTER's inter-stock multi-head attention -- the module responsible for modeling cross-sectional stock relationships -- accounts for 42.5% of model parameters and 25% of predictive value. We systematically decompose this module and uncover a surprising structure: the learned attention is near-uniform (perplexity 278/300), yet forcing exact uniformity eliminates all cross-sectional discrimination. Spectral analysis resolves this paradox: the deviation from uniformity is low-rank (effective rank ~65, top-10 modes capture 96.5% of energy), explaining why sparse approximations consistently fail while Nystrom low-rank attention (m=32 landmarks) matches full O(N^2) attention at O(mN) cost -- certified equivalent via TOST at both N=300 (5 seeds, Rank IC p=0.003) and N=800 (10 seeds, Rank IC p=0.034). Additional findings include: (i) attention anti-correlates with return similarity (Spearman rho = -0.614; on the industry-labeled subset, -0.645 unconditionally and -0.627 after controlling for industry, beta, and volatility), suggesting complementarity-seeking rather than correlation mining; (ii) all graph-based alternatives degrade performance, with hard masking worse than complete module removal; and (iii) at N ~ 3,500 with adapted architectures, no cross-stock module (GCN, Nystrom, or MASTER-style pipeline) significantly outperforms a per-stock LSTM baseline (n=4 seeds), indicating that the benefits observed at smaller scales do not trivially transfer. These results establish that the inter-stock attention's value resides in a compressible, dynamic, near-global redistribution that rewards low-rank approximation but resists sparsification.

Tue 8 SeptMachine Learning
The gist
Predicting stock movements involves understanding relationships between many stocks at once. The authors studied an important part of a stock prediction model that looks at these relationships, called attention, and found it behaves almost uniformly but with subtle patterns. They discovered a way to approximate this complex attention efficiently using a low-rank method called Nystrom attention without losing accuracy. This suggests that the key value in understanding stock relationships lies in a compressed, global way of sharing information rather than focusing on sparse or simple connections.
Open → 2609.08106v1

Structured communication speeds agent collaboration in finance alpha discovery

VST: Verifiable Structured Transport for Auditable Agent-to-Agent Alpha Discovery

Abstract: Agent-to-agent (A2A) alpha discovery is slowed by repeated feedback cycles between mining and evaluation agents, whose hand-offs, in contemporary LLM multi-agent systems, are free-form natural-language messages that carry no stable contract and cannot be replayed. We first restructure this communication as a structured agent-to-agent protocol of \emph{typed, causally addressable, unicast records}, so that the committed stream forms a causal trajectory. On that trajectory a single predictor with four typed heads forecasts the accumulated guidance the two miners would receive several cycles ahead; a transactional verify--leap controller then commits a multi-cycle speculative outcome only when it passes a four-level gate, and otherwise rolls back to the exact prior state. Structure is the enabling contribution, and its value is not accuracy. A controlled ablation shows an equal-information free-text channel reaches the same predictor hit rate. What typing provides is a state that can be schema-checked, replayed deterministically, and prevented by construction from leaking a forecast to an evaluator: auditability by construction, not an empirically stress-tested guarantee. On a CSI~1000 out-of-sample holdout, our single run is the only one among eight methods (seven baselines and ours) to hold a positive median annualized return and Sharpe at the factor level, though the median return \emph{in excess} of the benchmark stays negative for every method including ours; its development-selected top-20 portfolios reach a $0.71$ median holdout Sharpe, selected on a split inside the optimization horizon. We report these single-run results descriptively, gross of costs, and are explicit about their limits throughout; in particular we do not isolate the effect of the leap machinery from the inherited search substrate, which we leave to future work.

Mon 7 SeptArtificial Intelligence
The gist
Finding good investment ideas through multiple computer agents talking to each other can be slow because they exchange unclear messages. The paper’s authors reorganize this communication into clear, typed records that can be checked and replayed step-by-step, making the whole process more auditable and reliable. Their approach uses a special controller that predicts outcomes several steps ahead and only commits changes if certain checks pass, otherwise returning to a previous state. This structured method does not increase prediction accuracy but improves trust and traceability in how agents work together. When tested on financial data, their method outperformed several alternatives on some key measures, though overall returns compared to benchmarks were negative for all.
Open → 2609.07065v1