Papers for

financial technology developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Benchmark tests how ai agents authorize payments securely

APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport

Abstract: APort Vault is a benchmark for payment authorization in tool-using AI agents. It replays 4,371 attacks written by humans against a live payment agent during a public capture-the-flag event, across 14 models from 8 labs, five policy configurations and two replay tracks, with and without a deterministic pre-action check implementing the Open Agent Passport (OAP) specification. 225,964 evaluations completed. We report five distinct events per evaluation, because collapsing them is how an agent benchmark produces a number that does not survive review. Requests are common and their rate differs far more across configurations than across models, though each attack exists at exactly one configuration so policy and attack cohort vary together: 10.9% of model-alone evaluations at Level 1, 3.0% at Level 2, 0.1% at Level 3, 79.4% at Level 4. On the 1,293 Level 4 prompts, each evaluated on every model, request rates run from 71.2% to 84.3%, and 809 prompts (62.6%) elicited a request from all fourteen models, each ending in a successful payment to the level's allowlisted recipient. The authorization boundary is where the conditions diverge. At Levels 2 to 4, transfers to recipients the passport did not permit number 140 of 76,842 with the model alone and 0 of 69,297 behind the layer, and 105 against 0 on 68,970 matched model, prompt and track triples. The zero spans 790 source sessions, giving a per-session upper bound of 0.38%. It was not obtained by refusing payments: 25,370 payments executed behind the layer, while the policy denied 187 of the 25,640 transfer calls it evaluated, 148 of them for a forbidden recipient. We release the 225,964 evaluations, the level passports, the scoring code and the analysis script at huggingface.co/datasets/aporthq/vault-benchmark-v1 .

Fri 18 SeptCryptography and Security
The gist
Payment systems powered by AI need to be safe from attackers trying to trick them into making unauthorized payments. The authors created a large set of tests called APort Vault, which uses thousands of real attack attempts to evaluate how well AI agents follow payment rules. They also tested a security layer called the Open Agent Passport, which stopped all unauthorized payment transfers in these tests. The study shows the effectiveness of this extra checking step in keeping AI-driven payments secure.
Open 2609.22076v1

How people judge AI financial advice depends on style and source

Trustworthy FinAInce: Unpacking How AI-Mediated Financial Advice is Judged

Abstract: As generative AI is increasingly used as a source of personal financial guidance, understanding how people appraise such advice is important for supporting appropriate reliance. We conducted a randomized vignette experiment with 285 U.S. adults across eight financial decisions, independently varying three advice styles---AI, expert, and online community---and displayed source labels while holding the underlying recommendation consistent. Advice style most strongly shaped message and safety appraisals, Expert labels selectively increased perceived source knowledge, and decision context primarily shaped risk and safety appraisals. These appraisals were associated with downstream judgments, with models explaining 69.2% of overall quality, 75.9% of trust, and 82.9% of intended reliance. Expert-style advice also remained most preferred when shown without source labels. Our findings have implications for understanding financial advice evaluation, distinguishing the roles of advice style and source labels, and designing financial AI that supports grounded evaluation rather than simply maximizing trust.

Thu 17 SeptHuman-Computer InteractionArtificial IntelligenceComputation and Language
The gist
Financial advice can come from AI, experts, or online communities, but people trust and rely on each source differently. The authors studied how labeling the advice source and the style of advice affect people's opinions on trustworthiness and safety. They found expert advice is generally preferred and seen as more knowledgeable, even if the advice content is the same. The decision context also changes how risky or safe the advice seems. These insights can help design financial AI tools that encourage careful evaluation rather than blind trust.
Open 2609.20989v1