TradeTrap:基于LLM的交易智能体真正可靠且忠实吗?
摘要
基于LLM的交易智能体日益被部署到实际金融市场,以执行 autonomous analysis (自主分析) 和 execution (执行)。然而,这些智能体在 adversarial (对抗) 或 faulty (故障) 条件下的 reliability (可靠性) 和 robustness (鲁棒性) 尚未得到充分考察,尽管它们正在 high-risk, irreversible (高风险、不可逆) 的金融环境中运行。我们提出TradeTrap,一个 unified evaluation framework (统一评估框架) 用于 systemically stress-testing (系统性压力测试) 所有 adaptive (自适应) 和 procedural (程序化) 自主交易智能体。TradeTrap 针对 autonomous trading agents (自主交易智能体) 的四个核心组件:market intelligence (市场情报)、strategy formulation (策略构建)、portfolio and ledger handling (投组合与账簿处理) 以及 trade execution (交易执行),在受控 system-level perturbations (系统层面扰动) 下评估其 robust (鲁棒性)。所有评估均在 real US equity market data (真实美国股票市场数据) 上进行的 closed-loop historical backtesting (封闭回测) 设置中完成,保持 identical initial conditions (相同的初始条件),实现 fair and reproducible comparisons (公平且可复现的比较) across agents and attacks (跨智能体与攻击)。广泛实验表明,单个组件的 small perturbations (小扰动) 可通过 agent decision loop (智能体决策循环) 传播,导致 extreme concentration (极端集中)、runaway exposure (暴露失控) 以及 large portfolio drawdowns (大型投组合回撤),从而在两种智能体类型中 demonstrate (展示) 当前自主交易智能体可在 system level (系统层面) 被 systemically misled (系统性误导)。我们的代码已在GitHub公开。
引用
@article{arxiv.2512.02261,
title = {TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?},
author = {Lewen Yan and Jilin Mei and Tianyi Zhou and Lige Huang and Jie Zhang and Dongrui Liu and Jing Shao},
journal= {arXiv preprint arXiv:2512.02261},
year = {2025}
}