F$^2$Agent: Financial Fusion of Agentic Intelligence for Multimodal Trading
Abstract
With increasingly diverse and heterogeneous information sources, effectively leveraging multimodal data is becoming pivotal for high-quality financial trading. Although recent advancements in Large Language Model (LLM)-based agents have enabled the ingestion of multimodal inputs, existing methods fail to capture nuanced cross-modal dependencies and remain vulnerable to market noise, due to limited multimodal modeling, ineffective fusion mechanisms, and inadequate robustness. To address these challenges, we propose FAgent, a novel multimodal agentic paradigm driven by the Financial Fusion of Agentic Intelligence. FAgent first deploys a hierarchy of specialized agents to comprehensively extract modality-specific signals. It further introduces a modality-aware adaptive fusion mechanism coupled with noise-robust consistency regularization to dynamically capture fine-grained inter-modality dependencies and generate noise-resilient trading signals. Extensive experiments on six stocks and cryptocurrency assets demonstrate that FAgent consistently outperforms 16 competitive baselines across multiple trading metrics, with over 20% relative improvement in annualized return on average. Notably, FAgent delivers returns of 120.48% on GOOG and 148.41% on TSLA, demonstrating its efficacy and robustness in varying market dynamics.
Keywords
Cite
@article{arxiv.2608.05668,
title = {F$^2$Agent: Financial Fusion of Agentic Intelligence for Multimodal Trading},
author = {Changshuo Liu and Yanzheng Jin and Shangfeng Cai and Peng Fang and Xiaokui Xiao and Beng Chin Ooi},
journal= {arXiv preprint arXiv:2608.05668},
year = {2026}
}
Comments
32 pages, 12 figures, 19 tables