English

From Text to Alpha: Can LLMs Track Evolving Signals in Corporate Disclosures?

Computational Engineering, Finance, and Science 2026-03-17 v5

Abstract

Natural language processing (NLP) has been widely used in quantitative finance, but traditional methods often struggle to capture rich narratives in corporate disclosures, leaving potentially informative signals under-explored. Large language models (LLMs) offer a promising alternative due to their ability to extract nuanced semantics. In this paper, we ask whether semantic signals extracted by LLMs from corporate disclosures predict alpha, defined as abnormal returns beyond broad market movements and common risk factors. We introduce a simple framework, LLM as extractor, embedding as ruler, which extracts context-aware, metric-focused textual spans and quantifies semantic changes across consecutive disclosure periods using embedding-based similarity. This allows us to measure the degree of metric shifting -- how much firms move away from previously emphasized metrics, referred as moving targets. In experiments with portfolio and cross-sectional regression tests against a recent NER-based baseline, our method achieves more than twice the risk-adjusted alpha and shows significantly stronger predictive power. Qualitative analysis suggests that these gains stem from preserving contextual qualifiers and filtering out non-metric terms that keyword-based approaches often miss.

Keywords

Cite

@article{arxiv.2510.03195,
  title  = {From Text to Alpha: Can LLMs Track Evolving Signals in Corporate Disclosures?},
  author = {Chanyeol Choi and Yoon Kim and Yu Yu and Young Cha and V. Zach Golkhou and Igor Halperin and Georgios Papaioannou and Minkyu Kim and Zhangyang Wang and Jihoon Kwon and Minjae Kim and Alejandro Lopez-Lira and Yongjae Lee},
  journal= {arXiv preprint arXiv:2510.03195},
  year   = {2026}
}

Comments

9 pages