English
Related papers

Related papers: Quantifying Semantic Shift in Financial NLP: Robus…

200 papers

The financial industry faces a critical dichotomy in AI adoption: deep learning often delivers strong empirical performance, while symbolic logic offers interpretability and rule adherence expected in regulated settings. We use Modal…

Machine Learning · Computer Science 2026-03-16 Antonin Sulc

AI systems are notorious for their fragility; minor input changes can potentially cause major output swings. When such systems are deployed in critical areas like finance, the consequences of their uncertain behavior could be severe. In…

Machine Learning · Computer Science 2024-06-21 Kausik Lakkaraju , Rachneet Kaur , Zhen Zeng , Parisa Zehtabi , Sunandita Patra , Biplav Srivastava , Marco Valtorta

This paper explores the topic of transportability, as a sub-area of generalisability. By proposing the utilisation of metrics based on well-established statistics, we are able to estimate the change in performance of NLP models in new…

Computation and Language · Computer Science 2021-05-04 Guy Marshall , Mokanarangan Thayaparan , Philip Osborne , Andre Freitas

Multi-agent LLM systems are increasingly used to solve complex tasks through decomposition, debate, specialization, and ensemble reasoning. However, these systems are usually evaluated in terms of robustness: whether performance is…

Multiagent Systems · Computer Science 2026-05-07 Jose Manuel de la Chica , Juan Manuel Vera , Jairo Rodríguez

We present semantic invariance testing, a method to test whether LLM self-explanations are faithful. A faithful self-report should remain stable when only the semantic context changes while the functional state stays fixed. We…

Computation and Language · Computer Science 2026-03-03 Stefan Szeider

The understanding of complex systems has become a central issue because complex systems exist in a wide range of scientific disciplines. Time series are typical experimental results we have about complex systems. In the analysis of such…

Statistical Finance · Quantitative Finance 2012-02-09 Michael C. Münnix , Takashi Shimada , Rudi Schäfer , Francois Leyvraz Thomas H. Seligman , Thomas Guhr , H. E. Stanley

Financial time series forecasting is fundamentally an information fusion challenge, yet most existing models rely on static architectures that struggle to integrate heterogeneous knowledge sources or adjust to rapid regime shifts.…

Artificial Intelligence · Computer Science 2025-12-23 Hafiz Saif Ur Rehman , Ling Liu , Kaleem Ullah Qasim

Evaluating robustness under temporal distribution shift remains an open challenge. Existing metrics quantify the average decline in performance, but fail to capture how models adapt to evolving data. As a result, temporal degradation is…

Machine Learning · Computer Science 2026-04-09 Lorenzo Iovine , Giacomo Ziffer , Emanuele Della Valle

Word embeddings are computed by a class of techniques within natural language processing (NLP), that create continuous vector representations of words in a language from a large text corpus. The stochastic nature of the training process of…

Computation and Language · Computer Science 2020-08-03 Lucas Rettenmeier

Text-based automated Cognitive Distortion detection is a challenging task due to its subjective nature, with low agreement scores observed even among expert human annotators, leading to unreliable annotations. We explore the use of Large…

Computation and Language · Computer Science 2026-05-21 Neha Sharma , Navneet Agarwal , Kairit Sirts

In many high-risk machine learning applications it is essential for a model to indicate when it is uncertain about a prediction. While large language models (LLMs) can reach and even surpass human-level accuracy on a variety of benchmarks,…

Computation and Language · Computer Science 2024-06-06 Evan Becker , Stefano Soatto

We introduce \textsc{CAT}, a framework designed to evaluate and visualize the \emph{interplay} of \emph{accuracy} and \emph{response consistency} of Large Language Models (LLMs) under controllable input variations, using multiple-choice…

Computation and Language · Computer Science 2026-01-01 Paulo Cavalin , Cassia Sanctos , Marcelo Grave , Claudio Pinhanez , Yago Primerano

Forecasting crude oil prices remains challenging because market-relevant information is embedded in large volumes of unstructured news and is not fully captured by traditional polarity-based sentiment measures. This paper examines whether…

Statistical Finance · Quantitative Finance 2026-03-18 Dehao Dai , Ding Ma , Dou Liu , Kerui Geng , Yiqing Wang

One of the long-standing goals in optimisation and constraint programming is to describe a problem in natural language and automatically obtain an executable, efficient model. Large language models appear to bring this vision closer,…

Artificial Intelligence · Computer Science 2025-11-20 Alessio Pellegrino , Jacopo Mauro

In this paper we present an exploratory research on quantifying the impact that data distribution has on the performance and evaluation of NLP models. We propose an automated framework that measures the data point distribution across 6…

Computation and Language · Computer Science 2024-04-02 Venelin Kovatchev , Matthew Lease

A systematic, comparative investigation into the effects of low-quality data reveals a stark spectrum of robustness across modern probabilistic models. We find that autoregressive language models, from token prediction to…

Artificial Intelligence · Computer Science 2025-12-16 Liu Peng , Yaochu Jin

Logic provides a controlled testbed for evaluating LLM-based reasoners, yet standard SAT-style benchmarks often conflate surface difficulty (length, wording, clause order) with the structural phenomena that actually determine…

Artificial Intelligence · Computer Science 2026-02-16 Naïm Es-sebbani , Esteban Marquer , Yakoub Salhi , Zied Bouraoui

Financial sentiment analysis is crucial for understanding the influence of news on stock prices. Recently, large language models (LLMs) have been widely adopted for this purpose due to their advanced text analysis capabilities. However,…

Computation and Language · Computer Science 2025-06-24 Yixuan Liang , Yuncong Liu , Neng Wang , Hongyang Yang , Boyu Zhang , Christina Dan Wang

This paper develops a prudential framework for assessing the reliability of large language models (LLMs) in reinsurance. A five-pillar architecture--governance, data lineage, assurance, resilience, and regulatory alignment--translates…

Artificial Intelligence · Computer Science 2025-11-12 Stella C. Dong

Large Language Models (LLMs) have shown promising performance in software vulnerability detection, particularly after domain-specific Supervised Fine-Tuning (SFT). However, it remains unclear whether these models genuinely internalize…

Cryptography and Security · Computer Science 2026-05-22 Feiyang Huang , Yuqiang Sun , Fan Zhang , Ziqi Yang , Han Liu , Yang Liu