English
Related papers

Related papers: Reasoning Models Ace the CFA Exams

200 papers

There is a growing literature on reasoning by large language models (LLMs), but the discussion on the uncertainty in their responses is still lacking. Our aim is to assess the extent of confidence that LLMs have in their answers and how it…

Computation and Language · Computer Science 2024-12-23 Yudi Pawitan , Chris Holmes

In an era dominated by Large Language Models (LLMs), understanding their capabilities and limitations, especially in high-stakes fields like law, is crucial. While LLMs such as Meta's LLaMA, OpenAI's ChatGPT, Google's Gemini, DeepSeek, and…

Computation and Language · Computer Science 2025-09-29 Antreas Ioannou , Andreas Shiamishis , Nora Hollenstein , Nezihe Merve Gürel

Leading large language models have demonstrated impressive capabilities in reasoning-intensive tasks, such as standardized educational testing. However, they often require extensive training in low-resource settings with inaccessible…

Computation and Language · Computer Science 2025-03-19 Mykyta Syromiatnikov , Victoria Ruvinskaya , Nataliia Komleva

We investigate consciousness-like behaviors in Large Language Models (LLMs) using the Maze Test, challenging models to navigate mazes from a first-person perspective. This test simultaneously probes spatial awareness, perspective-taking,…

Computation and Language · Computer Science 2025-08-26 Rui A. Pimenta , Tim Schlippe , Kristina Schaaff

AI chatbots are increasingly used by students as study tools in physics, raising practical questions about their reliability on conceptual tasks. Existing evaluations of large language models (LLMs) on physics concept inventories rely…

Physics Education · Physics 2026-05-12 Eugenio Tufino , Caterina Giovanzana , Andrea Zamboni , Pasquale Onorato , Stefano Oss

Evaluating large language models (LLMs) as judges is increasingly critical for building scalable and trustworthy evaluation pipelines. We present ScalingEval, a large-scale benchmarking study that systematically compares 36 LLMs, including…

Artificial Intelligence · Computer Science 2025-11-06 Tao Zhang , Kehui Yao , Luyi Ma , Jiao Chen , Reza Yousefi Maragheh , Kai Zhao , Jianpeng Xu , Evren Korpeoglu , Sushant Kumar , Kannan Achan

Automating the classification of negative treatment in legal precedent is a critical yet nuanced NLP task where misclassification carries significant risk. To address the shortcomings of standard accuracy, this paper introduces a more…

Computation and Language · Computer Science 2026-05-19 M. Mikail Demir , M. Abdullah Canbaz

Large Language Models (LLMs) have demonstrated remarkable abilities in scientific reasoning, yet their reasoning capabilities in materials science remain underexplored. To fill this gap, we introduce MatSciBench, a comprehensive…

Artificial Intelligence · Computer Science 2025-10-15 Junkai Zhang , Jingru Gan , Xiaoxuan Wang , Zian Jia , Changquan Gu , Jianpeng Chen , Yanqiao Zhu , Mingyu Derek Ma , Dawei Zhou , Ling Li , Wei Wang

Large language models (LLMs) are rapidly being adopted across various domains. However, their adoption in banking industry faces resistance due to demands for high accuracy, regulatory compliance, and the need for verifiable and grounded…

Large Language Models (LLMs) are increasingly used as evaluators of reasoning quality, yet their reliability and bias in payments-risk settings remain poorly understood. We introduce a structured multi-evaluator framework for assessing LLM…

Artificial Intelligence · Computer Science 2026-02-06 Liang Wang , Junpeng Wang , Chin-chia Michael Yeh , Yan Zheng , Jiarui Sun , Xiran Fan , Xin Dai , Yujie Fan , Yiwei Cai

In this paper, we explore how large language models (LLMs) approach financial decision-making by systematically comparing their responses to those of human participants across the globe. We posed a set of commonly used financial…

General Economics · Economics 2025-07-16 Orhan Erdem , Ragavi Pobbathi Ashok

This paper introduces AQA-Bench, a novel benchmark to assess the sequential reasoning capabilities of large language models (LLMs) in algorithmic contexts, such as depth-first search (DFS). The key feature of our evaluation benchmark lies…

Computation and Language · Computer Science 2025-06-23 Siwei Yang , Bingchen Zhao , Cihang Xie

Large language models (LLMs) have demonstrated strong capabilities in programming and mathematical reasoning tasks, but are constrained by limited high-quality training data. Synthetic data can be leveraged to enhance fine-tuning outcomes,…

Machine Learning · Computer Science 2025-04-28 Caia Costello , Simon Guo , Anna Goldie , Azalia Mirhoseini

The legal field already uses various large language models (LLMs) in actual applications, but their quantitative performance and reasons for it are underexplored. We evaluated several open-source and proprietary LLMs -- including…

Computers and Society · Computer Science 2025-09-12 Bhakti Khera , Rezvan Alamian , Pascal A. Scherz , Stephan M. Goetz

This paper explores the automatic classification of exam questions and learning outcomes according to Bloom's Taxonomy. A small dataset of 600 sentences labeled with six cognitive categories - Knowledge, Comprehension, Application,…

Computation and Language · Computer Science 2025-11-17 Ramya Kumar , Dhruv Gulwani , Sonit Singh

Large language models (LLMs) are emerging as valuable tools to support clinicians in routine decision-making. HIV management is a compelling use case due to its complexity, including diverse treatment options, comorbidities, and adherence…

Computation and Language · Computer Science 2025-07-28 Gonzalo Cardenal-Antolin , Jacques Fellay , Bashkim Jaha , Roger Kouyos , Niko Beerenwinkel , Diane Duroux

Large language models (LLMs) are increasingly integrated into high-stakes decision-making. Inspired by the theory of \emph{inattentional blindness} in human cognition, we investigate whether LLMs, trained on human-preferred corpora that…

Computation and Language · Computer Science 2026-05-20 Yuanqing Cai , Ziyi Huang , Minhao Liu , Lixin Duan , Wen Li , Yanru Zhang

This paper evaluates the knowledge and reasoning capabilities of Large Language Models in Islamic inheritance law, known as 'ilm al-mawarith. We assess the performance of seven LLMs using a benchmark of 1,000 multiple-choice questions…

Computation and Language · Computer Science 2025-09-18 Abdessalam Bouchekif , Samer Rashwani , Heba Sbahi , Shahd Gaben , Mutaz Al-Khatib , Mohammed Ghaly

Public leaderboards increasingly suggest that large language models (LLMs) surpass human experts on benchmarks spanning academic knowledge, law, and programming. Yet most benchmarks are fully public, their questions widely mirrored across…

Artificial Intelligence · Computer Science 2026-03-18 Eshwar Reddy M , Sourav Karmakar