English
Related papers

Related papers: Let the Results Speak: A Replication-First Paradig…

200 papers

The impressive performance of large language models (LLMs) has attracted considerable attention from the academic and industrial communities. Besides how to construct and train LLMs, how to effectively evaluate and compare the capacity of…

Information Retrieval · Computer Science 2024-06-04 Zhumin Chu , Qingyao Ai , Yiteng Tu , Haitao Li , Yiqun Liu

Large language models (LLMs) with reasoning capabilities have fueled a compelling narrative that reasoning universally improves performance across language tasks. We test this claim through a comprehensive evaluation of 504 configurations…

Computation and Language · Computer Science 2026-03-02 Donghao Huang , Zhaoxia Wang

Language models (LMs) should provide reliable confidence estimates to help users detect mistakes in their outputs and defer to human experts when necessary. Asking a language model to assess its confidence ("Score your confidence from…

Computation and Language · Computer Science 2025-02-04 Vaishnavi Shrivastava , Ananya Kumar , Percy Liang

The trustworthiness of Large Language Models (LLMs) refers to the extent to which their outputs are reliable, safe, and ethically aligned, and it has become a crucial consideration alongside their cognitive performance. In practice,…

Computation and Language · Computer Science 2024-12-24 Aaron J. Li , Satyapriya Krishna , Himabindu Lakkaraju

Large Language Models (LLMs) are increasingly adopted as evaluators, offering a scalable alternative to human annotation. However, existing supervised fine-tuning (SFT) approaches often fall short in domains that demand complex reasoning.…

Computation and Language · Computer Science 2025-11-04 Nuo Chen , Zhiyuan Hu , Qingyun Zou , Jiaying Wu , Qian Wang , Bryan Hooi , Bingsheng He

We characterize a compositional architecture of literary primitives in two instruction-tuned large language models (Llama 3.1 8B-Instruct and Gemma 2 9B-IT) via sparse autoencoders on mid-depth residual streams. Four feature classes emerge:…

Machine Learning · Computer Science 2026-05-20 Joao Paulo Cavalcante Presa , Savio Salvarino Teles de Oliveira

Self-assessment is a key aspect of reliable intelligence, yet evaluations of large language models (LLMs) focus mainly on task accuracy. We adapted the 10-item General Self-Efficacy Scale (GSES) to elicit simulated self-assessments from ten…

Artificial Intelligence · Computer Science 2025-11-27 Daniel I Jackson , Emma L Jensen , Syed-Amad Hussain , Emre Sezgin

Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchmarks evaluate this capability solely on numerical Big Five score prediction, leaving open…

Artificial Intelligence · Computer Science 2026-05-22 Caixin Kang , Tianyu Yan , Sitong Gong , Mingfang Zhang , Liangyang Ouyang , Ruicong Liu , Bo Zheng , Huchuan Lu , Kaipeng Zhang , Yoichi Sato , Yifei Huang

We present a principled approach to provide LLM-based evaluation with a rigorous guarantee of human agreement. We first propose that a reliable evaluation method should not uncritically rely on model preferences for pairwise evaluation, but…

Machine Learning · Computer Science 2024-07-29 Jaehun Jung , Faeze Brahman , Yejin Choi

Retrieval-augmented generation (RAG) enables large language models (LLMs) to generate answers with citations from source documents containing "ground truth", thereby reducing system hallucinations. A crucial factor in RAG evaluation is…

Computation and Language · Computer Science 2025-04-22 Nandan Thakur , Ronak Pradeep , Shivani Upadhyay , Daniel Campos , Nick Craswell , Jimmy Lin

Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Overcoming this challenge advances key applications, such as…

Reasoning language models can solve increasingly complex tasks, but struggle to produce the calibrated confidence estimates necessary for reliable deployment. Existing calibration methods usually depend on labels or repeated sampling at…

Machine Learning · Computer Science 2026-04-22 Thomas Zollo , Jimmy Wang , Richard Zemel

A critical component in the trustworthiness of LLMs is reliable uncertainty communication, yet LLMs often use assertive language when conveying false claims, leading to over-reliance and eroded trust. We present the first systematic study…

Computation and Language · Computer Science 2025-10-03 Gabrielle Kaili-May Liu , Gal Yona , Avi Caciularu , Idan Szpektor , Tim G. J. Rudner , Arman Cohan

The miscalibration of Large Reasoning Models (LRMs) undermines their reliability in high-stakes domains, necessitating methods to accurately estimate the confidence of their long-form, multi-step outputs. To address this gap, we introduce…

Computation and Language · Computer Science 2026-01-22 Reza Khanmohammadi , Erfan Miahi , Simerjot Kaur , Ivan Brugere , Charese H. Smiley , Kundan Thind , Mohammad M. Ghassemi

Large language models cannot estimate how long their own tasks take. We investigate this limitation through four experiments across 68 tasks and four model families. Pre-task estimates overshoot actual duration by 4--7$\times$ ($p <…

Computation and Language · Computer Science 2026-04-02 Aniketh Garikaparthi

Large language models can answer causal questions correctly for the wrong reasons. Current RL methods reward \emph{what} a model concludes but ignore \emph{why}, reinforcing correlational shortcuts -- a failure we call \emph{Reward…

Artificial Intelligence · Computer Science 2026-05-21 Edward Y. Chang , Longling Geng

As large language models (LLMs) advance, there is growing interest in using them to simulate human social behavior through generative agent-based modeling (GABM). However, validating these models remains a key challenge. We present a…

Multiagent Systems · Computer Science 2025-07-30 Logan Cross , Nick Haber , Daniel L. K. Yamins

We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments. Using TradeArena, an auditable trading-agent testbed with risk reports, execution simulation, memory, and…

Machine Learning · Computer Science 2026-05-29 Weicheng Xue

Inference-time scaling has attracted much attention which significantly enhance the performance of Large Language Models (LLMs) in complex reasoning tasks by increasing the length of Chain-of-Thought. These longer intermediate reasoning…

Computation and Language · Computer Science 2025-05-21 Hongru Wang , Deng Cai , Wanjun Zhong , Shijue Huang , Jeff Z. Pan , Zeming Liu , Kam-Fai Wong

Verbal confidence elicitation is widely used to extract uncertainty estimates from LLMs. We tested whether seven instruction-tuned open-weight models (3-9B parameters, four families) produce verbalised confidence that meets minimal validity…

Computation and Language · Computer Science 2026-04-27 Jon-Paul Cacioli