English
Related papers

Related papers: Asking Is Not Enough: Protocol Sensitivity in LLM …

200 papers

The open-ended nature of language generation makes the evaluation of autoregressive large language models (LLMs) challenging. One common evaluation approach uses multiple-choice questions (MCQ) to limit the response space. The model is then…

Computation and Language · Computer Science 2024-07-08 Xinpeng Wang , Bolei Ma , Chengzhi Hu , Leon Weber-Genzel , Paul Röttger , Frauke Kreuter , Dirk Hovy , Barbara Plank

Large language models (LLMs) are increasingly used in applications requiring factual accuracy, yet their outputs often contain hallucinated responses. While fact-checking can mitigate these errors, existing methods typically retrieve…

Computation and Language · Computer Science 2026-01-07 Haoran Wang , Maryam Khalid , Qiong Wu , Jian Gao , Cheng Cao

Large language models (LLMs) increasingly solve difficult problems by producing "reasoning traces" before emitting a final response. However, it remains unclear how accuracy and decision commitment evolve along a reasoning trajectory, and…

Machine Learning · Computer Science 2026-02-02 Marthe Ballon , Brecht Verbeken , Vincent Ginis , Andres Algaba

This paper systematically compares different methods of deriving item-level predictions of language models for multiple-choice tasks. It compares scoring methods for answer options based on free generation of responses, various…

Computation and Language · Computer Science 2024-03-05 Polina Tsvilodub , Hening Wang , Sharon Grosch , Michael Franke

Multimodal large language models (MLLMs) hold considerable promise for applications in healthcare. However, their deployment in safety-critical settings is hindered by two key limitations: (i) sensitivity to prompt design, and (ii) a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Anita Kriz , Elizabeth Laura Janes , Xing Shen , Tal Arbel

Multiple-choice question (MCQ) benchmarks have been a standard evaluation practice for measuring LLMs' ability to reason and answer knowledge-based questions. Through a synthetic NonsenseQA benchmark, we observe that different LLMs exhibit…

Computation and Language · Computer Science 2026-02-20 Mateusz Nowak , Xavier Cadet , Peter Chin

Large Language Models (LLMs) are increasingly applied to complex tasks that require extended reasoning. In such settings, models often benefit from diverse chains-of-thought to arrive at multiple candidate solutions. This requires two…

Machine Learning · Computer Science 2025-10-08 Xueyan Li , Guinan Su , Mrinmaya Sachan , Jonas Geiping

Large language models (LLMs) are increasingly deployed in high-stakes settings where good decisions require forming beliefs over the probability of unknown outcomes. However, it is unclear whether LLMs act as if they hold coherent beliefs…

Artificial Intelligence · Computer Science 2026-05-12 Khurram Yamin , Jingjing Tang , Santiago Cortes-Gomez , Amit Sharma , Eric Horvitz , Bryan Wilder

As evaluation designs of large language models may shape our trajectory toward artificial general intelligence, comprehensive and forward-looking assessment is essential. Existing benchmarks primarily assess static knowledge, while…

Computation and Language · Computer Science 2025-08-07 Jiayin Wang , Zhiquang Guo , Weizhi Ma , Min Zhang

Prompting is now a dominant method for evaluating the linguistic knowledge of large language models (LLMs). While other methods directly read out models' probability distributions over strings, prompting requires models to access this…

Computation and Language · Computer Science 2023-10-24 Jennifer Hu , Roger Levy

Despite warnings that LLMs can make mistakes, users often develop inappropriate trust and accept incorrect answers without critical evaluation. Uncertainty quantification (UQ), displaying LLMs' confidence, has emerged as a promising…

Human-Computer Interaction · Computer Science 2026-05-28 Mauricio Villavicencio , Sitong Pan , Qianwen Wang

Large language model (LLM) evaluations often assume there is a single correct response -- a gold label -- for each item in the evaluation corpus. However, some tasks can be ambiguous -- i.e., they provide insufficient information to…

Machine Learning · Computer Science 2024-11-22 Luke Guerdan , Hanna Wallach , Solon Barocas , Alexandra Chouldechova

Large Language Models deployed as question answering tools require robust calibration to avoid overconfidence. We systematically evaluate how reasoning capabilities and budget affect confidence assessment accuracy, using the ClimateX…

Artificial Intelligence · Computer Science 2025-08-22 Romain Lacombe , Kerrie Wu , Eddie Dilworth

Single-prompt first-token probabilities from zero-shot vision-language model (VLM) safety classifiers are treated as decision scores, but we show they are unreliable under semantically equivalent prompt reformulation: even when the binary…

Computation and Language · Computer Science 2026-05-04 Charles Weng , Dingwen Li , Alexander Martin

Large Language Models (LLMs) can produce surprisingly sophisticated estimates of their own uncertainty. However, it remains unclear to what extent this expressed confidence is tied to the reasoning, knowledge, or decision making of the…

Machine Learning · Computer Science 2026-01-13 Jiawei Wang , Yanfei Zhou , Siddartha Devic , Deqing Fu

Safety architectures for language models increasingly rely on external monitors to detect errors and inject corrective signals at inference time. For such systems to function in interactive settings, models must be able to incorporate…

Computation and Language · Computer Science 2026-01-15 Felipe Biava Cataneo

Large language models are increasingly deployed as protocols: structured multi-call procedures that spend additional computation to transform a baseline answer into a final one. These protocols are evaluated only by end-to-end accuracy,…

Machine Learning · Computer Science 2026-04-28 Fernando Reitich

Large Language Models (LLMs) are increasingly employed in various question-answering tasks. However, recent studies showcase that LLMs are susceptible to persuasion and could adopt counterfactual beliefs. We present a systematic evaluation…

Computation and Language · Computer Science 2026-03-20 Fan Huang , Haewoon Kwak , Jisun An

The versatility of Large Language Models (LLMs) on natural language understanding tasks has made them popular for research in social sciences. To properly understand the properties and innate personas of LLMs, researchers have performed…

Computation and Language · Computer Science 2024-04-03 Bangzhao Shu , Lechen Zhang , Minje Choi , Lavinia Dunagan , Lajanugen Logeswaran , Moontae Lee , Dallas Card , David Jurgens

Large language models (LLMs) achieve strong average performance yet remain unreliable at the instance level, with frequent hallucinations, brittle failures, and poorly calibrated confidence. We study reliability through the lens of…

Artificial Intelligence · Computer Science 2026-01-13 Pranav Kallem
‹ Prev 1 8 9 10 Next ›