English
Related papers

Related papers: CLASH: Evaluating Language Models on Judging High-…

200 papers

Large language models (LLMs) are increasingly used to meet user information needs, but their effectiveness in dealing with user queries that contain various types of ambiguity remains unknown, ultimately risking user trust and satisfaction.…

Computation and Language · Computer Science 2024-06-04 Tong Zhang , Peixin Qin , Yang Deng , Chen Huang , Wenqiang Lei , Junhong Liu , Dingnan Jin , Hongru Liang , Tat-Seng Chua

Mathematical reasoning serves as a cornerstone for assessing the fundamental cognitive capabilities of human intelligence. In recent times, there has been a notable surge in the development of Large Language Models (LLMs) geared towards the…

Computation and Language · Computer Science 2024-09-18 Janice Ahn , Rishu Verma , Renze Lou , Di Liu , Rui Zhang , Wenpeng Yin

Contradictory multimodal inputs are common in real-world settings, yet existing benchmarks typically assume input consistency and fail to evaluate cross-modal contradiction detection - a fundamental capability for preventing hallucinations…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Teodora Popordanoska , Jiameng Li , Matthew B. Blaschko

Large Language Models (LLMs) have achieved remarkable performance across a wide range of mathematical benchmarks. However, concerns remain as to whether these successes reflect genuine reasoning or superficial pattern recognition. Existing…

Artificial Intelligence · Computer Science 2026-04-21 Yujie Hou , Mei Wang , Yaoyao Zhong , Ting Zhang , Xuetao Ma , Hua Huang

We introduce LLM CHESS, an evaluation framework designed to probe the generalization of reasoning and instruction-following abilities in large language models (LLMs) through extended agentic interaction in the domain of chess. We rank over…

Artificial Intelligence · Computer Science 2025-12-02 Sai Kolasani , Maxim Saplin , Nicholas Crispino , Kyle Montgomery , Jared Quincy Davis , Matei Zaharia , Chi Wang , Chenguang Wang

Minimizing negative impacts of Artificial Intelligent (AI) systems on human societies without human supervision requires them to be able to align with human values. However, most current work only addresses this issue from a technical point…

Computation and Language · Computer Science 2024-08-13 Mehdi Khamassi , Marceau Nahon , Raja Chatila

High-stakes decision domains are increasingly exploring the potential of Large Language Models (LLMs) for complex decision-making tasks. However, LLM deployment in real-world settings presents challenges in data security, evaluation of its…

Computers and Society · Computer Science 2025-12-05 Swati Sachan , Theo Miller , Mai Phuong Nguyen

Governments are increasingly considering integrating autonomous AI agents in high-stakes military and foreign-policy decision-making, especially with the emergence of advanced generative AI models like GPT-4. Our work aims to scrutinize the…

Artificial Intelligence · Computer Science 2024-06-13 Juan-Pablo Rivera , Gabriel Mukobi , Anka Reuel , Max Lamparth , Chandler Smith , Jacquelyn Schneider

We investigate whether large language models (LLMs) can predict whether they will succeed on a given task and whether their predictions improve as they progress through multi-step tasks. We also investigate whether LLMs can learn from…

Computation and Language · Computer Science 2026-01-01 Casey O. Barkan , Sid Black , Oliver Sourbut

Supportive conversation depends on skills that go beyond language fluency, including reading emotions, adjusting tone, and navigating moments of resistance, frustration, or distress. Despite rapid progress in language models, we still lack…

Computation and Language · Computer Science 2026-02-26 Laya Iyer , Kriti Aggarwal , Sanmi Koyejo , Gail Heyman , Desmond C. Ong , Subhabrata Mukherjee

Large Language Models (LLMs) have shown strong performance on NLP classification tasks. However, they typically rely on aggregated labels-often via majority voting-which can obscure the human disagreement inherent in subjective annotations.…

Computation and Language · Computer Science 2025-06-09 Benedetta Muscato , Yue Li , Gizem Gezici , Zhixue Zhao , Fosca Giannotti

AI models are already deployed in societies affected by armed conflict, and journalists, humanitarian workers, governments and ordinary citizens rely on them for information or for their work processes. No established practice exists for…

Artificial Intelligence · Computer Science 2026-05-22 Andrii Kryshtal

Large Language Models (LLMs) are increasingly used in decision-making, yet their susceptibility to cognitive biases remains a pressing challenge. This study explores how personality traits influence these biases and evaluates the…

Artificial Intelligence · Computer Science 2025-02-21 Jiangen He , Jiqun Liu

The $\textit{LLM-as-a-judge}$ paradigm has become the operational backbone of automated AI evaluation pipelines, yet rests on an unverified assumption: that judges evaluate text strictly on its semantic content, impervious to surrounding…

Artificial Intelligence · Computer Science 2026-04-17 Manan Gupta , Inderjeet Nair , Lu Wang , Dhruv Kumar

Large language models (LLMs) can internally distinguish between evaluation and deployment contexts, a behaviour known as \emph{evaluation awareness}. This undermines AI safety evaluations, as models may conceal dangerous capabilities during…

Artificial Intelligence · Computer Science 2025-11-11 Maheep Chaudhary , Ian Su , Nikhil Hooda , Nishith Shankar , Julia Tan , Kevin Zhu , Ryan Lagasse , Vasu Sharma , Ashwinee Panda

In difficult decision-making scenarios, it is common to have conflicting opinions among expert human decision-makers as there may not be a single right answer. Such decisions may be guided by different attributes that can be used to…

Computation and Language · Computer Science 2024-06-11 Brian Hu , Bill Ray , Alice Leung , Amy Summerville , David Joy , Christopher Funk , Arslan Basharat

Can Large Language Models (LLMs) simulate humans in making important decisions? Recent research has unveiled the potential of using LLMs to develop role-playing language agents (RPLAs), mimicking mainly the knowledge and tones of various…

Artificial Intelligence · Computer Science 2024-11-19 Rui Xu , Xintao Wang , Jiangjie Chen , Siyu Yuan , Xinfeng Yuan , Jiaqing Liang , Zulong Chen , Xiaoqing Dong , Yanghua Xiao

When a dangerous international crisis begins, leaders need to know whether their next move is going to resolve the dispute or amplify it out of control. Theories of conflict have mainly served to deepen the confusion, revealing fighting,…

Physics and Society · Physics 2024-02-07 Rex W. Douglass , Erik Gartzke , Jon R. Lindsay , J. Andrés Gannon , Thomas Leo Scherer

Large Language Models (LLMs) are increasingly deployed in autonomous decision-making roles across high-stakes domains. However, since models are trained on human-generated data, they may inherit cognitive biases that systematically distort…

Artificial Intelligence · Computer Science 2025-08-08 Emilio Barkett , Olivia Long , Paul Kröger

The rapid advancement of Large Language Models (LLMs) has led to performance saturation on many established benchmarks, questioning their ability to distinguish frontier models. Concurrently, existing high-difficulty benchmarks often suffer…