English
Related papers

Related papers: Position: Stop Evaluating AI with Human Tests, Dev…

200 papers

AI-assisted decision making becomes increasingly prevalent, yet individuals often fail to utilize AI-based decision aids appropriately especially when the AI explanations are absent, potentially as they do not %understand reflect on AI's…

Human-Computer Interaction · Computer Science 2025-02-18 Zhuoyan Li , Hangxiao Zhu , Zhuoran Lu , Ziang Xiao , Ming Yin

As AI systems advance in capabilities, measuring their safety and alignment to human values is becoming paramount. A fast-growing field of AI research is devoted to developing such assessments. However, most current advances therein may be…

Computers and Society · Computer Science 2026-03-17 Max Hellrigel-Holderbaum , Edward James Young

When researchers claim AI systems possess ToM or mental models, they are fundamentally discussing behavioral predictions and bias corrections rather than genuine mental states. This position paper argues that the current discourse conflates…

Human-Computer Interaction · Computer Science 2025-10-06 Xiaoyun Yin , Elmira Zahmat Doost , Shiwen Zhou , Garima Arya Yadav , Jamie C. Gorman

The widespread adoption of Large Language Models (LLMs) and publicly available ChatGPT have marked a significant turning point in the integration of Artificial Intelligence (AI) into people's everyday lives. This study examines the ability…

Computation and Language · Computer Science 2025-10-28 Sandeep Kumar , Tirthankar Ghosal , Vinayak Goyal , Asif Ekbal

Objective and scalable measurement of teaching quality is a persistent challenge in education. While Large Language Models (LLMs) offer potential, general-purpose models have struggled to reliably apply complex, authentic classroom…

Computation and Language · Computer Science 2025-11-07 Michael Hardy

Evaluating the performance of Large Language Models (LLMs) is a critical yet challenging task, particularly when aiming to avoid subjective assessments. This paper proposes a framework for leveraging subjective metrics derived from the…

Computation and Language · Computer Science 2025-08-13 Haoze Du , Richard Li , Edward Gehringer

Have Large Language Models (LLMs) developed a personality? The short answer is a resounding "We Don't Know!". In this paper, we show that we do not yet have the right tools to measure personality in language models. Personality is an…

Computation and Language · Computer Science 2023-05-25 Xiaoyang Song , Akshat Gupta , Kiyan Mohebbizadeh , Shujie Hu , Anant Singh

Large language models (LLMs) distinguish themselves from previous technologies by functioning as collaborative ``thought partners,'' capable of engaging more fluidly in natural language on a range of tasks. As LLMs increasingly influence…

In the rapidly evolving landscape of Large Language Models (LLMs), introduction of well-defined and standardized evaluation methodologies remains a crucial challenge. This paper traces the historical trajectory of LLM evaluations, from the…

Computation and Language · Computer Science 2023-11-06 Alexey Tikhonov , Ivan P. Yamshchikov

Many existing benchmarks of large (multimodal) language models (LLMs) focus on measuring LLMs' academic proficiency, often with also an interest in comparing model performance with human test takers'. While such benchmarks have proven key…

Computation and Language · Computer Science 2025-06-25 Qixiang Fang , Daniel L. Oberski , Dong Nguyen

This paper investigates the ability of large language models (LLMs) to solve statistical tasks, as well as their capacity to assess the quality of reasoning. While state-of-the-art LLMs have demonstrated remarkable performance in a range of…

Computation and Language · Computer Science 2026-01-22 Crish Nagarkar , Leonid Bogachev , Serge Sharoff

Developments in the field of Artificial Intelligence (AI), and particularly large language models (LLMs), have created a 'perfect storm' for observing 'sparks' of Artificial General Intelligence (AGI) that are spurious. Like simpler models,…

Artificial Intelligence · Computer Science 2024-06-03 Patrick Altmeyer , Andrew M. Demetriou , Antony Bartlett , Cynthia C. S. Liem

Our paper argues that the majority of theory of mind benchmarks are broken because of their inability to directly test how large language models (LLMs) adapt to new partners. This problem stems from the fact that theory of mind benchmarks…

Artificial Intelligence · Computer Science 2025-06-13 Matthew Riemer , Zahra Ashktorab , Djallel Bouneffouf , Payel Das , Miao Liu , Justin D. Weisz , Murray Campbell

Large Language Models (LLMs) are increasingly used in everyday life and research. One of the most common use cases is conversational interactions, enabled by the language generation capabilities of LLMs. Just as between two humans, a…

Computation and Language · Computer Science 2024-11-12 Jingyao Zheng , Xian Wang , Simo Hosio , Xiaoxian Xu , Lik-Hang Lee

Traditional methods for eliciting people's opinions face a trade-off between depth and scale: structured surveys enable large-scale data collection but limit respondents' ability to voice their opinions in their own words, while…

Human-Computer Interaction · Computer Science 2025-03-13 Alexander Wuttke , Matthias Aßenmacher , Christopher Klamm , Max M. Lang , Quirin Würschinger , Frauke Kreuter

Psychometric tests are increasingly used to assess psychological constructs in large language models (LLMs). However, it remains unclear whether these tests -- originally developed for humans -- yield meaningful results when applied to…

Computation and Language · Computer Science 2026-01-28 Jana Jung , Marlene Lutz , Indira Sen , Markus Strohmaier

This study explores the potential of Large Language Models (LLMs), specifically GPT-4, to enhance objectivity in organizational task performance evaluations. Through comparative analyses across two studies, including various task…

Computation and Language · Computer Science 2024-08-13 Ning Li , Huaikang Zhou , Mingze Xu

Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current LLM-based human simulations remain insufficiently reliable, as evidenced by significant…

Computation and Language · Computer Science 2025-12-02 Qian Wang , Jiaying Wu , Zichen Jiang , Zhenheng Tang , Bingqiao Luo , Nuo Chen , Wei Chen , Bingsheng He

Artificial intelligence (AI), particularly in the form of large language models (LLMs) or chatbots, has become increasingly integrated into our daily lives. In the past five years, several LLMs have been introduced, including ChatGPT by…

Human-Computer Interaction · Computer Science 2026-05-06 Mouhacine Benosman

As evaluation designs of large language models may shape our trajectory toward artificial general intelligence, comprehensive and forward-looking assessment is essential. Existing benchmarks primarily assess static knowledge, while…

Computation and Language · Computer Science 2025-08-07 Jiayin Wang , Zhiquang Guo , Weizhi Ma , Min Zhang