English
Related papers

Related papers: Position: Stop Evaluating AI with Human Tests, Dev…

200 papers

Large language models (LLMs) are recognized as systems that closely mimic aspects of human intelligence. This capability has attracted attention from the social science community, who see the potential in leveraging LLMs to replace human…

Computers and Society · Computer Science 2025-03-04 Qiuejie Xie , Qiming Feng , Tianqi Zhang , Qingqiu Li , Linyi Yang , Yuejie Zhang , Rui Feng , Liang He , Shang Gao , Yue Zhang

Large language models (LLMs) are increasingly used as simulated participants in social science experiments, but their behavior is often unstable and highly sensitive to design choices. Prior evaluations frequently conflate base-model…

Artificial Intelligence · Computer Science 2026-02-03 Xuan Liu , Haoyang Shang , Zizhang Liu , Xinyan Liu , Yunze Xiao , Yiwen Tu , Haojian Jin

Large language models have recently been proposed as tools for automated essay scoring, but their agreement with human grading remains unclear. In this work, we evaluate how LLM-generated scores compare with human grades and analyze the…

Artificial Intelligence · Computer Science 2026-03-26 Jerin George Mathew , Sumayya Taher , Anindita Kundu , Denilson Barbosa

As large language models (LLMs) become more capable and agentic, the requirement for trust in their outputs grows significantly, yet at the same time concerns have been mounting that models may learn to lie in pursuit of their goals. To…

Large language models~(LLMs) have greatly advanced the frontiers of artificial intelligence, attaining remarkable improvement in model capacity. To assess the model performance, a typical approach is to construct evaluation benchmarks for…

Computation and Language · Computer Science 2023-11-06 Kun Zhou , Yutao Zhu , Zhipeng Chen , Wentong Chen , Wayne Xin Zhao , Xu Chen , Yankai Lin , Ji-Rong Wen , Jiawei Han

The field of artificial intelligence (AI) alignment aims to investigate whether AI technologies align with human interests and values and function in a safe and ethical manner. AI alignment is particularly relevant for large language models…

Human-Computer Interaction · Computer Science 2023-01-18 Thilo Hagendorff , Sarah Fabi

The explosion of high-performing conversational language models (LMs) has spurred a shift from classic natural language processing (NLP) benchmarks to expensive, time-consuming and noisy human evaluations - yet the relationship between…

The rapid adoption of large language models (LLMs) in education raises profound challenges for assessment design. To adapt assessments to the presence of LLM-based tools, it is crucial to characterize the strengths and weaknesses of LLMs in…

Human-Computer Interaction · Computer Science 2026-04-16 Licol Zeinfeld , Alona Strugatski , Ziva Bar-Dov , Ron Blonder , Shelley Rap , Giora Alexandron

The versatility of Large Language Models (LLMs) on natural language understanding tasks has made them popular for research in social sciences. To properly understand the properties and innate personas of LLMs, researchers have performed…

Computation and Language · Computer Science 2024-04-03 Bangzhao Shu , Lechen Zhang , Minje Choi , Lavinia Dunagan , Lajanugen Logeswaran , Moontae Lee , Dallas Card , David Jurgens

We present this article as a small gesture in an attempt to counter what appears to be exponentially growing hype around Artificial Intelligence (AI) and its capabilities, and the distraction provided by the associated talk of…

Computation and Language · Computer Science 2023-07-12 Michael O'Neill , Mark Connor

This paper introduces an approach to increasing the explainability of artificial intelligence (AI) systems by embedding Large Language Models (LLMs) within standardized analytical processes. While traditional explainable AI (XAI) methods…

Artificial Intelligence · Computer Science 2025-11-11 Marc Jansen , Marcel Pehlke

Achieving human-like perception and reasoning in Multimodal Large Language Models (MLLMs) remains a central challenge in artificial intelligence. While recent research has primarily focused on enhancing reasoning capabilities in MLLMs, a…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Hongcheng Gao , Zihao Huang , Lin Xu , Jingyi Tang , Xinhao Li , Yue Liu , Haoyang Li , Taihang Hu , Minhua Lin , Xinlong Yang , Ge Wu , Balong Bi , Hongyu Chen , Wentao Zhang

This paper addresses the conceptual, methodological and technical challenges in studying large language models (LLMs) and the texts they produce from a quantitative linguistics perspective. It builds on a theoretical framework that…

Computation and Language · Computer Science 2024-08-30 Jiří Milička

Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators. However, imperfect sensitivity and specificity of the LLM judges induce bias in naive evaluation scores. We propose a simple…

Machine Learning · Computer Science 2026-02-10 Chungpa Lee , Thomas Zeng , Jongwon Jeong , Jy-yong Sohn , Kangwook Lee

As Large Language Models (LLMs) are integrated with human daily applications rapidly, many societal and ethical concerns are raised regarding the behavior of LLMs. One of the ways to comprehend LLMs' behavior is to analyze their…

Computation and Language · Computer Science 2024-02-23 Xiaoyang Song , Yuta Adachi , Jessie Feng , Mouwei Lin , Linhao Yu , Frank Li , Akshat Gupta , Gopala Anumanchipalli , Simerjot Kaur

Grading assessments is time-consuming and prone to human bias. Students may experience delays in receiving feedback that may not be tailored to their expectations or needs. Harnessing AI in education can be effective for grading…

Physics Education · Physics 2025-12-01 Ryan Mok , Faraaz Akhtar , Louis Clare , Christine Li , Jun Ida , Lewis Ross , Mario Campanelli

Effective and safe human-machine collaboration requires the regulated and meaningful exchange of emotions between humans and artificial intelligence (AI). Current AI systems based on large language models (LLMs) can provide feedback that…

Computation and Language · Computer Science 2025-06-18 Xiuwen Wu , Hao Wang , Zhiang Yan , Xiaohan Tang , Pengfei Xu , Wai-Ting Siok , Ping Li , Jia-Hong Gao , Bingjiang Lyu , Lang Qin

Given the remarkable performance of Large Language Models (LLMs), an important question arises: Can LLMs conduct human-like scientific research and discover new knowledge, and act as an AI scientist? Scientific discovery is an iterative…

Machine Learning · Computer Science 2025-02-24 Tingting Chen , Srinivas Anumasa , Beibei Lin , Vedant Shah , Anirudh Goyal , Dianbo Liu

There is an increasing trend towards evaluating NLP models with LLMs instead of human judgments, raising questions about the validity of these evaluations, as well as their reproducibility in the case of proprietary models. We provide…

Large Language Models (LLMs) created new opportunities for generating personas, expected to streamline and accelerate the human-centered design process. Yet, AI-generated personas may not accurately represent actual user experiences, as…

Human-Computer Interaction · Computer Science 2025-08-06 Christopher Lazik , Christopher Katins , Charlotte Kauter , Jonas Jakob , Caroline Jay , Lars Grunske , Thomas Kosch
‹ Prev 1 8 9 10 Next ›