English
Related papers

Related papers: Beyond the Illusion of Consensus: From Surface Heu…

200 papers

LLMs have emerged as powerful evaluators in the LLM-as-a-Judge paradigm, offering significant efficiency and flexibility compared to human judgments. However, previous methods primarily rely on single-point evaluations, overlooking the…

Artificial Intelligence · Computer Science 2025-05-20 Luyu Chen , Zeyu Zhang , Haoran Tan , Quanyu Dai , Hao Yang , Zhenhua Dong , Xu Chen

Large language models (LLMs) can generate persuasive narratives at scale, raising concerns about their potential use in disinformation campaigns. Assessing this risk ultimately requires understanding how readers receive such content. In…

Artificial Intelligence · Computer Science 2026-04-09 Zonghuan Xu , Xiang Zheng , Yutao Wu , Xingjun Ma

LLM-as-judge systems promise scalable, consistent evaluation. We find the opposite: judges are consistent, but not with each other; they are consistent with themselves. Across 3,240 evaluations (9 judges x 120 unique video x pack items x 3…

Artificial Intelligence · Computer Science 2026-01-09 Wajid Nasser

Large Language Models (LLMs) are increasingly embedded in evaluative processes, from information filtering to assessing and addressing knowledge gaps through explanation and credibility judgments. This raises the need to examine how such…

LLM-as-a-judge approaches have emerged as a scalable solution for evaluating model behaviors, yet they rely on evaluation criteria often created by a single individual, embedding that person's assumptions, priorities, and interpretive lens.…

Human-Computer Interaction · Computer Science 2026-04-30 Charles Chiang , Simret Gebreegziabher , Annalisa Szymanski , Yukun Yang , Hyo Jin Do , Zahra Ashktorab , Werner Geyer , Toby Li , Diego Gomez-Zara

Evaluating large language models (LLMs) on open-ended tasks without ground-truth labels is increasingly done via the LLM-as-a-judge paradigm. A critical but under-modeled issue is that judge LLMs differ substantially in reliability;…

Machine Learning · Statistics 2026-01-30 Mingyuan Xu , Xinzi Tan , Jiawei Wu , Doudou Zhou

New Large Language Models (LLMs) become available every few weeks, and modern application developers confronted with the unenviable task of having to decide if they should switch to a new model. While human evaluation remains the gold…

Artificial Intelligence · Computer Science 2025-12-25 Suryaansh Jain , Umair Z. Ahmed , Shubham Sahai , Ben Leong

Offering a promising solution to the scalability challenges associated with human evaluation, the LLM-as-a-judge paradigm is rapidly gaining traction as an approach to evaluating large language models (LLMs). However, there are still many…

Computation and Language · Computer Science 2025-08-19 Aman Singh Thakur , Kartik Choudhary , Venkat Srinik Ramayapally , Sankaran Vaidyanathan , Dieuwke Hupkes

The growing scale of evaluation tasks has led to the widespread adoption of automated evaluation using LLMs, a paradigm known as "LLM-as-a-judge". However, improving its alignment with human preferences without complex prompts or…

Computation and Language · Computer Science 2025-10-17 Peng Lai , Jianjie Zheng , Sijie Cheng , Yun Chen , Peng Li , Yang Liu , Guanhua Chen

Large language models (LLMs) are increasingly used as automated evaluators (LLM-as-a-Judge). This work challenges its reliability by showing that trust judgments by LLMs are biased by disclosed source labels. Using a counterfactual design,…

Artificial Intelligence · Computer Science 2026-04-08 Xin Sun , Di Wu , Sijing Qin , Isao Echizen , Abdallah El Ali , Saku Sugawara

LLMs are increasingly employed both as judges for evaluating open-ended outputs and as co-creation partners in AI-assisted programming; yet rigorous evaluation in human-AI co-creation settings remains underdeveloped as judgments must be…

Software Engineering · Computer Science 2026-05-01 Md Faizul Ibne Amin , Yutaka Watanobe , Daniel M. Muepu , Haruto Suzuki , Kenta Nanaumi , Md Mostafizer Rahman

Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with human scoring standards remains challenging. We formulate this challenge as a criteria-transfer…

Computation and Language · Computer Science 2026-05-29 Yihan Hong , Huaiyuan Yao , Bolin Shen , Wanpeng Xu , Hua Wei , Yushun Dong

The adoption of Large Language Models (LLMs) as automated evaluators (LLM-as-a-judge) has revealed critical inconsistencies in current evaluation frameworks. We identify two fundamental types of inconsistencies: (1) Score-Comparison…

Artificial Intelligence · Computer Science 2025-09-29 Yidong Wang , Yunze Song , Tingyuan Zhu , Xuanwang Zhang , Zhuohao Yu , Hao Chen , Chiyu Song , Qiufeng Wang , Cunxiang Wang , Zhen Wu , Xinyu Dai , Yue Zhang , Wei Ye , Shikun Zhang

Accurate and consistent evaluation is crucial for decision-making across numerous fields, yet it remains a challenging task due to inherent subjectivity, variability, and scale. Large Language Models (LLMs) have achieved remarkable success…

This research introduces the Judge's Verdict Benchmark, a novel two-step methodology to evaluate Large Language Models (LLMs) as judges for response accuracy evaluation tasks. We assess how well 54 LLMs can replicate human judgment when…

Computation and Language · Computer Science 2025-10-14 Steve Han , Gilberto Titericz Junior , Tom Balough , Wenfei Zhou

The $\textit{LLM-as-a-judge}$ paradigm has become the operational backbone of automated AI evaluation pipelines, yet rests on an unverified assumption: that judges evaluate text strictly on its semantic content, impervious to surrounding…

Artificial Intelligence · Computer Science 2026-04-17 Manan Gupta , Inderjeet Nair , Lu Wang , Dhruv Kumar

Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse use cases. In practice, however, evaluating generative outputs typically relies on rigid lexical…

Computation and Language · Computer Science 2026-04-13 Hippolyte Gisserot-Boukhlef , Nicolas Boizard , Emmanuel Malherbe , Céline Hudelot , Pierre Colombo

LLM-as-a-judge has become a promising paradigm for using large language models (LLMs) to evaluate natural language generation (NLG), but the uncertainty of its evaluation remains underexplored. This lack of reliability may limit its…

Computation and Language · Computer Science 2025-09-24 Huanxin Sheng , Xinyi Liu , Hangfeng He , Jieyu Zhao , Jian Kang

Automated short-answer grading (ASAG) remains a challenging task due to the linguistic variability of student responses and the need for nuanced, rubric-aligned partial credit. While Large Language Models (LLMs) offer a promising solution,…

Computation and Language · Computer Science 2026-01-15 Haotian Deng , Chris Farber , Jiyoon Lee , David Tang

Rubric-based evaluation has become a prevailing paradigm for evaluating instruction following in large language models (LLMs). Despite its widespread use, the reliability of these rubric-level evaluations remains unclear, calling for…

Artificial Intelligence · Computer Science 2026-03-27 Tianjun Pan , Xuan Lin , Wenyan Yang , Qianyu He , Shisong Chen , Licai Qi , Wanqing Xu , Hongwei Feng , Bo Xu , Yanghua Xiao
‹ Prev 1 2 3 10 Next ›