English
Related papers

Related papers: Self-Preference Bias in Rubric-Based Evaluation of…

200 papers

Prompt sensitivity, referring to the phenomenon where paraphrasing (i.e., repeating something written or spoken using different words) leads to significant changes in large language model (LLM) performance, has been widely accepted as a…

Computation and Language · Computer Science 2025-09-03 Andong Hua , Kenan Tang , Chenhe Gu , Jindong Gu , Eric Wong , Yao Qin

While astonishingly capable, large Language Models (LLM) can sometimes produce outputs that deviate from human expectations. Such deviations necessitate an alignment phase to prevent disseminating untruthful, toxic, or biased information.…

Artificial Intelligence · Computer Science 2024-10-30 Long Tan Le , Han Shu , Tung-Anh Nguyen , Choong Seon Hong , Nguyen H. Tran

Automatic evaluation of generated textual content presents an ongoing challenge within the field of NLP. Given the impressive capabilities of modern language models (LMs) across diverse NLP tasks, there is a growing trend to employ these…

Computation and Language · Computer Science 2024-06-10 Yiqi Liu , Nafise Sadat Moosavi , Chenghua Lin

Cognitive biases are systematic deviations in thinking that lead to irrational judgments and problematic decision-making, extensively studied across various fields. Recently, large language models (LLMs) have shown advanced understanding…

Computation and Language · Computer Science 2024-10-10 Nuo Chen , Jiqun Liu , Xiaoyu Dong , Qijiong Liu , Tetsuya Sakai , Xiao-Ming Wu

While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their problem-solving abilities depend on the context and textual formatting. We introduce the Robust Reasoning Benchmark (RRB), a pipeline of…

Machine Learning · Computer Science 2026-05-22 Pavel Golikov , Evgenii Opryshko , Gennady Pekhimenko , Mark C. Jeffrey

Effective personalized feedback is critical to students' literacy development. Though LLM-powered tools now promise to automate such feedback at scale, LLMs are not language-neutral: they privilege standard academic English and reproduce…

Computation and Language · Computer Science 2026-03-16 Mei Tan , Lena Phalen , Dorottya Demszky

As large language models (LLMs) continue to advance, reliable evaluation methods are essential particularly for open-ended, instruction-following tasks. LLM-as-a-Judge enables automatic evaluation using LLMs as evaluators, but its…

Computation and Language · Computer Science 2025-06-17 Yusuke Yamauchi , Taro Yano , Masafumi Oyamada

The evaluation of large language model (LLM) outputs is increasingly performed by other LLMs, a setup commonly known as "LLM-as-a-judge", or autograders. While autograders offer a scalable alternative to human evaluation, they have shown…

Machine Learning · Computer Science 2026-02-27 Magda Dubois , Harry Coppock , Mario Giulianelli , Timo Flesch , Lennart Luettgau , Cozmin Ududec

Expert domain writing, such as scientific writing, typically demands extensive domain knowledge. Although large language models (LLMs) show promising potential in this task, evaluating the quality of automatically generated scientific…

Computation and Language · Computer Science 2026-01-12 Furkan Şahinuç , Subhabrata Dutta , Iryna Gurevych

Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available. The scale and versatility of such evaluations make the use of LLM-based…

Computation and Language · Computer Science 2025-06-11 Ariel Gera , Odellia Boni , Yotam Perlitz , Roy Bar-Haim , Lilach Eden , Asaf Yehudai

Rubrics have been extensively utilized for evaluating unverifiable, open-ended tasks, with recent research incorporating them into reward systems for reinforcement learning. However, existing frameworks typically treat rubrics only as…

Computation and Language · Computer Science 2026-05-11 Jiachen Yu , Zhihao Xu , Junjie Wang , Yujiu Yang

Large Language Models (LLMs) are increasingly used to evaluate information retrieval (IR) systems, generating relevance judgments traditionally made by human assessors. Recent empirical studies suggest that LLM-based evaluations often align…

Information Retrieval · Computer Science 2026-01-21 Laura Dietz , Oleg Zendel , Peter Bailey , Charles Clarke , Ellese Cotterill , Jeff Dalton , Faegheh Hasibi , Mark Sanderson , Nick Craswell

Large Language Models (LLMs) as judges and LLM-based data synthesis have emerged as two fundamental LLM-driven data annotation methods in model development. While their combination significantly enhances the efficiency of model training and…

Machine Learning · Computer Science 2026-03-05 Dawei Li , Renliang Sun , Yue Huang , Ming Zhong , Bohan Jiang , Jiawei Han , Xiangliang Zhang , Wei Wang , Huan Liu

Recent advances in reasoning with large language models (LLMs) have demonstrated strong performance on complex mathematical tasks, including combinatorial optimization. Techniques such as Chain-of-Thought and In-Context Learning have…

Artificial Intelligence · Computer Science 2025-09-17 Marylou Fauchard , Florian Carichon , Margarida Carvalho , Golnoosh Farnadi

The use of language models for automatically evaluating long-form text (LLM-as-a-judge) is becoming increasingly common, yet most LLM judges are optimized exclusively for English, with strategies for enhancing their multilingual evaluation…

Computation and Language · Computer Science 2025-10-31 José Pombal , Dongkeun Yoon , Patrick Fernandes , Ian Wu , Seungone Kim , Ricardo Rei , Graham Neubig , André F. T. Martins

Large multimodal models (LMMs) are increasingly adopted as judges in multimodal evaluation systems due to their strong instruction following and consistency with human preferences. However, their ability to follow diverse, fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Tianyi Xiong , Yi Ge , Ming Li , Zuolong Zhang , Pranav Kulkarni , Kaishen Wang , Qi He , Zeying Zhu , Chenxi Liu , Ruibo Chen , Tong Zheng , Yanshuo Chen , Xiyao Wang , Renrui Zhang , Wenhu Chen , Heng Huang

Recent research has explored LLMs as scalable tools for relevance labeling, but studies indicate they are susceptible to priming effects, where prior relevance judgments influence later ones. Although psychological theories link personality…

Computation and Language · Computer Science 2025-12-02 Nuo Chen , Hanpei Fang , Jiqun Liu , Wilson Wei , Tetsuya Sakai , Xiao-Ming Wu

Large language models (LLMs) are increasingly applied to clinical decision-making. However, their potential to exhibit bias poses significant risks to clinical equity. Currently, there is a lack of benchmarks that systematically evaluate…

Computation and Language · Computer Science 2024-11-18 Yubo Zhang , Shudi Hou , Mingyu Derek Ma , Wei Wang , Muhao Chen , Jieyu Zhao

Large Reasoning Models (LRMs) like DeepSeek-R1 and OpenAI-o1 have demonstrated remarkable reasoning capabilities, raising important questions about their biases in LLM-as-a-judge settings. We present a comprehensive benchmark comparing…

Computers and Society · Computer Science 2025-04-21 Qian Wang , Zhanzhi Lou , Zhenheng Tang , Nuo Chen , Xuandong Zhao , Wenxuan Zhang , Dawn Song , Bingsheng He

Large language models (LLMs) are increasingly used as automated evaluators, yet prior works demonstrate that these LLM judges often lack consistency in scoring when the prompt is altered. However, the effect of the grading scale itself…