Computation and Language · Computer Science
Exploring the Effects of Alignment on Numerical Bias in Large Language Models
Ayako Sato, Hwichan Kim, Zhousi Chen, Masato Mita +1
2026-01-27
Computation and Language · Computer Science
Mitigating the Bias of Large Language Model Evaluation
Hongli Zhou, Hui Huang, Yunfei Long, Bing Xu +4
2024-09-26
Computation and Language · Computer Science
Evaluating Scoring Bias in LLM-as-a-Judge
Qingquan Li, Shaoyu Dou, Kailai Shao, Chao Chen +1
2026-05-22
Artificial Intelligence · Computer Science
Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
Suryaansh Jain, Umair Z. Ahmed, Shubham Sahai, Ben Leong
2025-12-25
Computation and Language · Computer Science
Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models
Shachi H Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur +5
2024-08-08
Artificial Intelligence · Computer Science
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
Jiaxin Gao, Chen Chen, Yanwen Jia, Xueluan Gong +2
2026-03-03
Computation and Language · Computer Science
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
Weiyue Li, Minda Zhao, Weixuan Dong, Jiahui Cai +11
2026-01-08
Computation and Language · Computer Science
Enhancing Contextual Understanding in Large Language Models through Contrastive Decoding
Zheng Zhao, Emilio Monti, Jens Lehmann, Haytham Assem
2024-05-07
Machine Learning · Computer Science
How to Correctly Report LLM-as-a-Judge Evaluations
Chungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn +1
2026-02-10
Computation and Language · Computer Science
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan +1
2025-08-19
Computation and Language · Computer Science
Aligning Black-box Language Models with Human Judgments
Gerrit J. J. van den Burg, Gen Suzuki, Wei Liu, Murat Sensoy
2025-02-10
Computation and Language · Computer Science
Who can we trust? LLM-as-a-jury for Comparative Assessment
Mengjie Qian, Guangzhi Sun, Mark J. F. Gales, Kate M. Knill
2026-05-29
Computation and Language · Computer Science
Reasons to Reject? Aligning Language Models with Judgments
Weiwen Xu, Deng Cai, Zhisong Zhang, Wai Lam +1
2024-06-07
Machine Learning · Computer Science
Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment
Gregor Bachmann, Sotiris Anagnostidis, Albert Pumarola, Markos Georgopoulos +5
2025-02-03
Machine Learning · Computer Science
Interpreting Language Reward Models via Contrastive Explanations
Junqi Jiang, Tom Bewley, Saumitra Mishra, Freddy Lecue +1
2025-02-27
Computation and Language · Computer Science
Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction
Huanxin Sheng, Xinyi Liu, Hangfeng He, Jieyu Zhao +1
2025-09-24
Computation and Language · Computer Science
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
Xiao Wang, Daniil Larionov, Siwei Wu, Yiqi Liu +3
2025-11-25
Computation and Language · Computer Science
When Wording Steers the Evaluation: Framing Bias in LLM judges
Yerin Hwang, Dongryeol Lee, Taegwan Kang, Minwoo Lee +1
2026-01-21
Computation and Language · Computer Science
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
Xiaolin Zhou, Zheng Luo, Yicheng Gao, Qixuan Chen +3
2026-01-21