English
Related papers

Related papers: Using Elo Rating as a Metric for Comparative Judge…

200 papers

Equity is a core concern of learning analytics. However, applications that teach and assess equity skills, particularly at scale are lacking, often due to barriers in evaluating language. Advances in generative AI via large language models…

Human-Computer Interaction · Computer Science 2024-12-17 Danielle R. Thomas , Conrad Borchers , Sanjit Kakarla , Jionghao Lin , Shambhavi Bhushan , Boyuan Guo , Erin Gatz , Kenneth R. Koedinger

Reward models (RMs) play a crucial role in Reinforcement Learning from Human Feedback by serving as proxies for human preferences in aligning large language models. However, they suffer from various biases which could lead to reward…

Artificial Intelligence · Computer Science 2026-03-18 Xiao Zhu , Chenmien Tan , Pinzhen Chen , Rico Sennrich , Huiming Wang , Yanlin Zhang , Hanxu Hu

A broad current application of algorithms is in formal and quantitative measures of murky concepts -- like merit -- to make decisions. When people strategically respond to these sorts of evaluations in order to gain favorable decision…

Computers and Society · Computer Science 2023-10-06 Benjamin Laufer , Jon Kleinberg , Karen Levy , Helen Nissenbaum

ECHO (Evaluation of Chat, Human behavior, and Outcomes) is an open research platform designed to support reproducible, mixed-method studies of human interaction with both conversational AI systems and Web search engines. It enables…

Human-Computer Interaction · Computer Science 2026-02-12 Jiqun Liu , Nischal Dinesh , Ran Yu

Large language models (LLMs) have demonstrated remarkable advancements and have attracted significant efforts to develop LLMs into agents capable of executing intricate multi-step decision-making tasks beyond traditional NLP applications.…

Computation and Language · Computer Science 2025-06-10 Yining Ye , Xin Cong , Shizuo Tian , Yujia Qin , Chong Liu , Yankai Lin , Zhiyuan Liu , Maosong Sun

Reasoning is not just about solving problems -- it is also about evaluating which problems are worth solving at all. Evaluations of artificial intelligence (AI) systems primarily focused on problem solving, historically by studying how…

Comparative Judgement is an assessment method where item ratings are estimated based on rankings of subsets of the items. These rankings are typically pairwise, with ratings taken to be the estimated parameters from fitting a Bradley-Terry…

Methodology · Statistics 2024-05-22 Ian Hamilton , Nick Tawn

In peer review, reviewers are usually asked to provide scores for the papers. The scores are then used by Area Chairs or Program Chairs in various ways in the decision-making process. The scores are usually elicited in a quantized form to…

Information Retrieval · Computer Science 2022-04-13 Yusha Liu , Yichong Xu , Nihar B. Shah , Aarti Singh

In this work, we deal with the problem of rating in sports, where the skills of the players/teams are inferred from the observed outcomes of the games. Our focus is on the online rating algorithms which estimate the skills after each new…

Machine Learning · Statistics 2021-04-30 Leszek Szczecinski , Raphaëlle Tihon

Online Judge (OJ) systems are typically considered within programming-related courses as they yield fast and objective assessments of the code developed by the students. Such an evaluation generally provides a single decision based on a…

Computers and Society · Computer Science 2024-02-07 Juan Ramón Rico-Juan , Víctor M. Sánchez-Cartagena , Jose J. Valero-Mas , Antonio Javier Gallego

Speech emotions play a crucial role in human-computer interaction, shaping engagement and context-aware communication. Despite recent advances in spoken dialogue systems, a holistic system for evaluating emotional reasoning is still…

Large language models are often used as judges to score candidate responses, then validated with a single global metric such as correlation with reference labels. This can be misleading when the real deployment task is best-of-n selection…

Machine Learning · Computer Science 2026-03-16 Eddie Landesberg

Large language models (LLM) have shown remarkable abilities in text generation, question answering, language translation, reasoning and many other tasks. It continues to advance rapidly and is becoming increasingly influential in various…

Artificial Intelligence · Computer Science 2025-01-31 Yinqi Zhang , Xintian Han , Haolong Li , Kedi Chen , Shaohui Lin

Sentiment analysis AKA opinion mining is one of the most widely used NLP applications to identify human intentions from their reviews. In the education sector, opinion mining is used to listen to student opinions and enhance their…

Computation and Language · Computer Science 2023-02-10 Thanveer Shaik , Xiaohui Tao , Christopher Dann , Haoran Xie , Yan Li , Linda Galligan

Large Language Models (LLMs) are increasingly explored for educational tasks such as grading, yet their alignment with human evaluation in real classrooms remains underexamined. In this study, we investigate the feasibility of using an LLM…

Computation and Language · Computer Science 2025-11-19 Grace Byun , Swati Rajwal , Jinho D. Choi

Benchmarks for the evaluation of model performance play an important role in machine learning. However, there is no established way to describe and create new benchmarks. What is more, the most common benchmarks use performance measures…

Machine Learning · Computer Science 2022-09-23 Alicja Gosiewska , Katarzyna Woźnica , Przemysław Biecek

Prediction and modelling of competitive sports outcomes has received much recent attention, especially from the Bayesian statistics and machine learning communities. In the real world setting of outcome prediction, the seminal \'{E}l\H{o}…

Machine Learning · Statistics 2017-01-30 Franz J. Király , Zhaozhi Qian

We present the Multilingual Entity Linking of Occupations (MELO) Benchmark, a new collection of 48 datasets for evaluating the linking of entity mentions in 21 languages to the ESCO Occupations multilingual taxonomy. MELO was built using…

Computation and Language · Computer Science 2024-10-14 Federico Retyk , Luis Gasco , Casimiro Pio Carrino , Daniel Deniz , Rabih Zbib

With the accelerating development of Large Language Models (LLMs), many LLMs are beginning to be used in the Chinese K-12 education domain. The integration of LLMs and education is getting closer and closer, however, there is currently no…

Computation and Language · Computer Science 2024-01-30 Jinchang Hou , Chang Ao , Haihong Wu , Xiangtao Kong , Zhigang Zheng , Daijia Tang , Chengming Li , Xiping Hu , Ruifeng Xu , Shiwen Ni , Min Yang

Critique-guided reinforcement learning (RL) has emerged as a powerful paradigm for training LLM agents by augmenting sparse outcome rewards with natural-language feedback. However, current methods often rely on static or offline critic…

Artificial Intelligence · Computer Science 2026-04-15 Zhicong Li , Lingjie Jiang , Yulan Hu , Xingchen Zeng , Yixia Li , Xiangwen Zhang , Guanhua Chen , Zheng Pan , Xin Li , Yong Liu
‹ Prev 1 3 4 5 6 7 10 Next ›