English
Related papers

Related papers: Debate Helps Weak Judges Reward Stronger Models

200 papers

How do algorithmic decision aids introduced in business decision processes affect task performance? In a first experiment, we study effective collaboration. Faced with a decision, subjects alone have a success rate of 72%; Aided by a…

Human-Computer Interaction · Computer Science 2020-09-18 Thomas Baudel , Manon Verbockhaven , Guillaume Roy , Victoire Cousergue , Rida Laarach

As the most public component of the Supreme Court's decision-making process, oral argument receives an out-sized share of attention in the popular media. Despite its prominence, however, the basic function and operation of oral argument as…

Computers and Society · Computer Science 2023-06-09 Gregory M. Dickinson

Opinion modeling aims to capture individual or group political preferences, enabling applications such as digital democracies, where models could help shape fairer and more popular policies. Given their versatility, strong generalization…

Computation and Language · Computer Science 2026-03-13 Frédéric Berdoz , Yann Billeter , Yann Vonlanthen , Roger Wattenhofer

While state-of-the-art language models have achieved impressive results, they remain susceptible to inference-time adversarial attacks, such as adversarial prompts generated by red teams arXiv:2209.07858. One approach proposed to improve…

Computation and Language · Computer Science 2024-01-12 Steffi Chern , Zhen Fan , Andy Liu

In recent years, Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse NLP tasks. Extensive research has explored how to enhance the logical reasoning abilities such as Chain-of-Thought, Chain-of-Thought with…

Computation and Language · Computer Science 2025-12-29 Tongxuan Liu , Xingyu Wang , Weizhe Huang , Wenjiang Xu , Yuting Zeng , Lei Jiang , Hailong Yang , Jing Li

We introduce an approach to topic modelling with document-level covariates that remains tractable in the face of large text corpora. This is achieved by de-emphasizing the role of parameter estimation in an underlying probabilistic model,…

Methodology · Statistics 2025-11-05 Gabriel Phelan , David A. Campbell

One of the better studied properties for operators in judgment aggregation is independence, which essentially dictates that the collective judgment on one issue should not depend on the individual judgments given on some other issue(s) in…

Artificial Intelligence · Computer Science 2016-04-25 Jérôme Lang , Marija Slavkovik , Srdjan Vesic

Although peer code review is widely adopted in both commercial and open source development, existing studies suggest that such code reviews often contain a significant amount of non-useful review comments. Unfortunately, to date, no tools…

Software Engineering · Computer Science 2018-07-13 Mohammad Masudur Rahman , Chanchal K. Roy , Raula G. Kula

This paper reports on empirical work aimed at comparing evidential reasoning techniques. While there is prima facie evidence for some conclusions, this i6 work in progress; the present focus is methodology, with the goal that subsequent…

Artificial Intelligence · Computer Science 2013-04-10 Ronald P. Loui

Deliberation involves participants exchanging knowledge, arguments, and perspectives and has been shown to be effective at addressing polarization. The Stanford Online Deliberation Platform facilitates large-scale deliberations. It enables…

Artificial Intelligence · Computer Science 2024-08-23 Lodewijk Gelauff , Mohak Goyal , Bhargav Dindukurthi , Ashish Goel , Alice Siu

The evaluation of noisy binary classifiers on unlabeled data is treated as a streaming task: given a data sketch of the decisions by an ensemble, estimate the true prevalence of the labels as well as each classifier's accuracy on them. Two…

Machine Learning · Statistics 2023-09-11 Andrés Corrada-Emmanuel

Multi-agent debate improves LLM reasoning, yet agreement among agents is not evidence of correctness. When agents converge on a wrong answer through social reinforcement, consensus-based stopping commits that error to an automated action…

Artificial Intelligence · Computer Science 2026-04-10 Mengdie Flora Wang , Haochen Xie , Guanghui Wang , Aijing Gao , Guang Yang , Ziyuan Li , Qucy Wei Qiu , Fangwei Han , Hengzhi Qiu , Yajing Huang , Bing Zhu , Jae Oh Woo

A proposer requires the approval of a veto player to change a status quo. Preferences are single peaked. Proposer is uncertain about Vetoer's ideal point. We study Proposer's optimal mechanism without transfers. Vetoer is given a menu, or a…

Theoretical Economics · Economics 2022-12-29 Navin Kartik , Andreas Kleiner , Richard Van Weelden

Recent trends in test-time scaling for reasoning models (e.g., OpenAI o1, DeepSeek R1) have led to a popular belief that extending thinking traces using prompts like "Wait" or "Let me rethink" can improve performance. This raises a natural…

We develop models to classify desirable evidence and desirable reasoning revisions in student argumentative writing. We explore two ways to improve classifier performance - using the essay context of the revision, and using the feedback…

Computation and Language · Computer Science 2023-02-13 Tazin Afrin , Diane Litman

How can we construct an automated debate judge to evaluate an extensive, vibrant, multi-turn debate? This task is challenging, as judging a debate involves grappling with lengthy texts, intricate argument relationships, and…

Computation and Language · Computer Science 2024-06-21 Jingcong Liang , Rong Ye , Meng Han , Ruofei Lai , Xinyu Zhang , Xuanjing Huang , Zhongyu Wei

Large language models frequently encounter conflicts between their parametric knowledge and contextual input, often resulting in factual inconsistencies or hallucinations. We propose Self-Reflective Debate for Contextual Reliability…

Computation and Language · Computer Science 2025-06-09 Zeqi Zhou , Fang Wu , Shayan Talaei , Haokai Zhao , Cheng Meixin , Tinson Xu , Amin Saberi , Yejin Choi

Offering a promising solution to the scalability challenges associated with human evaluation, the LLM-as-a-judge paradigm is rapidly gaining traction as an approach to evaluating large language models (LLMs). However, there are still many…

Computation and Language · Computer Science 2025-08-19 Aman Singh Thakur , Kartik Choudhary , Venkat Srinik Ramayapally , Sankaran Vaidyanathan , Dieuwke Hupkes

Large language models (LLMs) with reasoning capabilities have fueled a compelling narrative that reasoning universally improves performance across language tasks. We test this claim through a comprehensive evaluation of 504 configurations…

Computation and Language · Computer Science 2026-03-02 Donghao Huang , Zhaoxia Wang

Pairwise evaluation of Large Language Models (LLMs) is a common paradigm, but it is prone to preference bias, where judges systematically favor certain outputs, such as their own. This bias leads to inconsistent and skewed rankings across…

Artificial Intelligence · Computer Science 2025-11-18 Yang Zhang , Cunxiang Wang , Lindong Wu , Wenbo Yu , Yidong Wang , Guangsheng Bao , Jie Tang