English
Related papers

Related papers: Debate Helps Weak Judges Reward Stronger Models

200 papers

In multi-agent debate (MAD) systems, performance gains are often reported; however, because the debate protocol (e.g., number of agents, rounds, and aggregation rule) is typically held fixed while model-related factors vary, it is difficult…

Multiagent Systems · Computer Science 2026-04-01 Ramtin Zargari Marandi

The remarkable growth in large language model (LLM) capabilities has spurred exploration into multi-agent systems, with debate frameworks emerging as a promising avenue for enhanced problem-solving. These multi-agent debate (MAD)…

Artificial Intelligence · Computer Science 2025-06-23 Yongjin Yang , Euiin Yi , Jongwoo Ko , Kimin Lee , Zhijing Jin , Se-Young Yun

Reasoning LLMs-as-Judges, which can benefit from inference-time scaling, provide a promising path for extending the success of reasoning models to non-verifiable domains where the output correctness/quality cannot be directly checked.…

Artificial Intelligence · Computer Science 2026-03-13 Yixin Liu , Yue Yu , DiJia Su , Sid Wang , Xuewei Wang , Song Jiang , Bo Liu , Arman Cohan , Yuandong Tian , Zhengxing Chen

How should two language models interact to produce better code than either can alone? The conventional approach -- a reasoning model plans, a code specialist implements -- seems natural but fails: on HumanEval+, plan-then-code degrades…

Software Engineering · Computer Science 2026-03-05 Jan Miller

As Large Language Models (LLMs) are increasingly adopted as automated judges in benchmarking and reward modeling, ensuring their reliability, efficiency, and robustness has become critical. In this work, we present a systematic comparison…

Artificial Intelligence · Computer Science 2026-05-12 Pratik Jayarao , Himanshu Gupta , Neeraj Varshney , Chaitanya Dwivedi

Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost. Studies…

Computation and Language · Computer Science 2026-01-29 Xiaochen Zhu , Caiqi Zhang , Yizhou Chi , Tom Stafford , Nigel Collier , Andreas Vlachos

We consider the classic veto bargaining model but allow the agenda setter to engage in persuasion to convince the veto player to approve her proposal. We fully characterize the optimal proposal and experiment when Vetoer has quadratic loss,…

Theoretical Economics · Economics 2023-10-23 Jenny S Kim , Kyungmin Kim , Richard Van Weelden

An informed Advisor and an uninformed Decision-Maker, with conflicting interests, engage in repeated cheap talk communication in always new decision problems. While the Decision-Maker's optimal payoff is attainable in some subgame-perfect…

Theoretical Economics · Economics 2025-08-04 Steven Kivinen , Christoph Kuzmics

Multiagent collaboration has emerged as a promising framework for enhancing the reasoning capabilities of large language models (LLMs). Despite improvements in reasoning, the approach introduces substantial computational overhead resulting…

Artificial Intelligence · Computer Science 2025-05-21 Sugyeong Eo , Hyeonseok Moon , Evelyn Hayoon Zi , Chanjun Park , Heuiseok Lim

The use of synthetic data has played a critical role in recent state-of-art breakthroughs. However, overly relying on a single oracle teacher model to generate data has been shown to lead to model collapse and invite propagation of biases.…

Computation and Language · Computer Science 2024-08-28 Ayomide Odumakinde , Daniel D'souza , Pat Verga , Beyza Ermis , Sara Hooker

Accurate detection of errors in large language models (LLM) responses is central to the success of scalable oversight, or providing effective supervision to superhuman intelligence. Yet, self-diagnosis is often unreliable on complex tasks…

Machine Learning · Computer Science 2025-10-27 Yongqiang Chen , Gang Niu , James Cheng , Bo Han , Masashi Sugiyama

We study the helpful product reviews identification problem in this paper. We observe that the evidence-conclusion discourse relations, also known as arguments, often appear in product reviews, and we hypothesise that some argument-based…

Computation and Language · Computer Science 2017-07-25 Haijing Liu , Yang Gao , Pin Lv , Mengxue Li , Shiqiang Geng , Minglan Li , Hao Wang

Philosophical accounts of persuasion often assume that shared evidence and rational argumentation should lead to a convergence of views between peers, yet everyday discourse often suggests otherwise. In this study, we use large language…

Computation and Language · Computer Science 2026-05-12 David Freeborn , Malihe Alikani , Anthony Sicilia

In many settings, an effective way of evaluating objects of interest is to collect evaluations from dispersed individuals and to aggregate these evaluations together. Some examples are categorizing online content and evaluating student…

Computer Science and Game Theory · Computer Science 2016-06-23 Alice Gao , James R. Wright , Kevin Leyton-Brown

As AI agents surpass human capabilities, scalable oversight -- the problem of effectively supplying human feedback to potentially superhuman AI models -- becomes increasingly critical to ensure alignment. While numerous scalable oversight…

Artificial Intelligence · Computer Science 2025-04-08 Abhimanyu Pallavi Sudhir , Jackson Kaunismaa , Arjun Panickssery

Evaluating the capabilities and risks of foundation models is paramount, yet current methods demand extensive domain expertise, hindering their scalability as these models rapidly evolve. We introduce SKATE: a novel evaluation framework in…

Artificial Intelligence · Computer Science 2026-02-13 Dewi S. W. Gould , Bruno Mlodozeniec , Samuel F. Brown

Self-improvement, where models improve beyond their current performance without external supervision, remains a challenge. The core difficulty is sourcing a training signal stronger than what the model itself can currently produce. Majority…

Artificial Intelligence · Computer Science 2026-02-02 Ankur Samanta , Akshayaa Magesh , Runzhe Wu , Ayush Jain , Youliang Yu , Daniel Jiang , Boris Vidolov , Paul Sajda , Yonathan Efroni , Kaveh Hassani

Inspired by e-participation systems, in this paper we propose a new model to represent human debates and methods to obtain collective conclusions from them. This model overcomes drawbacks of existing approaches by allowing users to…

Artificial Intelligence · Computer Science 2020-07-15 Jordi Ganzer , Natalia Criado , Maite Lopez-Sanchez , Simon Parsons , Juan A. Rodriguez-Aguilar

Multi-agent debate (MAD) systems improve LLM reasoning through iterative deliberation, but remain vulnerable to debate collapse, a failure type where final agent decisions are compromised on erroneous reasoning. Existing methods lack…

Multiagent Systems · Computer Science 2026-02-10 Luoxi Tang , Yuqiao Meng , Joseph Costa , Yingxue Zhang , Muchao Ye , Zhaohan Xi

An effective way to obtain different perspectives on any given topic is by conducting a debate, where participants argue for and against the topic. Here, we propose a novel debate framework for understanding and explaining a continuous…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Avinash Kori , Ben Glocker , Francesca Toni