English
Related papers

Related papers: On scalable oversight with weak LLMs judging stron…

200 papers

The field of AI research is advancing at an unprecedented pace, enabling automated hypothesis generation and experimental design across diverse domains such as biology, mathematics, and artificial intelligence. Despite these advancements,…

Machine Learning · Computer Science 2025-10-07 Yaowenqi Liu , Bingxu Meng , Rui Pan , Yuxing Liu , Jerry Huang , Jiaxuan You , Tong Zhang

Commonsense reasoning is a difficult task for a computer, but a critical skill for an artificial intelligence (AI). It can enhance the explainability of AI models by enabling them to provide intuitive and human-like explanations for their…

Artificial Intelligence · Computer Science 2024-07-08 Stefanie Krause , Frieder Stolzenburg

AI-assisted decision making becomes increasingly prevalent, yet individuals often fail to utilize AI-based decision aids appropriately especially when the AI explanations are absent, potentially as they do not %understand reflect on AI's…

Human-Computer Interaction · Computer Science 2025-02-18 Zhuoyan Li , Hangxiao Zhu , Zhuoran Lu , Ziang Xiao , Ming Yin

Smart recommendation algorithms have revolutionized content delivery and improved efficiency across various domains. However, concerns about user agency arise from the algorithms' inherent opacity (information asymmetry) and one-way output…

Human-Computer Interaction · Computer Science 2025-05-13 Mengke Wu , Weizi Liu , Yanyun Wang , Mike Yao

Large language models (LLMs) have shown promising accuracy in predicting survey responses and policy preferences, which has increased interest in their potential to represent human interests in various domains. Most existing research has…

Computers and Society · Computer Science 2025-11-18 Suyash Fulay , Jocelyn Zhu , Michiel Bakker

How can we construct an automated debate judge to evaluate an extensive, vibrant, multi-turn debate? This task is challenging, as judging a debate involves grappling with lengthy texts, intricate argument relationships, and…

Computation and Language · Computer Science 2024-06-21 Jingcong Liang , Rong Ye , Meng Han , Ruofei Lai , Xinyu Zhang , Xuanjing Huang , Zhongyu Wei

Just as people improve decision-making by consulting diverse human advisors, they can now also consult with multiple AI systems. Prior work on group decision-making shows that advice aggregation creates pressure to conform, leading to…

Human-Computer Interaction · Computer Science 2026-03-24 Yuta Tsuchiya , Yukino Baba

With the advance of Artificial Intelligence (AI), Large Language Models (LLMs) have gained prominence and been applied in diverse contexts. As they evolve into more sophisticated versions, it is essential to assess whether they reproduce…

Computation and Language · Computer Science 2025-08-15 Gustavo Bonil , Simone Hashiguti , Jhessica Silva , João Gondim , Helena Maia , Nádia Silva , Helio Pedrini , Sandra Avila

Multi-Agent Debate~(MAD) has emerged as a promising paradigm for improving the performance of large language models through collaborative reasoning. Despite recent advances, the key factors driving MAD's effectiveness remain unclear. In…

Computation and Language · Computer Science 2025-10-24 Hyeong Kyu Choi , Xiaojin Zhu , Sharon Li

We investigate how peer pressure influences the opinions of Large Language Model (LLM) agents across a spectrum of cognitive commitments by embedding them in social networks where they update opinions based on peer perspectives. Our…

Computers and Society · Computer Science 2025-10-23 Aliakbar Mehdizadeh , Martin Hilbert

The rapid advancements in large Language models (LLMs) have significantly enhanced their reasoning capabilities, driven by various strategies such as multi-agent collaboration. However, unlike the well-established performance improvements…

Artificial Intelligence · Computer Science 2026-04-23 Zihan Chen , Song Wang , Zhen Tan , Xingbo Fu , Zhenyu Lei , Peng Wang , Huan Liu , Cong Shen , Jundong Li

Large Language Models (LLMs) often exhibit sycophancy, distorting responses to align with user beliefs, notably by readily agreeing with user counterarguments. Paradoxically, LLMs are increasingly adopted as successful evaluative agents for…

Computation and Language · Computer Science 2025-09-23 Sungwon Kim , Daniel Khashabi

Equipping agents with the capacity to justify made decisions using supporting evidence represents a cornerstone of accountable decision-making. Furthermore, ensuring that justifications are in line with human expectations and societal norms…

Machine Learning · Computer Science 2024-02-27 Aleksa Sukovic , Goran Radanovic

Large Language Models (LLMs) have exhibited impressive capabilities across diverse application domains. Recent work has explored Multi-LLM Agent Debate (MAD) as a way to enhance performance by enabling multiple LLMs to discuss and refine…

Computation and Language · Computer Science 2026-05-27 Xuhang Chen , Zhifan Song , Deyi Ji , Shuo Gao , Lanyun Zhu

When Artificial Intelligence (AI) is used to replace consumers (e.g., synthetic data), it is often assumed that AI emulates established consumers, and more generally human behaviors. Ten experiments with Large Language Models (LLMs)…

Human-Computer Interaction · Computer Science 2025-10-10 Antonios Stamatogiannakis , Arsham Ghodsinia , Sepehr Etminanrad , Dilney Gonçalves , David Santos

Much of the success of multi-agent debates depends on carefully choosing the right parameters. The decision-making protocol stands out as it can highly impact final model answers, depending on how decisions are reached. Systematic…

Multiagent Systems · Computer Science 2025-10-01 Lars Benedikt Kaesberg , Jonas Becker , Jan Philip Wahle , Terry Ruas , Bela Gipp

Artificial Intelligence (AI) increasingly shows its potential to outperform predicate logic algorithms and human control alike. In automatically deriving a system model, AI algorithms learn relations in data that are not detectable for…

Artificial Intelligence · Computer Science 2022-10-12 Simon Daniel Duque Anton , Daniel Schneider , Hans Dieter Schotten

Large language models (LLMs) are increasingly integral to information retrieval (IR), powering ranking, evaluation, and AI-assisted content creation. This widespread adoption necessitates a critical examination of potential biases arising…

Information Retrieval · Computer Science 2025-07-11 Krisztian Balog , Donald Metzler , Zhen Qin

This paper explores the assumption that Large Language Models (LLMs) skilled in generation tasks are equally adept as evaluators. We assess the performance of three LLMs and one open-source LM in Question-Answering (QA) and evaluation tasks…

Computation and Language · Computer Science 2026-03-09 Juhyun Oh , Eunsu Kim , Inha Cha , Alice Oh

Despite the utility of Large Language Models (LLMs) across a wide range of tasks and scenarios, developing a method for reliably evaluating LLMs across varied contexts continues to be challenging. Modern evaluation approaches often use LLMs…

Computation and Language · Computer Science 2024-01-31 Steffi Chern , Ethan Chern , Graham Neubig , Pengfei Liu