English
Related papers

Related papers: Improving Your Model Ranking on Chatbot Arena by V…

200 papers

It is now common to evaluate Large Language Models (LLMs) by having humans manually vote to evaluate model outputs, in contrast to typical benchmarks that evaluate knowledge or skill at some particular task. Chatbot Arena, the most popular…

Large Language Models (LLMs) have unlocked new capabilities and applications; however, evaluating the alignment with human preferences still poses significant challenges. To address this issue, we introduce Chatbot Arena, an open platform…

Assessing the effectiveness of large language models (LLMs) presents substantial challenges. The method of conducting human-annotated battles in an online Chatbot Arena is a highly effective evaluative technique. However, this approach is…

Computation and Language · Computer Science 2024-07-16 Haipeng Luo , Qingfeng Sun , Can Xu , Pu Zhao , Qingwei Lin , Jianguang Lou , Shifeng Chen , Yansong Tang , Weizhu Chen

As LLMs continuously evolve, there is an urgent need for a reliable evaluation method that delivers trustworthy results promptly. Currently, static benchmarks suffer from inflexibility and unreliability, leading users to prefer human voting…

Computation and Language · Computer Science 2024-10-08 Ruochen Zhao , Wenxuan Zhang , Yew Ken Chia , Weiwen Xu , Deli Zhao , Lidong Bing

Similar to social media bots that shape public opinion, healthcare and financial decisions, LLM-based ChatBots like ChatGPT can persuade users to alter their behavior. Unlike prior work that persuades via overt-partisan bias or…

Human-Computer Interaction · Computer Science 2025-11-21 Anthony Wise , Xinyi Zhou , Martin Reimann , Anind Dey , Leilani Battle

Open community-driven platforms like Chatbot Arena that collect user preference data from site visitors have gained a reputation as one of the most trustworthy publicly available benchmarks for LLM performance. While now standard, it is…

Human-Computer Interaction · Computer Science 2024-12-06 Wenting Zhao , Alexander M. Rush , Tanya Goyal

Measuring progress is fundamental to the advancement of any scientific field. As benchmarks play an increasingly central role, they also grow more susceptible to distortion. Chatbot Arena has emerged as the go-to leaderboard for ranking the…

Reward models (RMs) play a crucial role in Reinforcement Learning from Human Feedback by serving as proxies for human preferences in aligning large language models. However, they suffer from various biases which could lead to reward…

Artificial Intelligence · Computer Science 2026-03-18 Xiao Zhu , Chenmien Tan , Pinzhen Chen , Rico Sennrich , Huiming Wang , Yanlin Zhang , Hanxu Hu

With the growing adoption and capabilities of vision-language models (VLMs) comes the need for benchmarks that capture authentic user-VLM interactions. In response, we create VisionArena, a dataset of 230K real-world conversations between…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Christopher Chou , Lisa Dunlap , Koki Mashita , Krishna Mandal , Trevor Darrell , Ion Stoica , Joseph E. Gonzalez , Wei-Lin Chiang

Battles, or side-by-side comparisons in so-called arenas that elicit human preferences, have emerged as a popular approach for assessing the output quality of LLMs. Recently, this idea has been extended to retrieval-augmented generation…

Information Retrieval · Computer Science 2025-05-27 Sahel Sharifymoghaddam , Shivani Upadhyay , Nandan Thakur , Ronak Pradeep , Jimmy Lin

We propose a method for evaluating the robustness of widely used LLM ranking systems -- variants of a Bradley--Terry model -- to dropping a worst-case very small fraction of preference data. Our approach is computationally fast and easy to…

Machine Learning · Statistics 2026-03-06 Jenny Y. Huang , Yunyi Shen , Dennis Wei , Tamara Broderick

Large language models (LLMs) have transformed natural language processing, with frameworks like Chatbot Arena providing pioneering platforms for evaluating these models. By facilitating millions of pairwise comparisons based on human…

Machine Learning · Statistics 2025-06-02 Siavash Ameli , Siyuan Zhuang , Ion Stoica , Michael W. Mahoney

Evaluating large language model (LLM) based chat assistants is challenging due to their broad capabilities and the inadequacy of existing benchmarks in measuring human preferences. To address this, we explore using strong LLMs as judges to…

With the advent of Large Language Models (LLM), conversational assistants have become prevalent for domain use cases. LLMs acquire the ability to contextual question answering through training, and Retrieval Augmented Generation (RAG)…

Computation and Language · Computer Science 2024-01-17 Mandar Kulkarni , Praveen Tangarajan , Kyung Kim , Anusua Trivedi

Search-augmented language models combine web search with Large Language Models (LLMs) to improve response groundedness and freshness. However, analyzing these systems remains challenging: existing datasets are limited in scale and narrow in…

Evaluating large language models (LLMs) is challenging. Traditional ground-truth-based benchmarks fail to capture the comprehensiveness and nuance of real-world queries, while LLM-as-judge benchmarks suffer from grading biases and limited…

Computation and Language · Computer Science 2024-10-15 Jinjie Ni , Fuzhao Xue , Xiang Yue , Yuntian Deng , Mahir Shah , Kabir Jain , Graham Neubig , Yang You

Evaluating the reasoning abilities of large language models (LLMs) is challenging. Existing benchmarks often depend on static datasets, which are vulnerable to data contamination and may get saturated over time, or on binary live human…

Artificial Intelligence · Computer Science 2025-02-18 Lanxiang Hu , Qiyu Li , Anze Xie , Nan Jiang , Ion Stoica , Haojian Jin , Hao Zhang

In arena-style evaluation of large language models (LLMs), two LLMs respond to a user query, and the user chooses the winning response or deems the "battle" a draw, resulting in an adjustment to the ratings of both models. The prevailing…

Computation and Language · Computer Science 2025-10-03 Raphael Tang , Crystina Zhang , Wenyan Li , Carmen Lai , Pontus Stenetorp , Yao Lu

Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have ushered in a new era of AI capabilities, demonstrating near-human-level performance across diverse scenarios. While numerous benchmarks (e.g., MMLU) and…

Artificial Intelligence · Computer Science 2025-09-03 Kangyu Wang , Hongliang He , Lin Liu , Ruiqi Liang , Zhenzhong Lan , Jianguo Li

The recent explosion of large language models (LLMs), each with its own general or specialized strengths, makes scalable, reliable benchmarking more urgent than ever. Standard practices nowadays face fundamental trade-offs: closed-ended…

‹ Prev 1 2 3 10 Next ›