中文
相关论文

相关论文: Strategic Candidacy in Generative AI Arenas

200 篇论文

Large language models are often ranked according to their level of alignment with human preferences -- a model is better than other models if its outputs are more frequently preferred by humans. One of the popular ways to elicit human…

机器学习 · 计算机科学 2024-12-05 Ivi Chatzi , Eleni Straitouri , Suhas Thejaswi , Manuel Gomez Rodriguez

We explore a new way to evaluate generative models using insights from evaluation of competitive games between human players. We show experimentally that tournaments between generators and discriminators provide an effective way to evaluate…

机器学习 · 统计学 2018-08-16 Catherine Olsson , Surya Bhupatiraju , Tom Brown , Augustus Odena , Ian Goodfellow

Recommender systems are one of the most pervasive applications of machine learning in industry, with many services using them to match users to products or information. As such it is important to ask: what are the possible fairness risks,…

计算机与社会 · 计算机科学 2019-03-12 Alex Beutel , Jilin Chen , Tulsee Doshi , Hai Qian , Li Wei , Yi Wu , Lukasz Heldt , Zhe Zhao , Lichan Hong , Ed H. Chi , Cristos Goodrow

The rapid advancement of visual generative models necessitates efficient and reliable evaluation methods. Arena platform, which gathers user votes on model comparisons, can rank models with human preferences. However, traditional Arena…

人工智能 · 计算机科学 2025-03-18 Zhikai Li , Xuewen Liu , Dongrong Joe Fu , Jianquan Li , Qingyi Gu , Kurt Keutzer , Zhen Dong

We propose a topic modeling approach to the prediction of preferences in pairwise comparisons. We develop a new generative model for pairwise comparisons that accounts for multiple shared latent rankings that are prevalent in a population…

机器学习 · 计算机科学 2015-01-27 Weicong Ding , Prakash Ishwar , Venkatesh Saligrama

Innovation in artificial intelligence (AI) has always been dependent on technological infrastructures, from code repositories to computing hardware. Yet industry -- rather than universities -- has become increasingly influential in shaping…

计算机与社会 · 计算机科学 2025-12-18 Sam Hind

Consider a marketplace of AI tools, each with slightly different strengths and weaknesses. By selecting the right model for the task at hand, a user can do better than simply committing to a single model for everything. Routers operate…

计算机科学与博弈论 · 计算机科学 2026-02-27 Kate Donahue , Manish Raghavan

Owing to the advancement of deep learning, artificial systems are now rival to humans in several pattern recognition tasks, such as visual recognition of object categories. However, this is only the case with the tasks for which correct…

机器学习 · 计算机科学 2019-06-03 Xing Liu , Takayuki Okatani

The problem of ranking is a multi-billion dollar problem. In this paper we present an overview of several production quality ranking systems. We show that due to conflicting goals of employing the most effective machine learning models and…

信息检索 · 计算机科学 2019-07-30 Murium Iqbal , Nishan Subedi , Kamelia Aryafar

Foundation models require fine-tuning to ensure their generative outputs align with intended results for specific tasks. Automating this fine-tuning process is challenging, as it typically needs human feedback that can be expensive to…

Mainstream ranking approaches typically follow a Generator-Evaluator two-stage paradigm, where a generator produces candidate lists and an evaluator selects the best one. Recent work has attempted to enhance performance by expanding the…

信息检索 · 计算机科学 2026-01-28 Kaike Zhang , Xiaobei Wang , Shuchang Liu , Hailan Yang , Xiang Li , Lantao Hu , Han Li , Qi Cao , Fei Sun , Kun Gai

Generative AI has made remarkable strides to revolutionize fields such as image and video generation. These advancements are driven by innovative algorithms, architecture, and data. However, the rapid proliferation of generative models has…

人工智能 · 计算机科学 2024-11-12 Dongfu Jiang , Max Ku , Tianle Li , Yuansheng Ni , Shizhuo Sun , Rongqi Fan , Wenhu Chen

The rapid advancement of visual generation models has outpaced traditional evaluation approaches, necessitating the adoption of Vision-Language Models as surrogate judges. In this work, we systematically investigate the reliability of the…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Ruihang Li , Leigang Qu , Jingxu Zhang , Dongnan Gui , Mengde Xu , Xiaosong Zhang , Han Hu , Wenjie Wang , Jiaqi Wang

The evaluation of large language models is a complex task, in which several approaches have been proposed. The most common is the use of automated benchmarks in which LLMs have to answer multiple-choice questions of different topics.…

人工智能 · 计算机科学 2025-07-18 Carlos Arriaga , Gonzalo Martínez , Eneko Sendin , Javier Conde , Pedro Reviriego

With the advent of highly capable instruction-tuned neural language models, benchmarking in natural language processing (NLP) is increasingly shifting towards pairwise comparison leaderboards, such as LMSYS Arena, from traditional global…

计算与语言 · 计算机科学 2025-09-24 Georgii Levtsov , Dmitry Ustalov

Generative AI models differ from traditional machine learning tools in that they allow users to provide as much or as little information as they choose in their inputs. This flexibility often leads users to omit certain details, relying on…

计算机科学与博弈论 · 计算机科学 2026-05-13 Charlotte Park , Kate Donahue , Manish Raghavan

We reformulate explanation quality assessment as a ranking problem rather than a generation problem. Instead of optimizing models to produce a single "best" explanation token-by-token, we train reward models to discriminate among multiple…

人工智能 · 计算机科学 2026-04-28 Thomas Bailleux , Tanmoy Mukherjee , Emmanuel Lonca , Pierre Marquis , Zied Bouraoui

Modern large-scale ranking systems operate within a sophisticated landscape of competing objectives, operational constraints, and evolving product requirements. Progress in this domain is increasingly bottlenecked by the engineering context…

Evaluating the quality of retrieval-augmented generation (RAG) and document reranking systems remains challenging due to the lack of scalable, user-centric, and multi-perspective evaluation tools. We introduce RankArena, a unified platform…

信息检索 · 计算机科学 2025-08-08 Abdelrahman Abdallah , Mahmoud Abdalla , Bhawna Piryani , Jamshid Mozafari , Mohammed Ali , Adam Jatowt

Aligning large visual generative models with human feedback is often performed through pairwise preference optimization. While such approaches are conceptually simple, they fundamentally rely on annotated pairs, limiting scalability in…

机器学习 · 计算机科学 2026-05-07 Jinbin Bai , Yu Lei , Qingyu Shi , Aosong Feng , Yi Xin , Zhuoran Zhao , Fei Shen , Kaidong Yu , Jason Li
‹ 上一页 1 2 3 10 下一页 ›