中文
相关论文

相关论文: Rethinking Incentives in Recommender Systems: Are …

200 篇论文

Driven by the new economic opportunities created by the creator economy, an increasing number of content creators rely on and compete for revenue generated from online content recommendation platforms. This burgeoning competition reshapes…

信息检索 · 计算机科学 2024-04-30 Fan Yao , Yiming Liao , Mingzhe Wu , Chuanhao Li , Yan Zhu , James Yang , Qifan Wang , Haifeng Xu , Hongning Wang

Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences. However, RM-based preference-tuning is vulnerable to reward hacking, whereby LM policies learn undesirable behaviors from flawed RMs. By…

计算与语言 · 计算机科学 2026-03-05 Daniel Fein , Max Lamparth , Violet Xiang , Mykel J. Kochenderfer , Nick Haber

Reward models (RMs) are central to the alignment of language models (LMs). An RM often serves as a proxy for human preferences to guide downstream LM behavior. However, our understanding of RM behavior is limited. Our work (i) formalizes a…

计算与语言 · 计算机科学 2025-10-09 Elle

We present CRM (Multi-Agent Collaborative Reward Model), a framework that replaces a single black-box reward model with a coordinated team of specialist evaluators to improve robustness and interpretability in RLHF. Conventional reward…

人工智能 · 计算机科学 2026-01-06 Pei Yang , Ke Zhang , Ji Wang , Xiao Chen , Yuxin Tang , Eric Yang , Lynn Ai , Bill Shi

Reward models (RMs) play a crucial role in reinforcement learning from human feedback (RLHF), aligning model behavior with human preferences. However, existing benchmarks for reward models show a weak correlation with the performance of…

机器学习 · 计算机科学 2025-05-20 Sunghwan Kim , Dongjin Kang , Taeyoon Kwon , Hyungjoo Chae , Dongha Lee , Jinyoung Yeo

Reward models are critical for reinforcement learning from human feedback, as they determine the alignment quality and reliability of generative models. For complex tasks such as image editing, reward models are required to capture global…

Content creators compete for exposure on recommendation platforms, and such strategic behavior leads to a dynamic shift over the content distribution. However, how the creators' competition impacts user welfare and how the relevance-driven…

计算机科学与博弈论 · 计算机科学 2023-05-04 Fan Yao , Chuanhao Li , Denis Nekipelov , Hongning Wang , Haifeng Xu

Online platforms in the Internet Economy commonly incorporate recommender systems that recommend products (or "arms") to users (or "agents"). A key challenge in this domain arises from myopic agents who are naturally incentivized to exploit…

信息检索 · 计算机科学 2024-06-19 Xiaowu Dai , Wenlu Xu , Yuan Qi , Michael I. Jordan

The prevalence of low-quality content on online platforms is often attributed to the absence of meaningful entry requirements. This motivates us to investigate whether implicit or explicit entry barriers, alongside appropriate reward…

计算机科学与博弈论 · 计算机科学 2025-09-03 Haiqing Zhu , Lexing Xie , Yun Kuen Cheung

Reward models play a critical role in guiding large language models toward outputs that align with human expectations. However, an open challenge remains in effectively utilizing test-time compute to enhance reward model performance. In…

计算与语言 · 计算机科学 2025-05-21 Jiaxin Guo , Zewen Chi , Li Dong , Qingxiu Dong , Xun Wu , Shaohan Huang , Furu Wei

Reward-based fine-tuning steers a pretrained diffusion or flow-based generative model toward higher-reward samples while remaining close to the pretrained model. Although existing methods are derived from different perspectives, we show…

机器学习 · 计算机科学 2026-05-08 Jeongjae Lee , Jinho Chang , Jeongsol Kim , Jong Chul Ye

Understanding the bias-variance tradeoff in user representation learning is essential for improving recommendation quality in modern content platforms. While well studied in static settings, this tradeoff becomes significantly more complex…

计算机科学与博弈论 · 计算机科学 2026-03-03 Kang Wang , Renzhe Xu , Bo Li

In mechanism design it is typical to impose incentive compatibility and then derive an optimal mechanism subject to this constraint. By replacing the incentive compatibility requirement with the goal of minimizing expected ex post regret,…

计算机科学与博弈论 · 计算机科学 2012-08-07 Paul Duetting , Felix Fischer , Pitchayut Jirapinyo , John K. Lai , Benjamin Lubin , David C. Parkes

Reward models (RMs) guide the alignment of large language models (LLMs), steering them toward behaviors preferred by humans. Evaluating RMs is the key to better aligning LLMs. However, the current evaluation of RMs may not directly…

Reward models (RMs) are inherently non-neutral value functions designed and trained to encode specific objectives, such as human preferences or text-image alignment. RMs have become crucial components of text-to-image (T2I) generation…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Salma Abdel Magid , Grace Guo , Esin Tureci , Amaya Dharmasiri , Vikram V. Ramaswamy , Hanspeter Pfister , Olga Russakovsky

The rapid advancement of Multimodal Large Language Models (MLLMs) has made aligning them with human preferences a critical challenge. Reward Models (RMs) are a core technology for achieving this goal, but a systematic guide for building…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Yi-Fan Zhang , Haihua Yang , Huanyu Zhang , Yang Shi , Zezhou Chen , Haochen Tian , Chaoyou Fu , Haotian Wang , Kai Wu , Bo Cui , Xu Wang , Jianfei Pan , Haotian Wang , Zhang Zhang , Liang Wang

Consumer generated media (CGM), such as social networking services rely on the voluntary activity of users to prosper, garnering the psychological rewards of feeling connected with other people through comments and reviews received online.…

社会与信息网络 · 计算机科学 2023-06-02 Shintaro Ueki , Fujio Toriumi , Toshiharu Sugawara

Nowadays, rating systems play a crucial role in the attraction of customers for different services. However, as it is difficult to detect a fake rating, attackers can potentially impact the rating's aggregated score unfairly. This malicious…

计算机科学与博弈论 · 计算机科学 2022-08-05 Iman Vakilinia , Peyman Faizian , Mohammad Mahdi Khalili

Mechanism design is a well-established game-theoretic paradigm for designing games to achieve desired outcomes. This paper addresses a closely related but distinct concept, equilibrium design. Unlike mechanism design, the designer's…

计算机科学与博弈论 · 计算机科学 2024-08-20 Muhammad Najib , Giuseppe Perelli

Crowdsourcing has emerged as a paradigm for leveraging human intelligence and activity to solve a wide range of tasks. However, strategic workers will find enticement in their self-interest to free-ride and attack in a crowdsourcing contest…

计算机科学与博弈论 · 计算机科学 2018-01-01 Jianfeng Lu , Yun Xin , Zhao Zhang , Shaojie Tang , Songyuan Yan , Changbing Tang
‹ 上一页 1 2 3 10 下一页 ›