English
Related papers

Related papers: Carrot and Stick: Eliciting Comparison Data and Be…

200 papers

Recent approaches in personalized reward modeling have primarily focused on leveraging user interaction history to align model judgments with individual preferences. However, existing approaches largely treat user context as a static or…

Computation and Language · Computer Science 2026-04-21 Kwangwook Seo , Dongha Lee

Humans often specify tasks incompletely, so assistants must know when and how to ask clarifying questions. However, effective clarification remains challenging in software engineering tasks as not all missing information is equally…

Software Engineering · Computer Science 2026-04-17 Sanidhya Vijayvargiya , Vijay Viswanathan , Graham Neubig

Crowdsourcing is now widely used to replace judgement by an expert authority with an aggregate evaluation from a number of non-experts, in applications ranging from rating and categorizing online content to evaluation of student assignments…

Computer Science and Game Theory · Computer Science 2013-03-05 Anirban Dasgupta , Arpita Ghosh

Current alignment pipelines presume a single, universal notion of desirable behavior. However, human preferences often diverge across users, contexts, and cultures. As a result, disagreement collapses into the majority signal and minority…

Machine Learning · Computer Science 2025-06-10 Daniel Halpern , Evi Micha , Ariel D. Procaccia , Itai Shapira

We consider situations where a user feeds her attributes to a machine learning method that tries to predict her best option based on a random sample of other users. The predictor is incentive-compatible if the user has no incentive to…

Econometrics · Economics 2021-09-07 Mehmet Caner , Kfir Eliaz

To ensure that large language model (LLM) responses are helpful and non-toxic, a reward model trained on human preference data is usually used. LLM responses with high rewards are then selected through best-of-$n$ (BoN) sampling or the LLM…

Machine Learning · Computer Science 2024-07-04 Adam X. Yang , Maxime Robeyns , Thomas Coste , Zhengyan Shi , Jun Wang , Haitham Bou-Ammar , Laurence Aitchison

Bayesian persuasion, an extension of cheap-talk communication, involves an informed sender committing to a signaling scheme to influence a receiver's actions. Compared to cheap talk, this sender's commitment enables the receiver to verify…

Computer Science and Game Theory · Computer Science 2025-06-10 Yue Lin , Shuhui Zhu , William A Cunningham , Wenhao Li , Pascal Poupart , Hongyuan Zha , Baoxiang Wang

Suppose a decision maker wants to predict weather tomorrow by eliciting and aggregating information from crowd. How can the decision maker incentivize the crowds to report their information truthfully? Many truthful peer prediction…

Computer Science and Game Theory · Computer Science 2021-07-22 Qishen Han , Sikai Ruan , Yuqing Kong , Ao Liu , Farhad Mohsin , Lirong Xia

Decision makers often need to rely on imperfect probabilistic forecasts. While average performance metrics are typically available, it is difficult to assess the quality of individual forecasts and the corresponding utilities. To convey…

Machine Learning · Statistics 2021-03-03 Shengjia Zhao , Stefano Ermon

Neural models for response generation produce responses that are semantically plausible but not necessarily factually consistent with facts describing the speaker's persona. These models are trained with fully supervised learning where the…

Computation and Language · Computer Science 2021-02-16 Mohsen Mesgar , Edwin Simpson , Iryna Gurevych

Truth discovery is to resolve conflicts and find the truth from multiple-source statements. Conventional methods mostly research based on the mutual effect between the reliability of sources and the credibility of statements, however, pay…

Computation and Language · Computer Science 2016-11-08 Luyang Li , Bing Qin , Wenjing Ren , Ting Liu

Appropriate credit assignment for delay rewards is a fundamental challenge for reinforcement learning. To tackle this problem, we introduce a delay reward calibration paradigm inspired from a classification perspective. We hypothesize that…

Machine Learning · Computer Science 2021-08-26 Yixuan Liu , Hu Wang , Xiaowei Wang , Xiaoyue Sun , Liuyue Jiang , Minhui Xue

Recent alignment techniques, such as reinforcement learning from human feedback, have been widely adopted to align large language models with human preferences by learning and leveraging reward models. In practice, these models often…

Machine Learning · Computer Science 2025-10-29 Ignavier Ng , Patrick Blöbaum , Siddharth Bhandari , Kun Zhang , Shiva Kasiviswanathan

The correct specification of reward models is a well-known challenge in reinforcement learning. Hand-crafted reward functions often lead to inefficient or suboptimal policies and may not be aligned with user values. Reinforcement learning…

Artificial Intelligence · Computer Science 2024-10-24 Muhan Lin , Shuyang Shi , Yue Guo , Behdad Chalaki , Vaishnav Tadiparthi , Ehsan Moradi Pari , Simon Stepputtis , Joseph Campbell , Katia Sycara

We study the design of decision-making mechanism for resource allocations over a multi-agent system in a dynamic environment. Agents' privately observed preference over resources evolves over time and the population is dynamic due to the…

Systems and Control · Electrical Eng. & Systems 2020-05-20 Tao Zhang , Quanyan Zhu

Ranking is fundamental to many areas, such as search engine optimization, human feedback for language models, as well as peer grading. Crowdsourcing, which is often used for these tasks, requires proper incentivization to ensure accurate…

Computer Science and Game Theory · Computer Science 2024-01-26 Kiriaki Frangias , Andrew Lin , Ellen Vitercik , Manolis Zampetakis

Comparative Judgement is an assessment method where item ratings are estimated based on rankings of subsets of the items. These rankings are typically pairwise, with ratings taken to be the estimated parameters from fitting a Bradley-Terry…

Methodology · Statistics 2024-05-22 Ian Hamilton , Nick Tawn

We study truthful mechanisms for matching and related problems in a partial information setting, where the agents' true utilities are hidden, and the algorithm only has access to ordinal preference information. Our model is motivated by the…

Computer Science and Game Theory · Computer Science 2016-10-20 Elliot Anshelevich , Shreyas Sekar

Recent approaches to question generation have used modifications to a Seq2Seq architecture inspired by advances in machine translation. Models are trained using teacher forcing to optimise only the one-step-ahead prediction. However, at…

Computation and Language · Computer Science 2019-06-04 Tom Hosking , Sebastian Riedel

When users lack specific knowledge of various system parameters, their uncertainty may lead them to make undesirable deviations in their decision making. To alleviate this, an informed system operator may elect to signal information to…

Computer Science and Game Theory · Computer Science 2023-03-31 Bryce L. Ferguson , Philip N. Brown , Jason R. Marden
‹ Prev 1 3 4 5 6 7 10 Next ›