English
Related papers

Related papers: Influencing Bandits: Arm Selection for Preference …

200 papers

We advance the study of incentivized bandit exploration, in which arm choices are viewed as recommendations and are required to be Bayesian incentive compatible. Recent work has shown under certain independence assumptions that after…

Computer Science and Game Theory · Computer Science 2024-09-25 Mark Sellke

In Batched Multi-Armed Bandits (BMAB), the policy is not allowed to be updated at each time step. Usually, the setting asserts a maximum number of allowed policy updates and the algorithm schedules them so that to minimize the expected…

Machine Learning · Computer Science 2021-10-01 Romain Laroche , Othmane Safsafi , Raphael Feraud , Nicolas Broutin

We consider an adversarial online learning setting where a decision maker can choose an action in every stage of the game. In addition to observing the reward of the chosen action, the decision maker gets side observations on the reward he…

Machine Learning · Computer Science 2011-10-26 Shie Mannor , Ohad Shamir

The focus of this work is on designing influencing strategies to shape the collective opinion of a network of individuals. We consider a variant of the voter model where opinions evolve in one of two ways. In the absence of external…

Social and Information Networks · Computer Science 2020-02-04 Anmol Gupta , Sharayu Moharir , Neeraja Sahasrabudhe

In many biomedical, science, and engineering problems, one must sequentially decide which action to take next so as to maximize rewards. One general class of algorithms for optimizing interactions with the world, while simultaneously…

Machine Learning · Statistics 2021-05-05 Iñigo Urteaga , Chris H. Wiggins

Opinion dynamics on social networks have been received considerable attentions in recent years. Nevertheless, just a few works have theoretically analyzed the condition in which a certain opinion can spread in the whole structured…

Computer Science and Game Theory · Computer Science 2022-08-31 Zhifang Li , Xiaojie Chen , Han-Xin Yang , Attila Szolnoki

Digital educational technologies offer the potential to customize students' experiences and learn what works for which students, enhancing the technology as more students interact with it. We consider whether and when attempting to discover…

Artificial Intelligence · Computer Science 2023-09-07 ZhaoBin Li , Luna Yee , Nathaniel Sauerberg , Irene Sakson , Joseph Jay Williams , Anna N. Rafferty

A common explanation for negative user impacts of content recommender systems is misalignment between the platform's objective and user welfare. In this work, we show that misalignment in the platform's objective is not the only potential…

Machine Learning · Computer Science 2024-01-26 Jessica Dai , Bailey Flanigan , Nika Haghtalab , Meena Jagadeesan , Chara Podimata

We study a stochastic model for the diffusion of competing opinions in a population composed of three types of agents: trend-followers, opposers, and indifferent individuals. The decision dynamics are driven by reinforcement mechanisms,…

Probability · Mathematics 2025-06-24 Manuel González-Navarrete

We consider a scenario where an agent has multiple available strategies to explore an unknown environment. For each new interaction with the environment, the agent must select which exploration strategy to use. We provide a new…

Machine Learning · Computer Science 2018-08-24 Fabien C. Y. Benureau , Pierre-Yves Oudeyer

Adaptive experiments such as multi-arm bandits adapt the treatment-allocation policy and/or the decision to stop the experiment to the data observed so far. This has the potential to improve outcomes for study participants within the…

Methodology · Statistics 2024-05-03 Aurélien Bibaut , Nathan Kallus

The housing market, also known as one-sided matching market, is a classic exchange economy model where each agent on the demand side initially owns an indivisible good (a house) and has a personal preference over all goods. The goal is to…

Computer Science and Game Theory · Computer Science 2026-01-08 Shiyun Lin

Two-sided online matching platforms are employed in various markets. However, agents' preferences in the current market are usually implicit and unknown, thus needing to be learned from data. With the growing availability of dynamic side…

Machine Learning · Computer Science 2024-05-30 Yuantong Li , Chi-hua Wang , Guang Cheng , Will Wei Sun

We consider bandit problems involving a large (possibly infinite) collection of arms, in which the expected reward of each arm is a linear function of an $r$-dimensional random vector $\mathbf{Z} \in \mathbb{R}^r$, where $r \geq 2$. The…

Machine Learning · Computer Science 2010-02-24 Paat Rusmevichientong , John N. Tsitsiklis

We study a cooperative multi-agent bandit setting in the distributed GOSSIP model: in every round, each of $n$ agents chooses an action from a common set, observes the action's corresponding reward, and subsequently exchanges information…

Machine Learning · Computer Science 2024-10-21 John Lazarsfeld , Dan Alistarh

We present a novel approach to deformable object manipulation that does not rely on highly-accurate modeling. The key contribution of this paper is to formulate the task as a Multi-Armed Bandit problem, with each arm representing a model of…

Robotics · Computer Science 2020-06-02 Dale McConachie , Dmitry Berenson

Experimentation with interference poses a significant challenge in contemporary online platforms. Prior research on experimentation with interference has concentrated on the final output of a policy. The cumulative performance, while…

Machine Learning · Computer Science 2024-07-17 Su Jia , Peter Frazier , Nathan Kallus

Multi-Armed-Bandit frameworks have often been used by researchers to assess educational interventions, however, recent work has shown that it is more beneficial for a student to provide qualitative feedback through preference elicitation…

Machine Learning · Computer Science 2021-11-02 Nayan Saxena , Pan Chen , Emmy Liu

The dueling bandits problem is an online learning framework for learning from pairwise preference feedback, and is particularly well-suited for modeling settings that elicit subjective or implicit human feedback. In this paper, we study the…

Machine Learning · Computer Science 2017-05-02 Yanan Sui , Vincent Zhuang , Joel W. Burdick , Yisong Yue

Online decision-making can be formulated as the popular stochastic multi-armed bandit problem where a learner makes decisions (or takes actions) to maximize cumulative rewards collected from an unknown environment. This paper proposes to…

Systems and Control · Electrical Eng. & Systems 2025-11-26 Jonathan Gornet , Mehdi Hosseinzadeh , Bruno Sinopoli