中文
相关论文

相关论文: Online Learning for Dynamic Vickrey-Clarke-Groves …

200 篇论文

We study cooperative online learning in stochastic and adversarial Markov decision process (MDP). That is, in each episode, $m$ agents interact with an MDP simultaneously and share information in order to minimize their individual regret.…

机器学习 · 计算机科学 2022-09-02 Tal Lancewicki , Aviv Rosenberg , Yishay Mansour

We consider the problem of learning by demonstration from agents acting in unknown stochastic Markov environments or games. Our aim is to estimate agent preferences in order to construct improved policies for the same task that the agents…

机器学习 · 计算机科学 2014-08-12 Aristide Tossou , Christos Dimitrakakis

We consider the problem of learning by demonstration from agents acting in unknown stochastic Markov environments or games. Our aim is to estimate agent preferences in order to construct improved policies for the same task that the agents…

机器学习 · 统计学 2013-07-16 Aristide C. Y. Tossou , Christos Dimitrakakis

Reinforcement learning algorithms are typically designed for generic Markov Decision Processes (MDPs), where any state-action pair can lead to an arbitrary transition distribution. In many practical systems, however, only a subset of the…

机器学习 · 计算机科学 2026-03-05 Davide Maran , Davide Salaorni , Marcello Restelli

Incrementality, which is used to measure the causal effect of showing an ad to a potential customer (e.g. a user in an internet platform) versus not, is a central object for advertisers in online advertising platforms. This paper…

机器学习 · 计算机科学 2023-01-18 Ashwinkumar Badanidiyuru , Zhe Feng , Tianxi Li , Haifeng Xu

We consider the problem of a single seller repeatedly selling a single item to a single buyer (specifically, the buyer has a value drawn fresh from known distribution $D$ in every round). Prior work assumes that the buyer is fully rational…

计算机科学与博弈论 · 计算机科学 2017-11-28 Mark Braverman , Jieming Mao , Jon Schneider , S. Matthew Weinberg

We study reserve price optimization in multi-phase second price auctions, where the seller's prior actions affect the bidders' later valuations through a Markov Decision Process (MDP). Compared to the bandit setting in existing works, the…

机器学习 · 计算机科学 2026-03-04 Rui Ai , Boxiang Lyu , Zhaoran Wang , Zhuoran Yang , Michael I. Jordan

While significant advancements have been made in the field of fair machine learning, the majority of studies focus on scenarios where the decision model operates on a static population. In this paper, we study fairness in dynamic systems…

机器学习 · 计算机科学 2024-01-15 Yaowei Hu , Jacob Lear , Lu Zhang

In this paper, a rather general online problem called dynamic resource allocation with capacity constraints (DRACC) is introduced and studied in the realm of posted price mechanisms. This problem subsumes several applications of stateful…

计算机科学与博弈论 · 计算机科学 2020-06-30 Yuval Emek , Ron Lavi , Rad Niazadeh , Yangguang Shi

Most learning algorithms with formal regret guarantees essentially rely on trying all possible behaviors, which is problematic when some errors cannot be recovered from. Instead, we allow the learning agent to ask for help from a mentor and…

机器学习 · 计算机科学 2025-09-17 Benjamin Plaut , Juan Liévano-Karim , Hanlin Zhu , Stuart Russell

Motivated by pricing in ad exchange markets, we consider the problem of robust learning of reserve prices against strategic buyers in repeated contextual second-price auctions. Buyers' valuations for an item depend on the context that…

机器学习 · 计算机科学 2020-02-27 Negin Golrezaei , Adel Javanmard , Vahab Mirrokni

We study the problem of selling identical goods to n unit-demand bidders in a setting in which the total supply of goods is unknown to the mechanism. Items arrive dynamically, and the seller must make the allocation and payment decisions…

计算机科学与博弈论 · 计算机科学 2009-05-22 Moshe Babaioff , Liad Blumrosen , Aaron L. Roth

Average-reward Markov decision processes (MDPs) provide a foundational framework for sequential decision-making under uncertainty. However, average-reward MDPs have remained largely unexplored in reinforcement learning (RL) settings, with…

机器学习 · 计算机科学 2025-08-29 Juan Sebastian Rojas , Chi-Guhn Lee

In Bayesian single-item auctions, a monotone bidding strategy--one that prescribes a higher bid for a higher value type--can be equivalently represented as a partition of the quantile space into consecutive intervals corresponding to…

计算机科学与博弈论 · 计算机科学 2026-02-10 Junyao Zhao

We propose a new architecture to approximately learn incentive compatible, revenue-maximizing auctions from sampled valuations. Our architecture uses the Sinkhorn algorithm to perform a differentiable bipartite matching which allows the…

计算机科学与博弈论 · 计算机科学 2021-06-16 Michael J. Curry , Uro Lyi , Tom Goldstein , John Dickerson

We consider an agent who is involved in a Markov decision process and receives a vector of outcomes every round. Her objective is to maximize a global concave reward function on the average vectorial outcome. The problem models applications…

机器学习 · 计算机科学 2019-05-17 Wang Chi Cheung

Inferring the underlying graph topology that characterizes structured data is pivotal to many graph-based models when pre-defined graphs are not available. This paper focuses on learning graphs in the case of sequential data in dynamic…

机器学习 · 计算机科学 2022-02-25 Xiang Zhang

Mechanism design, a branch of economics, aims to design rules that can autonomously achieve desired outcomes in resource allocation and public decision making. The research on mechanism design using machine learning is called automated…

计算机科学与博弈论 · 计算机科学 2024-12-17 Tsuyoshi Suehara , Koh Takeuchi , Hisashi Kashima , Satoshi Oyama , Yuko Sakurai , Makoto Yokoo

We address the challenging problem of dynamically pricing complementary items that are sequentially displayed to customers. An illustrative example is the online sale of flight tickets, where customers navigate through multiple web pages.…

A natural optimization model that formulates many online resource allocation and revenue management problems is the online linear program (LP) in which the constraint matrix is revealed column by column along with the corresponding…

数据结构与算法 · 计算机科学 2014-04-10 Shipra Agrawal , Zizhuo Wang , Yinyu Ye