English
Related papers

Related papers: Policy Learning with $\alpha$-Expected Welfare

200 papers

Collaborative Edge Computing (CEC) is an effective method that improves the performance of Mobile Edge Computing (MEC) systems by offloading computation tasks from busy edge servers (ESs) to idle ones. However, ESs usually belong to…

Networking and Internet Architecture · Computer Science 2022-11-15 Xingqiu He , Yuhang Shen , Hongxi Zhu , Sheng Wang , Chaoqun You , Tony Q. S. Quek

In a Walrasian equilibrium (WE), all bidders are envy-free (EF), meaning that their allocation maximizes their utility; and the market clears (MC), meaning that the price of unallocated goods is zero. EF is desirable to ensure the long-term…

Computer Science and Game Theory · Computer Science 2017-08-11 Enrique Areyan Viqueira , Amy Greenwald , Victor Naroditskiy

We introduce Adversarial Policy Optimization (AdvPO), a novel solution to the pervasive issue of reward over-optimization in Reinforcement Learning from Human Feedback (RLHF) for Large Language Models (LLMs). Over-optimization occurs when a…

Machine Learning · Computer Science 2024-07-10 Xiaoying Zhang , Jean-Francois Ton , Wei Shen , Hongning Wang , Yang Liu

We study offline reinforcement learning (RL) which seeks to learn a good policy based on a fixed, pre-collected dataset. A fundamental challenge behind this task is the distributional shift due to the dataset lacking sufficient exploration,…

Machine Learning · Computer Science 2023-10-11 Wenzhuo Zhou

Searching the space of policies directly for the optimal policy has been one popular method for solving partially observable reinforcement learning problems. Typically, with each change of the target policy, its value is estimated from the…

Artificial Intelligence · Computer Science 2007-05-23 Leonid Peshkin , Christian R. Shelton

We present a recommender system based on the Random Utility Model. Online shoppers are modeled as rational decision makers with limited information, and the recommendation task is formulated as the problem of optimally enriching the…

Computer Science and Game Theory · Computer Science 2024-09-24 Benjamin Heymann , Flavian Vasile , David Rohde

When faced with sequential decision-making problems, it is often useful to be able to predict what would happen if decisions were made using a new policy. Those predictions must often be based on data collected under some previously used…

Machine Learning · Computer Science 2021-11-03 Yash Chandak , Scott Niekum , Bruno Castro da Silva , Erik Learned-Miller , Emma Brunskill , Philip S. Thomas

We study multi-objective reinforcement learning with nonlinear preferences over trajectories. That is, we maximize the expected value of a nonlinear function over accumulated rewards (expected scalarized return or ESR) in a multi-objective…

Machine Learning · Computer Science 2025-02-19 Nianli Peng , Muhang Tian , Brandon Fain

Consider a setting in which a policy maker assigns subjects to treatments, observing each outcome before the next subject arrives. Initially, it is unknown which treatment is best, but the sequential nature of the problem permits learning…

Econometrics · Economics 2020-08-13 Anders Bredahl Kock , David Preinerstorfer , Bezirgen Veliyev

Consider the seller's problem of finding optimal prices for her $n$ (divisible) goods when faced with a set of $m$ consumers, given that she can only observe their purchased bundles at posted prices, i.e., revealed preferences. We study…

Computer Science and Game Theory · Computer Science 2018-10-09 Ziwei Ji , Ruta Mehta , Matus Telgarsky

We propose and axiomatize a new model of incomplete preferences under uncertainty, which we call \textit{hope-and-prepare preferences}. An act is considered more desirable than an other act when, and only when, both an optimistic…

Theoretical Economics · Economics 2025-07-29 Pierre Bardier , Bach Dong-Xuan , Van-Quy Nguyen

The optimal policy of a reinforcement learning problem is often discontinuous and non-smooth. I.e., for two states with similar representations, their optimal policies can be significantly different. In this case, representing the entire…

Machine Learning · Computer Science 2020-02-10 Zhimin Hou , Kuangen Zhang , Yi Wan , Dongyu Li , Chenglong Fu , Haoyong Yu

This paper studies identification and inference of the welfare gain that results from switching from one policy (such as the status quo policy) to another policy. The welfare gain is not point identified in general when data are obtained…

Econometrics · Economics 2022-07-12 Undral Byambadalai

Safe reinforcement learning (RL) aims to learn policies that satisfy certain constraints before deploying them to safety-critical applications. Previous primal-dual style approaches suffer from instability issues and lack optimality…

Machine Learning · Computer Science 2022-06-20 Zuxin Liu , Zhepeng Cen , Vladislav Isenbaev , Wei Liu , Zhiwei Steven Wu , Bo Li , Ding Zhao

For any $\varepsilon>0$, we give a simple, deterministic $(4+\varepsilon)$-approximation algorithm for the Nash social welfare (NSW) problem under submodular valuations. We also consider the asymmetric variant of the problem, where the…

Computer Science and Game Theory · Computer Science 2026-03-31 Jugal Garg , Edin Husić , Wenzheng Li , László A. Végh , Jan Vondrák

This paper studies the statistical theory of batch data reinforcement learning with function approximation. Consider the off-policy evaluation problem, which is to estimate the cumulative value of a new target policy from logged history…

Machine Learning · Computer Science 2020-02-25 Yaqi Duan , Mengdi Wang

This paper presents the Stata community-distributed command "opl_ma_fb" (and the companion command "opl_ma_vf"), for implementing the first-best Optimal Policy Learning (OPL) algorithm to estimate the best treatment assignment given the…

Econometrics · Economics 2025-09-09 Giovanni Cerulli

In a sequential decision-making problem, off-policy evaluation estimates the expected cumulative reward of a target policy using logged trajectory data generated from a different behavior policy, without execution of the target policy.…

Machine Learning · Computer Science 2022-11-04 Jie Wang , Rui Gao , Hongyuan Zha

We state the problem of inverse reinforcement learning in terms of preference elicitation, resulting in a principled (Bayesian) statistical formulation. This generalises previous work on Bayesian inverse reinforcement learning and allows us…

Machine Learning · Statistics 2011-06-30 Constantin Rothkopf , Christos Dimitrakakis

The fundamental assignment problem is in search of welfare maximization mechanisms to allocate items to agents when the private preferences over indivisible items are provided by self-interested agents. The mainstream mechanism…

Computer Science and Game Theory · Computer Science 2019-06-04 Yansong Gao , Jie Zhang