English
Related papers

Related papers: Policy Learning with $\alpha$-Expected Welfare

200 papers

Many reinforcement learning algorithms, particularly those that rely on return estimates for policy improvement, can suffer from poor sample efficiency and training instability due to high-variance return estimates. In this paper we…

Machine Learning · Computer Science 2026-01-06 Alexander W. Goodall , Edwin Hamel-De le Court , Francesco Belardinelli

We study distributional off-policy evaluation (OPE), of which the goal is to learn the distribution of the return for a target policy using offline data generated by a different policy. The theoretical foundation of many existing work…

Machine Learning · Statistics 2025-03-13 Sungee Hong , Zhengling Qi , Raymond K. W. Wong

This paper develops theoretical criteria and econometric methods to rank policy interventions in terms of welfare when individuals are loss-averse. Our new criterion for "loss aversion-sensitive dominance" defines a weak partial ordering of…

Econometrics · Economics 2023-09-07 Sergio Firpo , Antonio F. Galvao , Martyna Kobus , Thomas Parker , Pedro Rosa-Dias

There is increasing interest in allocating treatments based on observed individual characteristics: examples include targeted marketing, individualized credit offers, and heterogeneous pricing. Treatment personalization introduces…

Econometrics · Economics 2023-04-06 Evan Munro

We study the problem of choosing optimal policy rules in uncertain environments using models that may be incomplete and/or partially identified. We consider a policymaker who wishes to choose a policy to maximize a particular counterfactual…

Econometrics · Economics 2020-12-22 Thomas M. Russell

Large-scale online recommendation systems must facilitate the allocation of a limited number of items among competing users while learning their preferences from user feedback. As a principled way of incorporating market constraints and…

Machine Learning · Computer Science 2022-12-15 Yigit Efe Erginbas , Soham Phade , Kannan Ramchandran

In this paper, we study distributional reinforcement learning from the perspective of statistical efficiency. We investigate distributional policy evaluation, aiming to estimate the complete return distribution (denoted $\eta^\pi$) attained…

Machine Learning · Statistics 2025-11-13 Liangyu Zhang , Yang Peng , Jiadong Liang , Wenhao Yang , Zhihua Zhang

Model-based reinforcement learning (RL) algorithms allow us to combine model-generated data with those collected from interaction with the real system in order to alleviate the data efficiency problem in RL. However, designing such…

Machine Learning · Computer Science 2020-06-25 Yinlam Chow , Brandon Cui , MoonKyung Ryu , Mohammad Ghavamzadeh

We develop an axiomatic framework to evaluate income distributions from the perspective of an opportunity-egalitarian social planner. Building on a formal link with the literature on decision theory under ambiguity, we characterize a class…

Theoretical Economics · Economics 2026-03-31 T. Wienand , B. Magdalou , R. Nock , P. Hufe

Improving the sample efficiency of reinforcement learning algorithms requires effective exploration. Following the principle of $\textit{optimism in the face of uncertainty}$ (OFU), we train a separate exploration policy to maximize the…

Machine Learning · Computer Science 2022-11-23 Jiachen Li , Shuo Cheng , Zhenyu Liao , Huayan Wang , William Yang Wang , Qinxun Bai

Data-driven decision making plays an important role even in high stakes settings like medicine and public policy. Learning optimal policies from observed data requires a careful formulation of the utility function whose expected value is…

Machine Learning · Statistics 2023-11-29 Eli Ben-Michael , Kosuke Imai , Zhichao Jiang

We consider a learning system based on the conventional multiplicative weight (MW) rule that combines experts' advice to predict a sequence of true outcomes. It is assumed that one of the experts is malicious and aims to impose the maximum…

Machine Learning · Computer Science 2020-09-21 S. Rasoul Etesami , Negar Kiyavash , Vincent Leon , H. Vincent Poor

Policymakers often face the decision of how to allocate resources across many different policies using noisy estimates of policy impacts. This paper develops a framework for optimal policy choices under statistical uncertainty. I consider a…

Econometrics · Economics 2026-02-03 Sarah Moon

Large language models (LLMs) are increasingly entrusted with high-stakes decisions that affect human welfare. However, the principles and values that guide these models when distributing scarce societal resources remain largely unexamined.…

Improving social welfare is a complex challenge requiring policymakers to optimize objectives across multiple time horizons. Evaluating the impact of such policies presents a fundamental challenge, as those that appear suboptimal in the…

Computers and Society · Computer Science 2026-02-17 Jiduan Wu , Rediet Abebe , Moritz Hardt , Ana-Andreea Stoica

Policy evaluation estimates the performance of a policy by (1) collecting data from the environment and (2) processing raw data into a meaningful estimate. Due to the sequential nature of reinforcement learning, any improper data-collecting…

Machine Learning · Computer Science 2025-03-21 Shuze Daniel Liu , Claire Chen , Shangtong Zhang

Medical "Crisis Standards of Care" call for a utilitarian allocation of scarce resources in emergencies, while favoring the worst-off under normal conditions. Inspired by such triage rules, we introduce social welfare functions whose…

Theoretical Economics · Economics 2026-03-31 Federico Echenique , Teddy Mekonnen , M. Bumin Yenmez

We consider a social planner faced with a stream of myopic selfish agents. The goal of the social planner is to maximize the social welfare, however, it is limited to using only information asymmetry (regarding previous outcomes) and cannot…

Computer Science and Game Theory · Computer Science 2019-05-15 Lee Cohen , Yishay Mansour

This paper studies offline policy learning, which aims at utilizing observations collected a priori (from either fixed or adaptively evolving behavior policies) to learn an optimal individualized decision rule that achieves the best overall…

Machine Learning · Computer Science 2025-06-06 Ying Jin , Zhimei Ren , Zhuoran Yang , Zhaoran Wang

One of the major concerns of targeting interventions on individuals in social welfare programs is discrimination: individualized treatments may induce disparities across sensitive attributes such as age, gender, or race. This paper…

Econometrics · Economics 2022-07-01 Davide Viviano , Jelena Bradic