English
Related papers

Related papers: Multi-Attribute Utility Preference Robust Optimiza…

200 papers

Bayesian experimental design involves the optimal allocation of resources in an experiment, with the aim of optimising cost and performance. For implicit models, where the likelihood is intractable but sampling from the model is possible,…

Machine Learning · Statistics 2019-02-26 Steven Kleinegesse , Michael Gutmann

Aligning generative models with human preference via RLHF typically suffers from overoptimization, where an imperfectly learned reward model can misguide the generative model to output undesired responses. We investigate this problem in a…

Machine Learning · Computer Science 2024-12-05 Zhihan Liu , Miao Lu , Shenao Zhang , Boyi Liu , Hongyi Guo , Yingxiang Yang , Jose Blanchet , Zhaoran Wang

We propose multi-type probabilistic serial (MPS) and multi-type random priority (MRP) as extensions of the well known PS and RP mechanisms to the multi-type resource allocation problem (MTRA) with partial preferences. In our setting, there…

Artificial Intelligence · Computer Science 2020-10-30 Haibin Wang , Sujoy Sikdar , Xiaoxi Guo , Lirong Xia , Yongzhi Cao , Hanpin Wang

This paper considers facility location problems in which a firm entering a market seeks to open facilities on a subset of candidate locations so as to maximize its expected market share, assuming that customers choose the available…

Optimization and Control · Mathematics 2024-02-19 Robin Legault , Emma Frejinger

The maximum entropy principle can be used to assign utility values when only partial information is available about the decision maker's preferences. In order to obtain such utility values it is necessary to establish an analogy between…

Statistical Finance · Quantitative Finance 2009-11-13 Andreia Dionisio , A. Heitor Reis

Deep neural networks have shown great success in prediction quality while reliable and robust uncertainty estimation remains a challenge. Predictive uncertainty supplements model predictions and enables improved functionality of downstream…

Machine Learning · Computer Science 2021-12-02 Johanna Rock , Tiago Azevedo , René de Jong , Daniel Ruiz-Muñoz , Partha Maji

Recent alignment methods based on Direct Preference Optimization (DPO) reformulate preference learning as supervised optimization over pairwise comparisons, offering improved efficiency and stability over reinforcement learning from human…

Machine Learning · Computer Science 2026-01-22 Yuhui Sun , Xiyao Wang , Zixi Li , YiTian Ding , Tianyang Ling , Jialuo Chen , Tianyi Yu , Zhenlong Yuan , Jinman Zhao

This paper studies a one-sector optimal growth model with i.i.d. productivity shocks that are allowed to be unbounded. The utility function is assumed to be non-negative and unbounded from above. The novel feature in our framework is that…

Economics · Quantitative Finance 2021-07-21 Nicole Bäuerle , Anna Jaśkiewicz

Bootstrapping large language models (LLMs) through preference-based policy optimization offers a promising direction for aligning model behavior with human preferences without relying on extensive manual annotations. In this work, we…

Artificial Intelligence · Computer Science 2025-12-25 Chen Jia

Robust Optimization has traditionally taken a pessimistic, or worst-case viewpoint of uncertainty which is motivated by a desire to find sets of optimal policies that maintain feasibility under a variety of operating conditions. In this…

Machine Learning · Statistics 2017-11-22 Matthew Norton , Akiko Takeda , Alexander Mafusalov

Forecasting multi-step user behavior trajectories requires reasoning over structured preferences across future actions, a challenge overlooked by traditional sequential recommendation. This problem is critical for applications such as…

Information Retrieval · Computer Science 2025-11-04 Hongtao Huang , Chengkai Huang , Junda Wu , Tong Yu , Julian McAuley , Lina Yao

Multi-preference optimization enriches language-model alignment beyond pairwise preferences by contrasting entire sets of helpful and undesired responses, thereby enabling richer training signals for large language models. During self-play…

Machine Learning · Computer Science 2025-06-10 Taneesh Gupta , Rahul Madhavan , Xuchao Zhang , Chetan Bansal , Saravan Rajmohan

In this paper, we settle the problem of learning optimal linear contracts from data in the offline setting, where agent types are drawn from an unknown distribution and the principal's goal is to design a contract that maximizes her…

Computer Science and Game Theory · Computer Science 2026-05-28 Mikael Møller Høgsgaard

We consider the optimal investment and marginal utility pricing problem of a risk averse agent and quantify their exposure to a small amount of model uncertainty. Specifically, we compute explicitly the first-order sensitivity of their…

Mathematical Finance · Quantitative Finance 2021-11-15 Jan Obloj , Johannes Wiesel

Aligning large language models (LLMs) with human preferences is a critical challenge in AI research. While methods like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) are widely used, they often…

Computation and Language · Computer Science 2026-05-19 Xuan Qi , Rongwu Xu , Zhijing Jin

In this work, we consider an optimal control problem subject to a nonlinear PDE constraint and apply it to the regularized $p$-Laplace equation. To this end, a reduced unconstrained optimization problem in terms of the control variable is…

Numerical Analysis · Mathematics 2020-06-29 Bernhard Endtmayer , Ulrich Langer , Ira Neitzel , Winnifried Wollner , Thomas Wick

We consider ordinal approximation algorithms for a broad class of utility maximization problems for multi-agent systems. In these problems, agents have utilities for connecting to each other, and the goal is to compute a maximum-utility…

Multiagent Systems · Computer Science 2017-11-30 Ben Abramowitz , Elliot Anshelevich

Driven by green communications, energy efficiency (EE) has become a new important criterion for designing wireless communication systems. However, high EE often leads to low spectral efficiency (SE), which spurs the research on EE-SE…

Networking and Internet Architecture · Computer Science 2016-05-10 Lei Deng , Wenjie Zhang , Yun Rui , Yeo Chai Kiat

In this paper, we consider a financial market with assets exposed to some risks inducing jumps in the asset prices, and which can still be traded after default times. We use a default-intensity modeling approach, and address in this…

Portfolio Management · Quantitative Finance 2015-10-21 Thomas Lim , Marie-Claire Quenez

The shortcomings of maximum likelihood estimation in the context of model-based reinforcement learning have been highlighted by an increasing number of papers. When the model class is misspecified or has a limited representational capacity,…

Machine Learning · Computer Science 2021-06-08 Evgenii Nikishin , Romina Abachi , Rishabh Agarwal , Pierre-Luc Bacon