English
Related papers

Related papers: Multi-Attribute Utility Preference Robust Optimiza…

200 papers

Since the debut of DPO, it has been shown that aligning a target LLM with human preferences via the KL-constrained RLHF loss is mathematically equivalent to a special kind of reward modeling task. Concretely, the task requires: 1) using the…

Machine Learning · Computer Science 2024-12-19 Yuzhong Hong , Hanshan Zhang , Junwei Bao , Hongfei Jiang , Yang Song

Reinforcement Learning (RL) has emerged as a powerful tool for neural combinatorial optimization, enabling models to learn heuristics that solve complex problems without requiring expert knowledge. Despite significant progress, existing RL…

Machine Learning · Computer Science 2025-05-14 Mingjun Pan , Guanquan Lin , You-Wei Luo , Bin Zhu , Zhien Dai , Lijun Sun , Chun Yuan

This paper describes a novel approach to planning which takes advantage of decision theory to greatly improve robustness in an uncertain environment. We present an algorithm which computes conditional plans of maximum expected utility. This…

Artificial Intelligence · Computer Science 2013-02-28 Stephen G. Pimentel , Lawrence M. Brem

We consider a discrete time financial market with proportional transaction costs under model uncertainty, and study a num\'eraire-based semi-static utility maximization problem with an exponential utility preference. The randomization…

Mathematical Finance · Quantitative Finance 2019-08-02 Shuoqing Deng , Xiaolu Tan , Xiang Yu

Direct alignment methods typically train large language models (LLMs) by contrasting the likelihoods of preferred and dispreferred responses. While effective at capturing relative preferences, these methods are widely observed to suppress…

Computation and Language · Computer Science 2025-12-04 Kaiyang Guo , Yinchuan Li , Zhitang Chen

This paper solves a utility maximization problem under utility-based shortfall risk constraint, by proposing an approach using Lagrange multiplier and convex duality. Under mild conditions on the asymptotic elasticity of the utility…

Mathematical Finance · Quantitative Finance 2016-06-28 Oliver Janke , Qinghua Li

In this paper, we develop a new method for finding an optimal biddingstrategy in sequential auctions, using a dynamic programming technique. Theexisting method assumes that the utility of a user is represented in anadditive form. Thus, the…

Computer Science and Game Theory · Computer Science 2013-01-14 Hiromitsu Hattori , Makoto Yokoo , Yuko Sakurai , Toramatsu Shintani

Offline reinforcement learning aims to learn from pre-collected datasets without active exploration. This problem faces significant challenges, including limited data availability and distributional shifts. Existing approaches adopt a…

Machine Learning · Computer Science 2024-10-01 Yue Wang , Jinjun Xiong , Shaofeng Zou

Language model (LM) post-training (or alignment) involves maximizing a reward function that is derived from preference annotations. Direct Preference Optimization (DPO) is a popular offline alignment method that trains a policy directly on…

Machine Learning · Computer Science 2025-03-04 Adam Fisch , Jacob Eisenstein , Vicky Zayats , Alekh Agarwal , Ahmad Beirami , Chirag Nagpal , Pete Shaw , Jonathan Berant

Simulation-based optimal design techniques are a convenient tool for solving a particular class of optimal design problems. The goal is to find the optimal configuration of factor settings with respect to an expected utility criterion. This…

Methodology · Statistics 2013-05-21 Markus Hainy , Werner G. Müller , Helga Wagner

Recent advancements have established Reinforcement Learning (RL) as a pivotal paradigm for aligning generative models with human intent. However, group-based optimization frameworks (e.g., GRPO) face a critical limitation: the rapid decay…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Sujie Hu , Chubin Chen , Jiashu Zhu , Jiahong Wu , Xiangxiang Chu , Xiu Li

We study the dynamic assortment planning problem, where for each arriving customer, the seller offers an assortment of substitutable products and customer makes the purchase among offered products according to an uncapacitated multinomial…

Machine Learning · Statistics 2019-02-11 Xi Chen , Yining Wang , Yuan Zhou

In this paper we consider multi-objective reinforcement learning where the objectives are balanced using preferences. In practice, the preferences are often given in an adversarial manner, e.g., customers can be picky in many applications.…

Machine Learning · Computer Science 2021-10-29 Jingfeng Wu , Vladimir Braverman , Lin F. Yang

Preference based alignment objectives implicitly assume that all human preferences are expressed with equal strength. In practice, however, preference strength varies across individuals and contexts -- a phenomenon established in behavioral…

Machine Learning · Computer Science 2026-01-13 Saki Imai , Pedram Heydari , Anthony Sicilia , Asteria Kaeberlein , Katherine Atwell , Malihe Alikhani

Optimization problems find widespread use in both single-objective and multi-objective scenarios. In practical applications, users aspire for solutions that converge to the region of interest (ROI) along the Pareto front (PF). While the…

Artificial Intelligence · Computer Science 2025-03-10 Tian Huang , Shengbo Wang , Ke Li

Prediction deviations of different uncertainties have varying impacts on downstream decision-making. Improving the prediction accuracy of critical uncertainties with significant impacts on decision-making quality yields better optimization…

Systems and Control · Electrical Eng. & Systems 2025-10-17 Yingrui Zhuang , Lin Cheng , Can Wan , Rui Xie , Ning Qi , Yue Chen

We study a problem of utility maximization under model uncertainty with information including jumps. We prove first that the value process of the robust stochastic control problem is described by the solution of a quadratic-exponential…

Probability · Mathematics 2016-10-11 Monique Jeanblanc , Anis Matoussi , Armand Ngoupeyou

This paper studies optimal motion planning subject to motion and environment uncertainties. By modeling the system as a probabilistic labeled Markov decision process (PL-MDP), the control objective is to synthesize a finite-memory policy,…

Robotics · Computer Science 2022-01-03 Mingyu Cai , Shaoping Xiao , Zhijun Li , Zhen Kan

We consider the system-wide utility maximization problem in the downlink of a cell-free massive multiple-input multiple-output (MIMO) system whereby a very large number of access points (APs) simultaneously serve a group of users.…

Signal Processing · Electrical Eng. & Systems 2020-09-22 Muhammad Farooq , Hien Quoc Ngo , Een-Kee Hong , Le-Nam Tran

In this paper, we consider an adaptive approach to address optimization problems with uncertain cost parameters. Here, the decision maker selects an initial decision, observes the realization of the uncertain cost parameters, and then is…

Computational Complexity · Computer Science 2013-12-17 Ebrahim Nasrabadi , James B. Orlin