English
Related papers

Related papers: Inference of Utilities and Time Preference in Sequ…

200 papers

Pervasive and ubiquitous computing facilitates immediate access to information in the sense of always-on. Information such as news, messages, or reminders can significantly enhance our daily routines but are rendered useless or disturbing…

Human-Computer Interaction · Computer Science 2022-12-20 Christoph Anderson , Judith Simone Heinisch , Shohreh Deldari , Flora D. Salim , Sandra Ohly , Klaus David , Veljko Pejovic

In classic reinforcement learning (RL) and decision making problems, policies are evaluated with respect to a scalar reward function, and all optimal policies are the same with regards to their expected return. However, many real-world…

Machine Learning · Computer Science 2023-11-02 Han Shao , Lee Cohen , Avrim Blum , Yishay Mansour , Aadirupa Saha , Matthew R. Walter

Decision-making problems often feature uncertainty stemming from heterogeneous and context-dependent human preferences. To address this, we propose a sequential learning-and-optimization pipeline to learn preference distributions and…

Machine Learning · Computer Science 2026-03-19 Benjamin Hudson , Laurent Charlin , Emma Frejinger

This paper studies a continuous-time market {under stochastic environment} where an agent, having specified an investment horizon and a target terminal mean return, seeks to minimize the variance of the return with multiple stocks and a…

Portfolio Management · Quantitative Finance 2013-02-28 Wan-Kai Pang , Yuan-Hua Ni , Xun Li , Ka-Fai Cedric Yiu

Continuous-time Markov decision processes are an important class of models in a wide range of applications, ranging from cyber-physical systems to synthetic biology. A central problem is how to devise a policy to control the system in order…

Systems and Control · Computer Science 2016-06-01 Ezio Bartocci , Luca Bortolussi , Tomǎš Brázdil , Dimitrios Milios , Guido Sanguinetti

Forecasting multi-step user behavior trajectories requires reasoning over structured preferences across future actions, a challenge overlooked by traditional sequential recommendation. This problem is critical for applications such as…

Information Retrieval · Computer Science 2025-11-04 Hongtao Huang , Chengkai Huang , Junda Wu , Tong Yu , Julian McAuley , Lina Yao

In domains where users tend to develop long-term preferences that do not change too frequently, the stability of recommendations is an important factor of the perceived quality of a recommender system. In such cases, unstable…

Information Retrieval · Computer Science 2021-04-13 Oluwafemi Olaleke , Ivan Oseledets , Evgeny Frolov

An internet network service provider manages its network with multiple objectives, such as high quality of service (QoS) and minimum computing resource usage. To achieve these objectives, a reinforcement learning-based (RL) algorithm has…

Networking and Internet Architecture · Computer Science 2025-06-17 DongNyeong Heo , Daniela Noemi Rim , Heeyoul Choi

Autonomous robots are increasingly utilized in realistic scenarios with multiple complex tasks. In these scenarios, there may be a preferred way of completing all of the given tasks, but it is often in conflict with optimal execution.…

Robotics · Computer Science 2023-06-26 Peter Amorese , Morteza Lahijanian

We present a scheme for sequential decision making with a risk-sensitive objective and constraints in a dynamic environment. A neural network is trained as an approximator of the mapping from parameter space to space of risk and policy with…

Artificial Intelligence · Computer Science 2019-07-10 Shuai Ma , Jia Yuan Yu , Ahmet Satir

Service supply chain management is to prepare spare parts for failed products under warranty. Their goal is to reach agreed service level at the minimum cost. We convert this business problem into a preference based multi-objective…

Artificial Intelligence · Computer Science 2019-06-20 Wenli Ouyang

We study the learning problem of revealed preference in a stochastic setting: a learner observes the utility-maximizing actions of a set of agents whose utility follows some unknown distribution, and the learner aims to infer the…

Optimization and Control · Mathematics 2022-06-06 John R. Birge , Xiaocheng Li , Chunlin Sun

Reinforcement Learning with Verifiable Rewards (RLVR) demonstrates significant potential in enhancing the reasoning capabilities of Large Language Models (LLMs). However, existing RLVR methods are often constrained by issues such as…

Artificial Intelligence · Computer Science 2026-01-14 Jinpeng Wang , Chao Li , Ting Ye , Mengyuan Zhang , Wei Liu , Jian Luan

We study a continuous-time portfolio choice problem for an investor whose state-dependent preferences are determined by an exogenous factor that evolves as an It\^o diffusion process. Since risk attitudes at the end of the investment…

Mathematical Finance · Quantitative Finance 2025-12-25 Luca De Gennaro Aquino , Sascha Desmettre , Yevhen Havrylenko , Mogens Steffensen

This paper investigates portfolio selection within a continuous-time financial market with regime-switching and beliefs-dependent utilities. The market coefficients and the investor's utility function both depend on the market regime, which…

Optimization and Control · Mathematics 2024-10-23 Xiaochen Chen , Guohui Guan , Zongxia Liang

In this paper, we consider the revealed preferences problem from a learning perspective. Every day, a price vector and a budget is drawn from an unknown distribution, and a rational agent buys his most preferred bundle according to some…

Computer Science and Game Theory · Computer Science 2012-11-20 Morteza Zadimoghaddam , Aaron Roth

Time series models predict numbers; decision-makers need advisory -- directional signals with reasoning, actionable suggestions, and risk management. Training language models for such predictive advisory faces a fundamental challenge:…

Machine Learning · Computer Science 2026-04-28 Yanwei Cui , Guanghui Wang , Xing Zhang , Peiyang He , Ziyuan Li , Bing Zhu , Wei Qiu , Xusheng Wang , Zheng Yu , Anqi Xin

Large-scale online recommendation systems must facilitate the allocation of a limited number of items among competing users while learning their preferences from user feedback. As a principled way of incorporating market constraints and…

Machine Learning · Computer Science 2022-12-15 Yigit Efe Erginbas , Soham Phade , Kannan Ramchandran

The paper [12] examines a concept of equilibrium policies instead of optimal controls in stochastic optimization to analyze a mean-variance portfolio selection problem. We follow the same approach in order to investigate the Merton…

Optimization and Control · Mathematics 2020-04-23 I. Alia , F. Chighoub , N. Khelfallah , J. Vives

This work extends a previous work in regime detection, which allowed trading positions to be profitably adjusted when a new regime was detected, to ex ante prediction of regimes, leading to substantial performance improvements over the…

Risk Management · Quantitative Finance 2023-10-10 Piotr Pomorski , Denise Gorse