English
Related papers

Related papers: Insurance Pricing Optimization via Off-Policy Eval…

200 papers

We study off-policy evaluation (OPE) in the problem of slate contextual bandits where a policy selects multi-dimensional actions known as slates. This problem is widespread in recommender systems, search engines, marketing, to medical…

Machine Learning · Statistics 2024-02-20 Haruka Kiyohara , Masahiro Nomura , Yuta Saito

In some applications of reinforcement learning, a dataset of pre-collected experience is already available but it is also possible to acquire some additional online data to help improve the quality of the policy. However, it may be…

Machine Learning · Computer Science 2023-07-11 Ruiqi Zhang , Andrea Zanette

Policy optimization is a core component of reinforcement learning (RL), and most existing RL methods directly optimize parameters of a policy based on maximizing the expected total reward, or its surrogate. Though often achieving…

Machine Learning · Computer Science 2018-08-10 Ruiyi Zhang , Changyou Chen , Chunyuan Li , Lawrence Carin

We propose a general method for deriving prognostics-based predictive maintenance policies. The method takes into account the available decision options at hand, the information on the future state of the system provided by a prognostic…

Optimization and Control · Mathematics 2025-10-10 Daniel Koutas , Daniel Straub

In the present work we tackle the problem of finding the optimal price tariff to be set by a risk-averse electric retailer participating in the pool and whose customers are price-sensitive. We assume that the retailer has access to a…

Optimization and Control · Mathematics 2022-02-24 Román Pérez-Santalla , Miguel Carrión , Carlos Ruiz

The net-premium principle is considered to be the most genuine and fair premium principle in actuarial applications. However, an insurance company, applying the net-premium principle, goes bankrupt with probability one in the long run, even…

Risk Management · Quantitative Finance 2013-04-03 Alois Pichler

We propose a simple randomized rule for the optimization of prices in revenue management with contextual information. It is known that the certainty equivalent pricing rule, albeit popular, is sub-optimal. We show that, by allowing a small…

Computer Science and Game Theory · Computer Science 2020-10-26 Neil Walton , Yuqing Zhang

Inverse optimization refers to the inference of unknown parameters of an optimization problem based on knowledge of its optimal solutions. This paper considers inverse optimization in the setting where measurements of the optimal solutions…

Optimization and Control · Mathematics 2017-12-27 Anil Aswani , Zuo-Jun Max Shen , Auyon Siddiq

Randomized trials, also known as A/B tests, are used to select between two policies: a control and a treatment. Given a corresponding set of features, we can ideally learn an optimized policy P that maps the A/B test data features to action…

Machine Learning · Computer Science 2018-06-08 Elon Portugaly , Joseph J. Pfeiffer

Offline reinforcement learning (RL) presents distinct challenges as it relies solely on observational data. A central concern in this context is ensuring the safety of the learned policy by quantifying uncertainties associated with various…

Machine Learning · Computer Science 2025-07-03 Xiaocong Chen , Siyu Wang , Tong Yu , Lina Yao

Abstract In this work, we build two environments, namely the modified QLBS and RLOP models, from a mathematics perspective which enables RL methods in option pricing through replicating by portfolio. We implement the environment…

Pricing of Securities · Quantitative Finance 2022-05-12 Ziheng Chen

We consider dynamic pricing schemes in online settings where selfish agents generate online events. Previous work on online mechanisms has dealt almost entirely with the goal of maximizing social welfare or revenue in an auction settings.…

Computer Science and Game Theory · Computer Science 2015-04-07 Ilan Reuven Cohen , Alon Eden , Amos Fiat , Łukasz Jeż

Our goal is to compute a policy that guarantees improved return over a baseline policy even when the available MDP model is inaccurate. The inaccurate model may be constructed, for example, by system identification techniques when the true…

Optimization and Control · Mathematics 2015-06-17 Yinlam Chow , Marek Petrik , Mohammad Ghavamzadeh

Demand response (DR) has been demonstrated to be an effective method for reducing peak load and mitigating uncertainties on both the supply and demand sides of the electricity market. One critical question for DR research is how to…

Machine Learning · Computer Science 2023-06-27 Jun Song , Chaoyue Zhao

Probabilistic learning to rank (LTR) has been the dominating approach for optimizing the ranking metric, but cannot maximize long-term rewards. Reinforcement learning models have been proposed to maximize user long-term rewards by…

Machine Learning · Computer Science 2024-01-18 Teng Xiao , Suhang Wang

Model-based algorithms, which learn a dynamics model from logged experience and perform some sort of pessimistic planning under the learned model, have emerged as a promising paradigm for offline reinforcement learning (offline RL).…

Machine Learning · Computer Science 2022-01-28 Tianhe Yu , Aviral Kumar , Rafael Rafailov , Aravind Rajeswaran , Sergey Levine , Chelsea Finn

Off-policy evaluation (OPE) estimates the value of a target treatment policy (e.g., a recommender system) using data collected by a different logging policy. It enables high-stakes experimentation without live deployment, yet in practice…

Machine Learning · Statistics 2026-05-18 Connor Douglas , Joel Persson , Foster Provost

Traditional reinforcement learning methods optimize agents without considering safety, potentially resulting in unintended consequences. In this paper, we propose an optimal actor-free policy that optimizes a risk-sensitive criterion based…

Machine Learning · Computer Science 2023-07-04 Ruoqi Zhang , Jens Sjölund

Recommender systems predict what items a user will interact with next, based on their past interactions. The problem is often approached through supervised learning, but recent advancements have shifted towards policy optimization of…

Machine Learning · Computer Science 2023-04-28 Dawen Liang , Nikos Vlassis

We study the optimal investment and proportional reinsurance problem of an insurance company, whose investment preferences are described via a forward dynamic utility of exponential type in a stochastic factor model allowing for a possible…

Mathematical Finance · Quantitative Finance 2022-10-20 Katia Colaneri , Alessandra Cretarola , Benedetta Salterini