Related papers: Variable Annuity with GMWB: surrender or not, that…
The multi-armed bandit (MAB) model is one of the most classical models to study decision-making in an uncertain environment. In this model, a player chooses one of $K$ possible arms of a bandit machine to play at each time step, where the…
We find the variance-optimal equivalent martingale measure when multivariate assets are modeled by a regime-switching geometric Brownian motion, and the regimes are represented by a homogeneous continuous time Markov chain. Under this new…
Robust mechanism design is a rising alternative to Bayesian mechanism design, which yields designs that do not rely on assumptions like full distributional knowledge. We apply this approach to mechanisms for selling a single item, assuming…
We consider the problem of stochastic optimal control, where the state-feedback control policies take the form of a probability distribution and where a penalty on the entropy is added. By viewing the cost function as a Kullback- Leibler…
Inspired by real-time ad exchanges for online display advertising, we consider the problem of inferring a buyer's value distribution for a good when the buyer is repeatedly interacting with a seller through a posted-price mechanism. We…
A combinatorial market consists of a set of indivisible items and a set of agents, where each agent has a valuation function that specifies for each subset of items its value for the given agent. From an optimization point of view, the goal…
Withdrawal guarantees ensure the periodical deduction of a constant dollar-amount from a fund investment for a fixed number of periods. If the fund depletes before the last withdrawal, the guarantor has to finance the outstanding…
We consider a stochastic multi-armed bandit setting where reward must be actively queried for it to be observed. We provide tight lower and upper problem-dependent guarantees on both the regret and the number of queries. Interestingly, we…
This paper explores optimal insurance solutions based on the Lambda-Value-at-Risk ($\Lambda\VaR$). If the expected value premium principle is used, our findings confirm that, similar to the VaR model, a truncated stop-loss indemnity is…
We study the stochastic multi-armed bandit (MAB) problem in the presence of side-observations across actions that occur as a result of an underlying network structure. In our model, a bipartite graph captures the relationship between…
We investigate an optimal reinsurance problem for an insurance company facing a constant fixed cost when the reinsurance contract is signed. The insurer needs to optimally choose both the starting time of the reinsurance contract and the…
Stochastic gradient descent is the method of choice for large-scale machine learning problems, by virtue of its light complexity per iteration. However, it lags behind its non-stochastic counterparts with respect to the convergence rate,…
We consider dynamic pricing with covariates under a generalized linear demand model: a seller can dynamically adjust the price of a product over a horizon of $T$ time periods, and at each time period $t$, the demand of the product is…
We introduce algorithms that achieve state-of-the-art \emph{dynamic regret} bounds for non-stationary linear stochastic bandit setting. It captures natural applications such as dynamic pricing and ads allocation in a changing environment.…
In this paper we address the problem of optimal dividend payout strategies from a surplus process governed by Brownian motion with drift under a drawdown constraint, i.e. the dividend rate can never decrease below a given fraction $a$ of…
We propose an analytically tractable variation of the minority game in which rational agents use probabilistic strategies. In our model, $N$ agents choose between two alternatives repeatedly, and those who are in the minority get a pay-off…
In this paper we study the problem of optimal dividend payment strategy which maximizes the expected discounted sum of dividends to a multidimensional set up of n associated insurance companies where the surplus process follows an…
This paper studies a class of constrained restless multi-armed bandits (CRMAB). The constraints are in the form of time varying set of actions (set of available arms). This variation can be either stochastic or semi-deterministic. Given a…
Multi-armed bandit (MAB) is a widely adopted framework for sequential decision-making under uncertainty. Traditional bandit algorithms rely solely on online data, which tends to be scarce as it must be gathered during the online phase when…
Multi-armed bandit models have proven to be useful in modeling many real world problems in the areas of control and sequential decision making with partial information. However, in many scenarios, such as those prevalent in healthcare and…