Related papers: Optimal consumption under adjustment costs with re…
Preference-based feedback is important for many applications in machine learning where evaluation of a reward function is not feasible. Notable recent examples arise in preference alignment for large language models, including in…
We analyze the consumption-portfolio selection problem of an investor facing both Brownian and jump risks. We bring new tools, in the form of orthogonal decompositions, to bear on the problem in order to determine the optimal portfolio in…
Adaptive optimal control using value iteration initiated from a stabilizing control policy is theoretically analyzed in terms of stability of the system during the learning stage without ignoring the effects of approximation errors. This…
This paper explores a new class of incomplete preferences -- termed ``connected preferences'' -- in which maximal domains of comparability are topologically connected. We provide necessary and sufficient conditions for continuous…
A long-standing question about consumer behavior is whether individuals' observed purchase decisions satisfy the revealed preference (RP) axioms of the utility maximization theory (UMT). Researchers using survey or experimental panel data…
We revisit the classical Merton consumption--investment problem when risky-asset returns are modeled by stochastic differential equations interpreted through a general $\alpha$-integral, interpolating between It\^{o}, Stratonovich, and…
A speculative agent with Prospect Theory preference chooses the optimal time to purchase and then to sell an indivisible risky asset to maximize the expected utility of the round-trip profit net of transaction costs. The optimization…
A rational behavior of a consumer is analyzed when the user participates in a Peak Time Rebate (PTR) mechanism, which is a demand response (DR) incentive program based on a baseline. A multi-stage stochastic programming is proposed from the…
We study the optimal control of battery energy storage under a general "pay-for-performance" setup such as providing frequency regulation and renewable integration. In these settings, batteries need to carefully balance the trade-off…
Smoothed online combinatorial optimization considers a learner who repeatedly chooses a combinatorial decision to minimize an unknown changing cost function with a penalty on switching decisions in consecutive rounds. We study smoothed…
We establish when the two problems of minimizing a function of lifetime minimum wealth and of maximizing utility of lifetime consumption result in the same optimal investment strategy on a given open interval $O$ in wealth space. To answer…
Coupon allocation drives customer purchases and boosts revenue. However, it presents a fundamental trade-off between exploiting the current optimal policy to maximize immediate revenue and exploring alternative policies to collect data for…
In this paper we study simulation based optimization algorithms for solving discrete time optimal stopping problems. This type of algorithms became popular among practioneers working in the area of quantitative finance. Using large…
Cultures around the world show varying levels of conservatism. While maintaining traditional ideas prevents wrong ones from being embraced, it also slows or prevents adaptation to new times. Without exploration there can be no improvement,…
We consider the impulse control of Levy processes under the infinite horizon, discounted cost criterion. Our motivating example is the cash management problem in which a controller is charged a fixed plus proportional cost for adding to or…
In this paper, we study how a budget-constrained bidder should learn to bid adaptively in repeated first-price auctions to maximize cumulative payoff. This problem arises from the recent industry-wide shift from second-price auctions to…
In this work, we consider an optimal control problem subject to a nonlinear PDE constraint and apply it to the regularized $p$-Laplace equation. To this end, a reduced unconstrained optimization problem in terms of the control variable is…
In this paper, we study the optimal investment problem of an insurer whose surplus process follows the diffusion approximation of the classical Cramer-Lundberg model. Investment in the foreign market is allowed, and therefore, the foreign…
This paper studies a continuous-time optimal portfolio selection problem in the complete market for a behavioral investor whose preference is of the prospect type with probability distortion. The investor concerns about the terminal…
As is well known, average-cost optimality inequalities imply the existence of stationary optimal policies for Markov Decision Processes with average costs per unit time, and these inequalities hold under broad natural conditions. This paper…