Related papers: Refraction-reflection strategies in the dual model
Selective prediction, where a model has the option to abstain from making a decision, is crucial for machine learning applications in which mistakes are costly. In this work, we focus on distributional regression and introduce a framework…
We study approximations of reflected It\^o diffusions on convex subsets $D$ of $\Rd$ by solutions of stochastic differential equations with penalization terms. We assume that the diffusion coefficients are merely measurable (possibly…
This work develops a distributed optimization strategy with guaranteed exact convergence for a broad class of left-stochastic combination policies. The resulting exact diffusion strategy is shown in Part II to have a wider stability range…
In this paper we study the joint distributions of the telegraph process and its maximum conditioned on the number of changes of direction and the initial velocity. We prove that in the case of positive starting velocity, a form of the…
The problem of reinforcement learning is considered where the environment or the model undergoes a change. An algorithm is proposed that an agent can apply in such a problem to achieve the optimal long-time discounted reward. The algorithm…
Bridging the gap between diffusion models and human preferences is crucial for their integration into practical generative workflows. While optimizing downstream reward models has emerged as a promising alignment strategy, concerns arise…
The optimization criterion for dividends from a risky business is most often formalized in terms of the expected present value of future dividends. That criterion disregards a potential, explicit demand for stability of dividends. In…
This paper investigates a robust optimal consumption, investment, and reinsurance problem for an insurer with Epstein-Zin recursive preferences operating under model uncertainty. The insurer's surplus follows the diffusion approximation of…
In this paper, we use replica analysis to determine the investment strategy that can maximize the net present value for portfolios containing multiple development projects. Replica analysis was developed in statistical mechanical…
We present a formulation of an optimal control problem for a two-dimensional diffusion process governed by a Fokker-Planck equation to achieve a nonequilibrium steady state with a desired circulation while accelerating convergence toward…
Reinforcement learning has long struggled with poor sample efficiency. One promising approach to mitigate this problem is leveraging group-invariant Markov Decision Processes ($G$-invariant MDPs). Existing works in this direction have…
Motivated by recent developments in risk management based on the U.S. bankruptcy code, we revisit the De Finetti's optimal dividend problem by incorporating the reorganization process and regulator's intervention documented in Chapter 11…
Model-based reinforcement learning methods often use learning only for the purpose of estimating an approximate dynamics model, offloading the rest of the decision-making work to classical trajectory optimizers. While conceptually simple,…
We study the behavior of deterministic methods for solving inverse problems in imaging. These methods are commonly designed to achieve two goals: (1) attaining high perceptual quality, and (2) generating reconstructions that are consistent…
We study optimal buying and selling strategies in target zone models. In these models the price is modeled by a diffusion process which is reflected at one or more barriers. Such models arise for example when a currency exchange rate is…
We consider the problem of optimal reactive power compensation for the minimization of power distribution losses in a smart microgrid. We first propose an approximate model for the power distribution network, which allows us to cast the…
In the last few years there has been renewed interest in the classical control problem of de Finetti for the case that underlying source of randomness is a spectrally negative Levy process. In particular a significant step forward is made…
We study the dynamic portfolio selection of an investor who uses deep learning methods to forecast stock market excess returns. In a two-asset allocation problem, deep neural networks -- both feedforward and long short-term memory (LSTM)…
This paper studies stochastic control problems motivated by optimal consumption with wealth benchmark tracking. The benchmark process is modeled by a combination of a geometric Brownian motion and a running maximum process, indicating its…
Diffusion models excel at modeling complex data distributions, including those of images, proteins, and small molecules. However, in many cases, our goal is to model parts of the distribution that maximize certain properties: for example,…