Related papers: Executive stock option exercise with full and part…
In this paper we study optimal trading strategies in a financial market in which stock returns depend on a hidden Gaussian mean reverting drift process. Investors obtain information on that drift by observing stock returns. Moreover, expert…
The extended state observer (ESO) is an inherent element of robust observer-based control systems that allows estimating the impact of disturbance on system dynamics. Proper tuning of ESO parameters is necessary to ensure a good quality of…
This paper is concerned with an optimal reinsurance and investment problem for an insurance firm under the criterion of mean-variance. The driving Brownian motion and the rate in return of the risky asset price dynamic equation cannot be…
This is the last part of four series papers, aiming at stabilization for signal-input-signaloutput (SISO) linear finite-dimensional systems corrupted by general input disturbances. A new observer, referred to as Extended Dynamics Observer…
Safe reinforcement learning (safe RL) aims to respect safety requirements while optimizing long-term performance. In many practical applications, however, the problem involves an infinite number of constraints, known as semi-infinite safe…
We consider an insurance company modelling its surplus process by a Brownian motion with drift. Our target is to maximise the expected exponential utility of discounted dividend payments, given that the dividend rates are bounded by some…
On-policy reinforcement learning (RL) algorithms are widely used for their strong asymptotic performance and training stability, but they struggle to scale with larger batch sizes, as additional parallel environments yield redundant data…
People often express opinions that differ from their privately held views, a phenomenon known in economy as preference falsification. Expressed-private opinion (EPO) models capture this by assigning each agent two dynamical variables: a…
This paper solves a Bayes sequential impulse control problem for a diffusion, whose drift has an unobservable parameter with a change point. The partially-observed problem is reformulated into one with full observations, via a change of…
The construction of an efficient portfolio with a good level of return and minimal risk depends on selecting the optimal combination of stocks. This paper introduces a novel decision-making framework for stock selection based on fractional…
Although resource-limited networked autonomous systems must be able to efficiently and effectively accomplish tasks, better conservation of resources often results in worse task performance. We specifically address the problem of finding…
Learning to perform tasks by leveraging a dataset of expert observations, also known as imitation learning from observations (ILO), is an important paradigm for learning skills without access to the expert reward function or the expert…
We study optimal liquidation strategies under partial information for a single asset within a finite time horizon. We propose a model tailored for high-frequency trading, capturing price formation driven solely by order flow through…
We propose a discrete time algorithm for the valuation of employee stock options based on exponential indifference prices and taking into account both the possibility of partial exercise of a fraction of the options and the use of a…
Standard models in economics stress the role of intelligent agents who maximize utility. However, there may be situations where, for some purposes, constraints imposed by market institutions dominate intelligent agent behavior. We use data…
In this work we study the optimal execution problem with multiplicative price impact in algorithm trading, when an agent holds an initial position of shares of a financial asset. The inter-selling-decision times are modelled by the arrival…
The scaled-dot-product attention (SDPA) mechanism is a core component of modern deep learning, but its mathematical form is often motivated by heuristics. This work provides a first-principles justification for SDPA. We first show that the…
We consider an investor who is dynamically informed about the future evolution of one of the independent Brownian motions driving a stock's price fluctuations. With linear temporary price impact the resulting optimal investment problem with…
We model continuous-time information flows generated by a number of information sources that switch on and off at random times. By modulating a multi-dimensional L\'evy random bridge over a random point field, our framework relates the…
We consider a stochastic game between a slow institutional investor and a high-frequency trader who are trading a risky asset and their aggregated order-flow impacts the asset price. We model this system by means of two coupled stochastic…