Related papers: Entropy-regularized penalization schemes and refle…
Bandit algorithms sequentially accumulate data using adaptive sampling policies, offering flexibility for real-world applications. However, excessive sampling can be costly, motivating the devolopment of early stopping methods and reliable…
Large Reasoning Models (LRMs) have demonstrated remarkable performance on complex reasoning tasks by employing test-time scaling. However, they often generate over-long chains-of-thought that, driven by substantial reflections such as…
We study a robust optimal stopping problem with respect to a set $\cP$ of mutually singular probabilities. This can be interpreted as a zero-sum controller-stopper game in which the stopper is trying to maximize its pay-off while an adverse…
Early stopping of iterative algorithms is a widely-used form of regularization in statistics, commonly used in conjunction with boosting and related gradient-type algorithms. Although consistency results have been established in some…
Evolutionary strategies have recently been shown to achieve competing levels of performance for complex optimization problems in reinforcement learning. In such problems, one often needs to optimize an objective function subject to a set of…
We first introduce the concept of $\mathscr{Y}^{g,\xi}$-submartingale systems, where the nonlinear operator $\mathscr{Y}^{g,\xi}$ corresponds to the first component of the solution of a reflected BSDE with generator $g$ and lower obstacle…
In constrained Markov decision processes (CMDPs) with adversarial rewards and constraints, a well-known impossibility result prevents any algorithm from attaining both sublinear regret and sublinear constraint violation, when competing…
In this paper, we provide a new algorithm for the problem of prediction in Reinforcement Learning, \emph{i.e.}, estimating the Value Function of a Markov Reward Process (MRP) using the linear function approximation architecture, with memory…
This paper proves the existence and uniqueness of a solution to doubly reflected backward stochastic differential equations where the coefficient is stochastic Lipschitz, by means of the penalization method.
In this paper we establish the convergence of a numerical scheme based, on the Finite Element Method, for a time-independent problem modelling the deformation of a linearly elastic elliptic membrane shell subjected to remaining confined in…
Sequential Bayesian experimental design typically assumes that the number of experiments is fixed before data collection begins. In practical campaigns, however, experimentation may need to terminate early because additional measurements…
In this paper, we address the problem of detecting anomalies among a given set of binary processes via learning-based controlled sensing. Each process is parameterized by a binary random variable indicating whether the process is anomalous.…
We propose an entropic approximation approach for optimal transportation problems with a supremal cost. We establish $\Gamma$-convergence for suitably chosen parameters for the entropic penalization and that this procedure selects…
We obtain existence and uniqueness in L^p, p>1 of the solutions of a backward stochastic differential equations (BSDEs for short) driven by a marked point process, on a bounded interval. We show that the solution of the BSDE can be…
We study a two armed-bandit algorithm with penalty. We show the convergence of the algorithm and establish the rate of convergence. For some choices of the parameters, we obtain a central limit theorem in which the limit distribution is…
We show a concise extension of the monotone stability approach to backward stochastic differential equations (BSDEs) that are jointly driven by a Brownian motion and a random measure for jumps, which could be of infinite activity with a…
We consider a discounted infinite horizon optimal stopping problem. If the underlying distribution is known a priori, the solution of this problem is obtained via dynamic programming (DP) and is given by a well known threshold rule. When…
We develop a covariant formalism to investigate the mixed state entanglement structure of time-dependent boosted subsystems in $\textrm{T}\bar{\textrm{T}}$ deformed CFT$_2$s through the reflected entropy. To this end we utilize the…
Maximum likelihood estimation of energy-based models is a challenging problem due to the intractability of the log-likelihood gradient. In this work, we propose learning both the energy function and an amortized approximate sampling…
In a noise driving by a multivariate point process $\mu$ with predictable compensator $\nu$, we prove existence and uniqueness of the reflected backward stochastic differential equation's solution with a lower obstacle…