English
Related papers

Related papers: Entropy-regularized penalization schemes and refle…

200 papers

Bandit algorithms sequentially accumulate data using adaptive sampling policies, offering flexibility for real-world applications. However, excessive sampling can be costly, motivating the devolopment of early stopping methods and reliable…

Statistics Theory · Mathematics 2025-02-06 Zihan Cui

Large Reasoning Models (LRMs) have demonstrated remarkable performance on complex reasoning tasks by employing test-time scaling. However, they often generate over-long chains-of-thought that, driven by substantial reflections such as…

Artificial Intelligence · Computer Science 2026-03-02 Zewei Yu , Lirong Gao , Yuke Zhu , Bo Zheng , Junbo Zhao , Sheng Guo , Haobo Wang

We study a robust optimal stopping problem with respect to a set $\cP$ of mutually singular probabilities. This can be interpreted as a zero-sum controller-stopper game in which the stopper is trying to maximize its pay-off while an adverse…

Probability · Mathematics 2016-04-12 Erhan Bayraktar , Song Yao

Early stopping of iterative algorithms is a widely-used form of regularization in statistics, commonly used in conjunction with boosting and related gradient-type algorithms. Although consistency results have been established in some…

Machine Learning · Statistics 2018-03-15 Yuting Wei , Fanny Yang , Martin J. Wainwright

Evolutionary strategies have recently been shown to achieve competing levels of performance for complex optimization problems in reinforcement learning. In such problems, one often needs to optimize an objective function subject to a set of…

Neural and Evolutionary Computing · Computer Science 2022-02-23 Youssef Diouane , Aurelien Lucchi , Vihang Patil

We first introduce the concept of $\mathscr{Y}^{g,\xi}$-submartingale systems, where the nonlinear operator $\mathscr{Y}^{g,\xi}$ corresponds to the first component of the solution of a reflected BSDE with generator $g$ and lower obstacle…

Optimization and Control · Mathematics 2023-05-26 Roxana Dumitrescu , Romuald Elie , Wissal Sabbagh , Chao Zhou

In constrained Markov decision processes (CMDPs) with adversarial rewards and constraints, a well-known impossibility result prevents any algorithm from attaining both sublinear regret and sublinear constraint violation, when competing…

Machine Learning · Computer Science 2024-09-27 Francesco Emanuele Stradi , Anna Lunghi , Matteo Castiglioni , Alberto Marchesi , Nicola Gatti

In this paper, we provide a new algorithm for the problem of prediction in Reinforcement Learning, \emph{i.e.}, estimating the Value Function of a Markov Reward Process (MRP) using the linear function approximation architecture, with memory…

Systems and Control · Computer Science 2016-09-30 Ajin George Joseph , Shalabh Bhatnagar

This paper proves the existence and uniqueness of a solution to doubly reflected backward stochastic differential equations where the coefficient is stochastic Lipschitz, by means of the penalization method.

Probability · Mathematics 2018-01-04 Mohamed Marzougue , Mohamed El Otmani

In this paper we establish the convergence of a numerical scheme based, on the Finite Element Method, for a time-independent problem modelling the deformation of a linearly elastic elliptic membrane shell subjected to remaining confined in…

Analysis of PDEs · Mathematics 2023-10-25 Aaron Meixner , Paolo Piersanti

Sequential Bayesian experimental design typically assumes that the number of experiments is fixed before data collection begins. In practical campaigns, however, experimentation may need to terminate early because additional measurements…

Methodology · Statistics 2026-05-29 Chen Cheng , Xun Huan

In this paper, we address the problem of detecting anomalies among a given set of binary processes via learning-based controlled sensing. Each process is parameterized by a binary random variable indicating whether the process is anomalous.…

Machine Learning · Computer Science 2023-12-04 Geethu Joseph , Chen Zhong , M. Cenk Gursoy , Senem Velipasalar , Pramod K. Varshney

We propose an entropic approximation approach for optimal transportation problems with a supremal cost. We establish $\Gamma$-convergence for suitably chosen parameters for the entropic penalization and that this procedure selects…

Analysis of PDEs · Mathematics 2023-02-24 Guillaume Carlier , Camilla Brizzi , Luigi De Pascale

We obtain existence and uniqueness in L^p, p>1 of the solutions of a backward stochastic differential equations (BSDEs for short) driven by a marked point process, on a bounded interval. We show that the solution of the BSDE can be…

Probability · Mathematics 2016-12-04 Fulvia Confortola

We study a two armed-bandit algorithm with penalty. We show the convergence of the algorithm and establish the rate of convergence. For some choices of the parameters, we obtain a central limit theorem in which the limit distribution is…

Probability · Mathematics 2016-08-16 Damien Lamberton , Gilles Pagès

We show a concise extension of the monotone stability approach to backward stochastic differential equations (BSDEs) that are jointly driven by a Brownian motion and a random measure for jumps, which could be of infinite activity with a…

Probability · Mathematics 2019-11-21 Dirk Becherer , Martin Büttner , Klebert Kentia

We consider a discounted infinite horizon optimal stopping problem. If the underlying distribution is known a priori, the solution of this problem is obtained via dynamic programming (DP) and is given by a well known threshold rule. When…

Machine Learning · Computer Science 2021-02-23 Daniel Russo , Assaf Zeevi , Tianyi Zhang

We develop a covariant formalism to investigate the mixed state entanglement structure of time-dependent boosted subsystems in $\textrm{T}\bar{\textrm{T}}$ deformed CFT$_2$s through the reflected entropy. To this end we utilize the…

High Energy Physics - Theory · Physics 2024-02-13 Debarshi Basu , Vinayak Raj

Maximum likelihood estimation of energy-based models is a challenging problem due to the intractability of the log-likelihood gradient. In this work, we propose learning both the energy function and an amortized approximate sampling…

Machine Learning · Computer Science 2019-05-29 Rithesh Kumar , Sherjil Ozair , Anirudh Goyal , Aaron Courville , Yoshua Bengio

In a noise driving by a multivariate point process $\mu$ with predictable compensator $\nu$, we prove existence and uniqueness of the reflected backward stochastic differential equation's solution with a lower obstacle…

Probability · Mathematics 2023-10-03 Brahim Baadi , Mohamed Marzougue