English
Related papers

Related papers: Approximation Theory for Lipschitz Continuous Tran…

200 papers

We propose a new scheme for the long time approximation of a diffusion when the drift vector field is not globally Lipschitz. Under this assumption, regular explicit Euler scheme --with constant or decreasing step-- may explode and implicit…

Probability · Mathematics 2018-02-20 Vincent Lemaire

Transformer models have emerged as fundamental tools across various scientific and engineering disciplines, owing to their outstanding performance in diverse applications. Despite this empirical success, the theoretical foundations of…

Machine Learning · Computer Science 2026-04-14 Zhen Qin , Jinxin Zhou , Jiachen Jiang , Zhihui Zhu

Relative Positional Encoding (RPE), which encodes the relative distance between any pair of tokens, is one of the most successful modifications to the original Transformer. As far as we know, theoretical understanding of the RPE-based…

Machine Learning · Computer Science 2022-10-31 Shengjie Luo , Shanda Li , Shuxin Zheng , Tie-Yan Liu , Liwei Wang , Di He

Designing neural networks with bounded Lipschitz constant is a promising way to obtain certifiably robust classifiers against adversarial examples. However, the relevant progress for the important $\ell_\infty$ perturbation setting is…

Machine Learning · Computer Science 2022-10-28 Bohang Zhang , Du Jiang , Di He , Liwei Wang

We consider (stochastic) subgradient methods for strongly convex but potentially nonsmooth non-Lipschitz optimization. We provide new equivalent dual descriptions (in the style of dual averaging) for the classic subgradient method, the…

Optimization and Control · Mathematics 2024-12-31 Benjamin Grimmer , Danlin Li

We propose a new length formula that governs the iterates of the momentum method when minimizing differentiable semialgebraic functions with locally Lipschitz gradients. It enables us to establish local convergence, global convergence, and…

Optimization and Control · Mathematics 2024-01-09 Cédric Josz , Lexiao Lai , Xiaopeng Li

Under general assumptions on the target distribution $p^\star$, we establish a sharp Lipschitz regularity theory for flow-matching vector fields and diffusion-model scores, with optimal dependence on time and dimension. As applications, we…

Statistics Theory · Mathematics 2026-04-08 Arthur Stéphanovitch

This paper proposes that Lipschitz continuity is a natural outcome of regularized least squares in kernel-based learning. Lipschitz continuity is an important proxy for robustness of input-output operators. It is also instrumental for…

Optimization and Control · Mathematics 2021-12-08 Henk J. van Waarde , Rodolphe Sepulchre

To understand the empirical success of approximate MAP inference, recent work (Lang et al., 2018) has shown that some popular approximation algorithms perform very well when the input instance is stable. The simplest stability condition…

Machine Learning · Statistics 2020-11-16 Hunter Lang , David Sontag , Aravindan Vijayaraghavan

Several recent works demonstrate that transformers can implement algorithms like gradient descent. By a careful construction of weights, these works show that multiple layers of transformers are expressive enough to simulate iterations of…

Machine Learning · Computer Science 2023-11-13 Kwangjun Ahn , Xiang Cheng , Hadi Daneshmand , Suvrit Sra

Phase transitions mark qualitative reorganizations of collective behavior, yet identifying their boundaries remains challenging whenever analytic solutions are absent and conventional simulations fail. Here we introduce learnability as a…

Materials Science · Physics 2025-10-10 Şener Özönder

The Lipschitz constant of a neural network is connected to several important properties of the network such as its robustness and generalization. It is thus useful in many settings to estimate the Lipschitz constant of a model. Prior work…

Machine Learning · Computer Science 2026-03-02 Giannis Nikolentzos , Konstantinos Skianis

Recent work has shown that state-of-the-art classifiers are quite brittle, in the sense that a small adversarial change of an originally with high confidence correctly classified input leads to a wrong classification again with high…

Machine Learning · Computer Science 2017-11-07 Matthias Hein , Maksym Andriushchenko

An existence and uniqueness theorem for a class of stochastic delay differential equations is presented, and the convergence of Euler approximations for these equations is proved under general conditions. Moreover, the rate of almost sure…

Probability · Mathematics 2012-12-17 Istvan Gyöngy , Sotirios Sabanis

We provide sufficient conditions for instability of the subgradient method with constant step size around a local minimum of a locally Lipschitz semi-algebraic function. They are satisfied by several spurious local minima arising in robust…

Optimization and Control · Mathematics 2023-06-30 Cédric Josz , Lexiao Lai

Based on the convergence of their infinitesimal generators in the mixed topology, we provide a stability result for strongly continuous convex monotone semigroups on spaces of continuous functions. In contrast to previous results, we do not…

Analysis of PDEs · Mathematics 2026-05-19 Jonas Blessing , Michael Kupper , Max Nendel

Transformer models have redefined sequence learning, yet dot-product self-attention introduces a quadratic token-mixing bottleneck for long-context time-series. We introduce the \textbf{Phasor Transformer} block, a phase-native alternative…

Machine Learning · Computer Science 2026-03-19 Dibakar Sigdel

We introduce the Graded Transformer framework, a new class of sequence models that embeds algebraic inductive biases through grading transformations on vector spaces. Extending Graded Neural Networks (GNNs), we propose two architectures:…

Machine Learning · Computer Science 2025-09-03 Tony Shaska

We consider first-order methods with constant step size for minimizing locally Lipschitz coercive functions that are tame in an o-minimal structure on the real field. We prove that if the method is approximated by subgradient trajectories,…

Optimization and Control · Mathematics 2023-08-03 Cédric Josz , Lexiao Lai

Transformers, which are state-of-the-art in most machine learning tasks, represent the data as sequences of vectors called tokens. This representation is then exploited by the attention function, which learns dependencies between tokens and…

Machine Learning · Computer Science 2025-01-31 Valérie Castin , Pierre Ablin , José Antonio Carrillo , Gabriel Peyré
‹ Prev 1 3 4 5 6 7 10 Next ›