English
Related papers

Related papers: Monotone, Bi-Lipschitz, and Polyak-Lojasiewicz Net…

200 papers

This article provides a comprehensive understanding of optimization in deep learning, with a primary focus on the challenges of gradient vanishing and gradient exploding, which normally lead to diminished model representational ability and…

Machine Learning · Computer Science 2023-11-14 Xianbiao Qi , Jianan Wang , Lei Zhang

The (global) Lipschitz smoothness condition is crucial in establishing the convergence theory for most optimization methods. Unfortunately, most machine learning and signal processing problems are not Lipschitz smooth. This motivates us to…

Optimization and Control · Mathematics 2019-04-23 Qiuwei Li , Zhihui Zhu , Gongguo Tang , Michael B. Wakin

Nonlinear Parametric Optimization Network (NLPOpt-Net) is an unsupervised learning architecture to solve constrained nonlinear programs (NLP). Given the structure of an NLP, it learns the parametric solution maps with guaranteed constraint…

Machine Learning · Computer Science 2026-05-04 Bimol Nath Roy , Rahul Golder , MM Faruque Hasan

The Jacobian matrix (or the gradient for single-output networks) is directly related to many important properties of neural networks, such as the function landscape, stationary points, (local) Lipschitz constants and robustness to…

Machine Learning · Statistics 2019-02-28 Huan Zhang , Pengchuan Zhang , Cho-Jui Hsieh

Existing bounds on the generalization error of deep networks assume some form of smooth or bounded dependence on the input variable, falling short of investigating the mechanisms controlling such factors in practice. In this work, we…

Machine Learning · Computer Science 2025-07-24 Matteo Gamba , Hossein Azizpour , Mårten Björkman

We introduce LilNetX, an end-to-end trainable technique for neural networks that enables learning models with specified accuracy-rate-computation trade-off. Prior works approach these problems one at a time and often require post-processing…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Sharath Girish , Kamal Gupta , Saurabh Singh , Abhinav Shrivastava

In decentralized optimization, several nodes connected by a network collaboratively minimize some objective function. For minimization of Lipschitz functions lower bounds are known along with optimal algorithms. We study a specific class of…

Optimization and Control · Mathematics 2023-03-15 Savelii Chezhegov , Alexander Rogozin , Alexander Gasnikov

Bilevel optimization is a hierarchical framework where an upper-level optimization problem is constrained by a lower-level problem, commonly used in machine learning applications such as hyperparameter optimization. Existing bilevel…

Optimization and Control · Mathematics 2026-03-03 Yuman Wu , Xiaochuan Gong , Jie Hao , Mingrui Liu

In this paper, a new method of H_infinity observer design for Lipschitz nonlinear systems is proposed in the form of an LMI optimization problem. The proposed observer has guaranteed decay rate (exponential convergence) and is robust…

Systems and Control · Computer Science 2010-10-06 Masoud Abbaszadeh , Horacio J. Marquez

Bilevel optimization is an important formulation for many machine learning problems. Current bilevel optimization algorithms assume that the gradient of the upper-level function is Lipschitz. However, recent studies reveal that certain…

Machine Learning · Computer Science 2024-01-19 Jie Hao , Xiaochuan Gong , Mingrui Liu

We demonstrate two new important properties of the 1-path-norm of shallow neural networks. First, despite its non-smoothness and non-convexity it allows a closed form proximal operator which can be efficiently computed, allowing the use of…

Machine Learning · Computer Science 2020-07-16 Fabian Latorre , Paul Rolland , Nadav Hallak , Volkan Cevher

Monotonic linear interpolation (MLI) - on the line connecting a random initialization with the minimizer it converges to, the loss and accuracy are monotonic - is a phenomenon that is commonly observed in the training of neural networks.…

Machine Learning · Statistics 2023-02-15 Xiang Wang , Annie N. Wang , Mo Zhou , Rong Ge

We introduce Invertible Dense Networks (i-DenseNets), a more parameter efficient extension of Residual Flows. The method relies on an analysis of the Lipschitz continuity of the concatenation in DenseNets, where we enforce invertibility of…

Machine Learning · Statistics 2021-10-26 Yura Perugachi-Diaz , Jakub M. Tomczak , Sandjai Bhulai

With the advancement of modern applications, an increasing number of composite optimization problems arise whose smooth component does not possess a globally Lipschitz continuous gradient. This setting prevents the direct use of the…

Optimization and Control · Mathematics 2026-05-11 Lei Yang , Jingjing Hu , Tianxiang Liu

Input gradients have a pivotal role in a variety of applications, including adversarial attack algorithms for evaluating model robustness, explainable AI techniques for generating Saliency Maps, and counterfactual explanations.However,…

Artificial Intelligence · Computer Science 2024-02-05 Mathieu Serrurier , Franck Mamalet , Thomas Fel , Louis Béthune , Thibaut Boissin

Residual neural networks are state-of-the-art deep learning models. Their continuous-depth analog, neural ordinary differential equations (ODEs), are also widely used. Despite their success, the link between the discrete and continuous…

Machine Learning · Statistics 2024-07-08 Pierre Marion , Yu-Han Wu , Michael E. Sander , Gérard Biau

Many types of neural network layers rely on matrix properties such as invertibility or orthogonality. Retaining such properties during optimization with gradient-based stochastic optimizers is a challenging task, which is usually addressed…

Machine Learning · Statistics 2020-12-02 Andreas Krämer , Jonas Köhler , Frank Noé

We propose a novel composite framework to find unknown fields in the context of inverse problems for partial differential equations (PDEs). We blend the high expressibility of deep neural networks as universal function estimators with the…

Numerical Analysis · Mathematics 2021-06-02 Samira Pakravan , Pouria A. Mistani , Miguel Angel Aragon-Calvo , Frederic Gibou

We propose a learnable variational model that learns the features and leverages complementary information from both image and measurement domains for image reconstruction. In particular, we introduce a learned alternating minimization…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Chi Ding , Qingchao Zhang , Ge Wang , Xiaojing Ye , Yunmei Chen

We consider a neural network architecture with randomized features, a sign-splitter, followed by rectified linear units (ReLU). We prove that our architecture exhibits robustness to the input perturbation: the output feature of the neural…

Machine Learning · Statistics 2018-03-14 Arun Venkitaraman , Alireza M. Javid , Saikat Chatterjee