English
Related papers

Related papers: Representing smooth functions as compositions of n…

200 papers

This note shows that, for a fixed Lipschitz constant $L > 0$, one layer neural networks that are $L$-Lipschitz are dense in the set of all $L$-Lipschitz functions with respect to the uniform norm on bounded sets.

Machine Learning · Statistics 2020-09-30 Stephan Eckstein

Many economic parameters are identified by ``thin sets'' (submanifolds with Lebesgue measure zero) and hence difficult to recover from data in an ambient space. This paper provides a unified theory for estimation and inference of such…

Econometrics · Economics 2026-03-09 Xiaohong Chen , Wayne Yuan Gao

This paper investigates the ability of finite samples to identify two-layer irreducible shallow networks with various nonlinear activation functions, including rectified linear units (ReLU) and analytic functions such as the logistic…

Machine Learning · Computer Science 2025-03-18 Yu Xia , Zhiqiang Xu

In this paper, in a multivariate setting we derive near optimal rates of convergence in the minimax sense for estimating partial derivatives of the mean function for functional data observed under a fixed synchronous design over H\"older…

Statistics Theory · Mathematics 2025-08-25 Max Berger , Hajo Holzmann

One prominent method of evaluating machine learning model trustworthiness is the notion of calibration. In the binary outcome setting, a probabilistic predictor is calibrated if outcomes are realized according to a model's distributional…

Machine Learning · Computer Science 2026-05-25 Jessica Finocchiaro , Victor Ganson , Drona Khurana

We study the approximation of the spectrum of a second-order elliptic differential operator by the Hybrid High-Order (HHO) method. The HHO method is formulated using cell and face unknowns which are polynomials of some degree $k\geq0$. The…

Numerical Analysis · Mathematics 2018-07-23 Victor Calo , Matteo Cicuttin , Quanling Deng , Alexandre Ern

We study the oracle complexity of nonsmooth nonconvex optimization, with the algorithm assumed to have access only to local function information. It has been shown by Davis, Drusvyatskiy, and Jiang (2023) that for nonsmooth Lipschitz…

Optimization and Control · Mathematics 2024-09-17 Guy Kornowski , Swati Padmanabhan , Ohad Shamir

We prove that solution operators of elliptic obstacle-type variational inequalities (or, more generally, locally Lipschitz continuous functions possessing certain pointwise-a.e. convexity properties) are Newton differentiable when…

Optimization and Control · Mathematics 2023-06-09 Constantin Christof , Gerd Wachsmuth

We study the problem of efficiently computing the derivative of the fixed-point of a parametric nondifferentiable contraction map. This problem has wide applications in machine learning, including hyperparameter optimization, meta-learning…

Machine Learning · Statistics 2024-06-05 Riccardo Grazzi , Massimiliano Pontil , Saverio Salzo

The possibility of approximating a continuous function on a compact subset of the real line by a feedforward single hidden layer neural network with a sigmoidal activation function has been studied in many papers. Such networks can…

Neural and Evolutionary Computing · Computer Science 2016-06-29 Namig J. Guliyev , Vugar E. Ismailov

Structured convex optimization problems typically involve a mix of smooth and nonsmooth functions. The common practice is to activate the smooth functions via their gradient and the nonsmooth ones via their proximity operator. We show that,…

Optimization and Control · Mathematics 2019-09-11 Patrick L. Combettes , Lilian E. Glaudin

Feature attributions are a popular tool for explaining the behavior of Deep Neural Networks (DNNs), but have recently been shown to be vulnerable to attacks that produce divergent explanations for nearby inputs. This lack of robustness is…

Machine Learning · Computer Science 2020-10-23 Zifan Wang , Haofan Wang , Shakul Ramkumar , Matt Fredrikson , Piotr Mardziel , Anupam Datta

Convergence and convergence rate analyses of adaptive methods, such as Adaptive Moment Estimation (Adam) and its variants, have been widely studied for nonconvex optimization. The analyses are based on assumptions that the expected or…

Machine Learning · Computer Science 2022-06-28 Hideaki Iiduka

We consider the problem of numerically approximating the solutions to an elliptic partial differential equation (PDE) for which the boundary conditions are lacking. To alleviate this missing information, we assume to be given measurement…

Numerical Analysis · Mathematics 2024-06-07 Andrea Bonito , Diane Guignard

In decentralized optimization, several nodes connected by a network collaboratively minimize some objective function. For minimization of Lipschitz functions lower bounds are known along with optimal algorithms. We study a specific class of…

Optimization and Control · Mathematics 2023-03-15 Savelii Chezhegov , Alexander Rogozin , Alexander Gasnikov

In this work, we consider a nonsmooth minimisation problem in which the objective function can be represented as the maximum of finitely many smooth ``subfunctions''. First, we study a smooth min-max reformulation of the problem. Due to…

Optimization and Control · Mathematics 2024-04-17 Charl Ras , Matthew Tam , Daniel Uteda

Orthogonally invariant functions of symmetric matrices often inherit properties from their diagonal restrictions: von Neumann's theorem on matrix norms is an early example. We discuss the example of "identifiability", a common property of…

Optimization and Control · Mathematics 2013-04-15 Aris Daniilidis , Dmitriy Drusvyatskiy , Adrian S. Lewis

The Lipschitz constant of the map between the input and output space represented by a neural network is a natural metric for assessing the robustness of the model. We present a new method to constrain the Lipschitz constant of dense deep…

Machine Learning · Computer Science 2023-08-22 Ouail Kitouni , Niklas Nolte , Mike Williams

We study local filters for the Lipschitz property of real-valued functions $f: V \to [0,r]$, where the Lipschitz property is defined with respect to an arbitrary undirected graph $G=(V,E)$. We give nearly optimal local Lipschitz filters…

Data Structures and Algorithms · Computer Science 2024-05-06 Jane Lange , Ephraim Linder , Sofya Raskhodnikova , Arsen Vasilyan

We study the convergence of gradient flows related to learning deep linear neural networks (where the activation function is the identity map) from data. In this case, the composition of the network layers amounts to simply multiplying the…

Optimization and Control · Mathematics 2020-10-16 Bubacarr Bah , Holger Rauhut , Ulrich Terstiege , Michael Westdickenberg