English
Related papers

Related papers: Optimal Nonlinearities Improve Generalization Perf…

200 papers

Deep learning models are often successfully trained using gradient descent, despite the worst case hardness of the underlying non-convex optimization problem. The key question is then under what conditions can one prove that optimization…

Machine Learning · Computer Science 2017-02-28 Alon Brutzkus , Amir Globerson

The linear functional strategy for the regularization of inverse problems is considered. For selecting the regularization parameter therein, we propose the heuristic quasi-optimality principle and some modifications including the smoothness…

Numerical Analysis · Mathematics 2018-05-23 Stefan Kindermann , Sergiy Pereverzyev , Andrey Pilipenko

Asymptotic equivalence in Le Cam's sense for nonparametric regression experiments is extended to the case of non-regular error densities, which have jump discontinuities at their endpoints. We prove asymptotic equivalence of such regression…

Statistics Theory · Mathematics 2011-01-28 Alexander Meister , Markus Reiß

We introduce a parametric nonlinear transformation that is well-suited for Gaussianizing data from natural images. The data are linearly transformed, and each component is then normalized by a pooled activity measure, computed by…

Machine Learning · Computer Science 2021-01-19 Johannes Ballé , Valero Laparra , Eero P. Simoncelli

Performance guarantees for compression in nonlinear models under non-Gaussian observations can be achieved through the use of distributional characteristics that are sensitive to the distance to normality, and which in particular return the…

Statistics Theory · Mathematics 2017-10-03 Larry Goldstein , Xiaohan Wei

Deep learning has been applied to various tasks in the field of machine learning and has shown superiority to other common procedures such as kernel methods. To provide a better theoretical understanding of the reasons for its success, we…

Machine Learning · Statistics 2023-05-31 Satoshi Hayakawa , Taiji Suzuki

Function approximation from input and output data pairs constitutes a fundamental problem in supervised learning. Deep neural networks are currently the most popular method for learning to mimic the input-output relationship of a general…

Machine Learning · Computer Science 2019-12-09 Nikos Kargas , Nicholas D. Sidiropoulos

In this work, we study the training and generalization performance of two-layer neural networks (NNs) after one gradient descent step under structured data modeled by Gaussian mixtures. While previous research has extensively analyzed this…

Machine Learning · Statistics 2025-05-20 Samet Demir , Zafer Dogan

The first purpose of this article is to obtain a.s. asymptotic properties of the maximum likelihood estimator in the autoregressive process driven by a stationary Gaussian noise. The second purpose is to show the local asymptotic normality…

Statistics Theory · Mathematics 2018-10-23 Marius Soltane

At the heart of machine learning lies the question of generalizability of learned rules over previously unseen data. While over-parameterized models based on neural networks are now ubiquitous in machine learning applications, our…

Machine Learning · Computer Science 2020-05-04 Melikasadat Emami , Mojtaba Sahraee-Ardakan , Parthe Pandit , Sundeep Rangan , Alyson K. Fletcher

Observations which are realizations from some continuous process are frequent in sciences, engineering, economics, and other fields. We consider linear models, with possible random effects, where the responses are random functions in a…

Statistics Theory · Mathematics 2016-11-30 Giacomo Aletti , Caterina May , Chiara Tommasi

We derive a novel deterministic equivalence for the two-point function of a random matrix resolvent. Using this result, we give a unified derivation of the performance of a wide variety of high-dimensional linear models trained with…

Disordered Systems and Neural Networks · Physics 2026-05-08 Alexander Atanasov , Blake Bordelon , Jacob A. Zavatone-Veth , Courtney Paquette , Cengiz Pehlevan

In this manuscript, we investigate the problem of how two-layer neural networks learn features from data, and improve over the kernel regime, after being trained with a single gradient descent step. Leveraging the insight from (Ba et al.,…

Machine Learning · Statistics 2024-09-06 Hugo Cui , Luca Pesce , Yatin Dandi , Florent Krzakala , Yue M. Lu , Lenka Zdeborová , Bruno Loureiro

Overparameterized ML models, including neural networks, typically induce underdetermined training objectives with multiple global minima. The implicit bias refers to the limiting global minimum that is attained by a common optimization…

Machine Learning · Statistics 2026-03-06 Kuo-Wei Lai , Guanghui Wang , Molei Tao , Vidya Muthukumar

The four-time correlation function of a general dynamical variable obeying Gaussian statistics is calculated for the trap model with a Gaussian density of states. It is argued that for energy-independent variables this function is…

Statistical Mechanics · Physics 2015-06-11 Gregor Diezemann

We present a general variational framework for the training of freeform nonlinearities in layered computational architectures subject to some slope constraints. The regularization that we add to the traditional training loss penalizes the…

Machine Learning · Statistics 2025-03-31 Michael Unser , Alexis Goujon , Stanislas Ducotterd

Though deep neural networks have achieved impressive success on various vision tasks, obvious performance degradation still exists when models are tested in out-of-distribution scenarios. In addressing this limitation, we ponder that the…

Computer Vision and Pattern Recognition · Computer Science 2023-01-18 Xiaotong Li , Zixuan Hu , Jun Liu , Yixiao Ge , Yongxing Dai , Ling-Yu Duan

Meta-learning has arisen as a successful method for improving training performance by training over many similar tasks, especially with deep neural networks (DNNs). However, the theoretical understanding of when and why overparameterized…

Machine Learning · Computer Science 2023-04-11 Peizhong Ju , Yingbin Liang , Ness B. Shroff

This note addresses the question of optimally estimating a linear functional of an object acquired through linear observations corrupted by random noise, where optimality pertains to a worst-case setting tied to a symmetric, convex, and…

Statistics Theory · Mathematics 2023-08-01 Simon Foucart , Grigoris Paouris

Periodic activation functions, often referred to as learned Fourier features have been widely demonstrated to improve sample efficiency and stability in a variety of deep RL algorithms. Potentially incompatible hypotheses have been made…

Machine Learning · Computer Science 2025-03-20 Augustine N. Mavor-Parker , Matthew J. Sargent , Caswell Barry , Lewis Griffin , Clare Lyle