English
Related papers

Related papers: Double Descent Risk and Volume Saturation Effects:…

200 papers

The classical bias-variance trade-off predicts that bias decreases and variance increase with model complexity, leading to a U-shaped risk curve. Recent work calls this into question for neural networks and other over-parameterized models,…

Machine Learning · Computer Science 2020-12-09 Zitong Yang , Yaodong Yu , Chong You , Jacob Steinhardt , Yi Ma

The dynamics of gradient-based training in neural networks often exhibit nontrivial structures; hence, understanding them remains a central challenge in theoretical machine learning. In particular, a concept of feature unlearning, in which…

Machine Learning · Computer Science 2026-02-10 Shota Imai , Sota Nishiyama , Masaaki Imaizumi

Double robustness (DR) is a widely-used property of estimators that provides protection against model misspecification and slow convergence of nuisance functions. Despite its widespread application, the theoretical foundation of DR remains…

Statistics Theory · Mathematics 2025-07-22 Andrew Ying

Over-parameterized models, such as large deep networks, often exhibit a double descent phenomenon, whereas a function of model size, error first decreases, increases, and decreases at last. This intriguing double descent behavior also…

Machine Learning · Computer Science 2020-09-22 Reinhard Heckel , Fatih Furkan Yilmaz

Crystal plasticity of sub-micron finite volumes is characterized by the flow of emergent dislocation defects, giving rise to size effects in mechanical properties and avalanche phenomena. In this chapter, we present a minimal model for…

Materials Science · Physics 2019-08-09 Stefanos Papanikolaou , Michail Tzimas

We investigate the test risk of continuous-time stochastic gradient flow dynamics in learning theory. Using a path integral formulation we provide, in the regime of a small learning rate, a general formula for computing the difference…

Machine Learning · Statistics 2025-03-05 Rodrigo Veiga , Anastasia Remizova , Nicolas Macris

Continual learning is motivated by the need to adapt to real-world dynamics in tasks and data distribution while mitigating catastrophic forgetting. Despite significant advances in continual learning techniques, the theoretical…

Methodology · Statistics 2025-08-22 Yihan Zhao , Wenqing Su , Ying Yang

Recent empirical and theoretical analyses of several commonly used prediction procedures reveal a peculiar risk behavior in high dimensions, referred to as double/multiple descent, in which the asymptotic risk is a non-monotonic function of…

Statistics Theory · Mathematics 2022-05-26 Pratik Patil , Arun Kumar Kuchibhotla , Yuting Wei , Alessandro Rinaldo

In this paper we consider the problem of learning variational models in the context of supervised learning via risk minimization. Our goal is to provide a deeper understanding of the two approaches of learning of variational models via…

Machine Learning · Statistics 2023-09-07 Christoph Brauer , Niklas Breustedt , Timo de Wolff , Dirk A. Lorenz

Stochastic gradient descent updates parameters with summation gradient computed from a random data batch. This summation will lead to unbalanced training process if the data we obtained is unbalanced. To address this issue, this paper takes…

Machine Learning · Computer Science 2019-05-22 Tao Yi , Xingxuan Wang

This work identifies the existence and cause of a type of posterior collapse that frequently occurs in the Bayesian deep learning practice. For a general linear latent variable model that includes linear variational autoencoders as a…

Machine Learning · Computer Science 2022-10-17 Zihao Wang , Liu Ziyin

Unrolled neural networks emerged recently as an effective model for learning inverse maps appearing in image restoration tasks. However, their generalization risk (i.e., test mean-squared-error) and its link to network design and train…

Machine Learning · Computer Science 2019-06-11 Morteza Mardani , Qingyun Sun , Vardan Papyan , Shreyas Vasanawala , John Pauly , David Donoho

In this paper we consider the isoperimetric problem with double density in an Euclidean space, that is, we study the minimisation of the perimeter among subsets of $\mathbb{R}^n$ with fixed volume, where volume and perimeter are relative to…

Analysis of PDEs · Mathematics 2018-11-08 Aldo Pratelli , Giorgio Saracco

The double-cone ascending an inclined V-rail is a common exhibit used for demonstrating concepts related to center-of-mass in introductory physics courses. While the conceptual explanation is well-known--the widening of the ramp allows the…

Classical Physics · Physics 2009-11-11 Sohang C. Gandhi , Costas J. Efthimiou

A major challenge in understanding the generalization of deep learning is to explain why (stochastic) gradient descent can exploit the network architecture to find solutions that have good generalization performance when using high capacity…

Machine Learning · Computer Science 2019-02-12 Yifan Wu , Barnabas Poczos , Aarti Singh

When multiple models are considered in regression problems, the model averaging method can be used to weigh and integrate the models. In the present study, we examined how the goodness-of-prediction of the estimator depends on the…

Statistics Theory · Mathematics 2023-08-21 Ryo Ando , Fumiyasu Komaki

Deep neural networks are known to exhibit a `double descent' behavior as the number of parameters increases. Recently, it has also been shown that an `epochwise double descent' effect exists in which the generalization error initially…

Machine Learning · Computer Science 2021-08-30 Cory Stephenson , Tyler Lee

Existing bounds on the generalization error of deep networks assume some form of smooth or bounded dependence on the input variable, falling short of investigating the mechanisms controlling such factors in practice. In this work, we…

Machine Learning · Computer Science 2025-07-24 Matteo Gamba , Hossein Azizpour , Mårten Björkman

Double descent refers to the phase transition that is exhibited by the generalization error of unregularized learning models when varying the ratio between the number of parameters and the number of training samples. The recent success of…

Machine Learning · Computer Science 2020-06-19 Michał Dereziński , Feynman Liang , Michael W. Mahoney

Logistic regression is commonly used for modeling dichotomous outcomes. In the classical setting, where the number of observations is much larger than the number of parameters, properties of the maximum likelihood estimator in logistic…

Machine Learning · Statistics 2019-11-14 Fariborz Salehi , Ehsan Abbasi , Babak Hassibi
‹ Prev 1 3 4 5 6 7 10 Next ›