中文
相关论文

相关论文: Double Descent Risk and Volume Saturation Effects:…

200 篇论文

The yielding transition of amorphous materials is studied with a two-dimensional Hamiltonian model that allows both shear and volume deformations. The model is investigated as a function of the relative value of the bulk modulus $B$ with…

软凝聚态物质 · 物理学 2022-01-05 E. A. Jagla

Recent large language models have been trained on vast datasets, but also often on repeated data, either intentionally for the purpose of upweighting higher quality data, or unintentionally because data deduplication is not perfect and the…

Bi-log-concavity of probability measures is a univariate extension of the notion of log-concavity that has been recently proposed in a statistical literature. Among other things, it has the nice property from a modelisation perspective to…

概率论 · 数学 2019-03-20 Adrien Saumard

Double descent presents a counter-intuitive aspect within the machine learning domain, and researchers have observed its manifestation in various models and tasks. While some theoretical explanations have been proposed for this phenomenon…

机器学习 · 计算机科学 2024-05-14 Yufei Gu

We develop a probabilistic framework for large-scale dimension bounds in metric geometry, based on padded decompositions, randomized ball carving on net graphs, and the Lov\'asz Local Lemma. For metric measure spaces with volume doubling…

度量几何 · 数学 2026-05-18 Jing Yu , Xingyu Zhu

Many Bayesian inference problems involve high dimensional models for which only a subset of the model variables are of actual interest. All other variables are just nuisance parameters that one would ideally like to integrate out…

统计计算 · 统计学 2025-08-13 Fabián González , Víctor Elvira , Joaquín Miguez

Modern deep neural networks exhibit strong generalization even in highly overparameterized regimes. Significant progress has been made to understand this phenomenon in the context of supervised learning, but for unsupervised tasks such as…

机器学习 · 统计学 2025-06-02 Jonghyun Ham , Maximilian Fleissner , Debarghya Ghoshdastidar

Recent theoretical studies illustrated that kernel ridgeless regression can guarantee good generalization ability without an explicit regularization. In this paper, we investigate the statistical properties of ridgeless regression with…

机器学习 · 计算机科学 2023-08-30 Jian Li , Yong Liu , Yingying Zhang

The robustness of risk measures to changes in underlying loss distributions (distributional uncertainty) is of crucial importance in making well-informed decisions. In this paper, we quantify, for the class of distortion risk measures with…

风险管理 · 定量金融 2023-03-14 Carole Bernard , Silvana M. Pesenti , Steven Vanduffel

Motivated by a recent literature on the double-descent phenomenon in machine learning, we consider highly over-parameterized models in causal inference, including synthetic control with many control units. In such models, there may be so…

计量经济学 · 经济学 2023-10-16 Jann Spiess , Guido Imbens , Amar Venugopal

Statistical uncertainties complicate engineering design -- confounding regulated design approaches, and degrading the performance of reliability efforts. The simplest means to tackle this uncertainty is double loop simulation; a nested…

统计方法学 · 统计学 2018-11-02 Zachary del Rosario , Richard W. Fenrich , Gianluca Iaccarino

Grokking, the unusual phenomenon for algorithmic datasets where generalization happens long after overfitting the training data, has remained elusive. We aim to understand grokking by analyzing the loss landscapes of neural networks,…

机器学习 · 计算机科学 2023-03-24 Ziming Liu , Eric J. Michaud , Max Tegmark

The problem of reinforcement learning in an unknown and discrete Markov Decision Process (MDP) under the average-reward criterion is considered, when the learner interacts with the system in a single stream of observations, starting from an…

机器学习 · 统计学 2018-03-06 Mohammad Sadegh Talebi , Odalric-Ambrym Maillard

LLM training is resource-intensive. Quantized training improves computational and memory efficiency but introduces quantization noise, which can hinder convergence and degrade model accuracy. Stochastic Rounding (SR) has emerged as a…

机器学习 · 计算机科学 2025-11-04 Taowen Liu , Marta Andronic , Deniz Gündüz , George A. Constantinides

Conventional wisdom attributes the mysterious generalization abilities of overparameterized neural networks to gradient descent (and its variants). The recent volume hypothesis challenges this view: it posits that these generalization…

机器学习 · 计算机科学 2025-12-19 Yotam Alexander , Yonatan Slutzky , Yuval Ran-Milo , Nadav Cohen

Adversarially trained models exhibit a large generalization gap: they can interpolate the training set even for large perturbation radii, but at the cost of large test error on clean samples. To investigate this gap, we decompose the test…

机器学习 · 计算机科学 2021-06-15 Yaodong Yu , Zitong Yang , Edgar Dobriban , Jacob Steinhardt , Yi Ma

We examine gradient descent on unregularized logistic regression problems, with homogeneous linear predictors on linearly separable datasets. We show the predictor converges to the direction of the max-margin (hard margin SVM) solution. The…

机器学习 · 统计学 2024-10-29 Daniel Soudry , Elad Hoffer , Mor Shpigel Nacson , Suriya Gunasekar , Nathan Srebro

The presence and evolution of defects that appear in the manufacturing process play a vital role in the failure mechanisms of engineering materials. In particular, the collective behavior of dislocation dynamics at the mesoscale leads to…

材料科学 · 物理学 2022-05-13 Eduardo Augusto Barros de Moraes , Marta D'Elia , Mohsen Zayernouri

Scaling recommendation models into large recommendation models has become one of the most widely discussed topics. Recent efforts focus on components beyond the scaling embedding dimension, as it is believed that scaling embedding may lead…

信息检索 · 计算机科学 2025-10-28 Yicheng He , Zhou Kaiyu , Haoyue Bai , Fengbin Zhu , Yonghui Yang

When training the parameters of a linear dynamical model, the gradient descent algorithm is likely to fail to converge if the squared-error loss is used as the training loss function. Restricting the parameter space to a smaller subset and…

机器学习 · 计算机科学 2020-07-13 Kamil Nar , Yuan Xue , Andrew M. Dai