中文
相关论文

相关论文: Temperature is All You Need for Generalization in …

200 篇论文

Thermodynamic principles can be employed to design parameter update laws that address challenges such as the exploration vs. exploitation dilemma. In this paper, inspired by the Langevin equation, an update law is developed for a…

系统与控制 · 电气工程与系统科学 2025-08-22 Saiedeh Akbari , Omkar Sudhir Patil , Warren E. Dixon

Understanding generalization in reinforcement learning (RL) is a significant challenge, as many common assumptions of traditional supervised learning theory do not apply. We focus on the special class of reparameterizable RL problems, where…

机器学习 · 计算机科学 2019-05-31 Huan Wang , Stephan Zheng , Caiming Xiong , Richard Socher

We present a general approach, based on exponential inequalities, to derive bounds on the generalization error of randomized learning algorithms. Using this approach, we provide bounds on the average generalization error as well as bounds…

机器学习 · 计算机科学 2023-03-10 Fredrik Hellström , Giuseppe Durisi

We introduce a new theoretical framework to analyze deep learning optimization with connection to its generalization error. Existing frameworks such as mean field theory and neural tangent kernel theory for neural network optimization…

机器学习 · 计算机科学 2020-10-28 Taiji Suzuki

In real word applications, data generating process for training a machine learning model often differs from what the model encounters in the test stage. Understanding how and whether machine learning models generalize under such…

机器学习 · 统计学 2022-02-08 Abdulkadir Canatar , Blake Bordelon , Cengiz Pehlevan

In this work, the probability of an event under some joint distribution is bounded by measuring it with the product of the marginals instead (which is typically easier to analyze) together with a measure of the dependence between the two…

信息论 · 计算机科学 2020-10-22 Amedeo Roberto Esposito , Michael Gastpar , Ibrahim Issa

Continuous-time models provide important insights into the training dynamics of optimization algorithms in deep learning. In this work, we establish a non-asymptotic convergence analysis of stochastic gradient Langevin dynamics (SGLD),…

机器学习 · 计算机科学 2026-01-30 Noah Oberweis , Semih Cayci

Using the methodology of conditional-probability density functional theory, and several mild assumptions, we calculate the temperature-dependence of the Perdew-Burke-Ernzerhof (PBE) generalized gradient approximation (GGA). This…

化学物理 · 物理学 2023-08-17 John Kozlowski , Dennis Perchak , Kieron Burke

We establish in-expectation and tail bounds on the generalization error of representation learning type algorithms. The bounds are in terms of the relative entropy between the distribution of the representations extracted from the training…

机器学习 · 统计学 2025-03-21 Milad Sefidgaran , Abdellatif Zaidi , Piotr Krasnowski

We consider information-theoretic bounds on expected generalization error for statistical learning problems in a networked setting. In this setting, there are $K$ nodes, each with its own independent dataset, and the models from each node…

信息论 · 计算机科学 2024-01-17 L. P. Barnes , Alex Dytso , H. V. Poor

To assess generalization, machine learning scientists typically either (i) bound the generalization gap and then (after training) plug in the empirical risk to obtain a bound on the true risk; or (ii) validate empirically on holdout data.…

机器学习 · 计算机科学 2021-11-09 Saurabh Garg , Sivaraman Balakrishnan , J. Zico Kolter , Zachary C. Lipton

In learned image compression, probabilistic models play an essential role in characterizing the distribution of latent variables. The Gaussian model with mean and scale parameters has been widely used for its simplicity and effectiveness.…

图像与视频处理 · 电气工程与系统科学 2025-04-24 Haotian Zhang , Li Li , Dong Liu

The estimation of the generalization error of classifiers often relies on a validation set. Such a set is hardly available in few-shot learning scenarios, a highly disregarded shortcoming in the field. In these scenarios, it is common to…

We establish a sharp uniform-in-time error estimate for the Stochastic Gradient Langevin Dynamics (SGLD), which is a widely-used sampling algorithm. Under mild assumptions, we obtain a uniform-in-time $O(\eta^2)$ bound for the KL-divergence…

概率论 · 数学 2025-03-20 Lei Li , Yuliang Wang

We consider the problem of regression learning for deterministic design and independent random errors. We start by proving a sharp PAC-Bayesian type bound for the exponentially weighted aggregate (EWA) under the expected squared empirical…

应用统计 · 统计学 2012-06-27 Arnak Dalalyan , Alexandre B. Tsybakov

We derive a novel information-theoretic analysis of the generalization property of meta-learning algorithms. Concretely, our analysis proposes a generic understanding of both the conventional learning-to-learn framework and the modern…

机器学习 · 计算机科学 2021-12-13 Qi Chen , Changjian Shui , Mario Marchand

We investigate the parameter space of transformer models trained on protein sequence data using a statistical mechanics framework, sampling the loss landscape at varying temperatures by Langevin dynamics to characterize the low-loss…

无序系统与神经网络 · 物理学 2026-04-01 L. Ghiringhelli , A. Zambon , G. Tiana

It has been shown that the nonreversible overdamped Langevin dynamics enjoy better convergence properties in terms of spectral gap and asymptotic variance than the reversible one. In this article we propose a variance reduction method for…

概率论 · 数学 2017-01-23 Romain Poncet

Physical scenarios that require a relativistic treatment are ubiquitous in nature, ranging from cosmological objects to charge carriers in Dirac materials. Interestingly all of these situations have in common that the systems typically…

统计力学 · 物理学 2020-08-11 P. S. Pal , Sebastian Deffner

A machine learning (ML) system must learn not only to match the output of a target function on a training set, but also to generalize to novel situations in order to yield accurate predictions at deployment. In most practical applications,…

机器学习 · 计算机科学 2022-12-13 Clare Lyle