中文
相关论文

相关论文: Approximate Heavy Tails in Offline (Multi-Pass) St…

200 篇论文

Stochastic gradient descent (SGD) is widely believed to perform implicit regularization when used to train deep neural networks, but the precise manner in which this occurs has thus far been elusive. We prove that SGD minimizes an average…

机器学习 · 计算机科学 2018-01-17 Pratik Chaudhari , Stefano Soatto

Recent theoretical and empirical successes in deep learning, including the celebrated neural scaling laws, are punctuated by the observation that many objects of interest tend to exhibit some form of heavy-tailed or power law behavior. In…

机器学习 · 统计学 2025-06-05 Liam Hodgkinson , Zhichao Wang , Michael W. Mahoney

This paper studies the high-dimensional scaling limits of online stochastic gradient descent (SGD). Building on the recent work of Ben Arous, Gheissari, and Jagannath on the effective dynamics of SGD, we study the critical scaling regime of…

机器学习 · 统计学 2026-05-01 Parsa Rangriz

Deep neural networks have revolutionized machine learning, yet their training dynamics remain theoretically unclear-we develop a continuous-time, matrix-valued stochastic differential equation (SDE) framework that rigorously connects the…

机器学习 · 计算机科学 2026-02-10 Brian Richard Olsen , Sam Fatehmanesh , Frank Xiao , Adarsh Kumarappan , Anirudh Gajula

Modelling excesses over a high threshold using the Pareto or generalized Pareto distribution (PD/GPD) is the most popular approach in extreme value statistics. This method typically requires high thresholds in order for the (G)PD to fit…

统计理论 · 数学 2009-01-13 Jan Beirlant , Elisabeth Joossens , Johan Segers

Neural scaling laws suggest that the test error of large language models trained online decreases polynomially as the model size and data size increase. However, such scaling can be unsustainable when running out of new data. In this work,…

机器学习 · 计算机科学 2025-09-26 Licong Lin , Jingfeng Wu , Peter L. Bartlett

Traditional implicit generative models are capable of learning highly complex data distributions. However, their training involves distinguishing real data from synthetically generated data using adversarial discriminators, which can lead…

机器学习 · 计算机科学 2025-09-05 José Manuel de Frutos , Manuel A. Vázquez , Pablo Olmos , Joaquín Míguez

Stochastic Gradient Descent (SGD) is an important algorithm in machine learning. With constant learning rates, it is a stochastic process that, after an initial phase of convergence, generates samples from a stationary distribution. We show…

机器学习 · 统计学 2017-09-12 Stephan Mandt , Matthew D. Hoffman , David M. Blei

Real-world data is extremely imbalanced and presents a long-tailed distribution, resulting in models that are biased towards classes with sufficient samples and perform poorly on rare classes. Recent methods propose to rebalance classes but…

计算机视觉与模式识别 · 计算机科学 2023-11-02 Weiqi Li , Fan Lyu , Fanhua Shang , Liang Wan , Wei Feng

Economically responsible mitigation of multivariate extreme risks-such as extreme rainfall over large areas, large simultaneous variations in many stock prices, or widespread breakdowns in transportation systems-requires assessing the…

机器学习 · 统计学 2026-01-13 Stéphane Lhaut , Holger Rootzén , Johan Segers

We study a heavily overloaded single-server queue with abandonment and derive bounds on stationary tail probabilities of the queue length. As the abandonment rate $\gamma \downarrow 0$, the centered-scaled queue length $\tilde{q}$ is known…

概率论 · 数学 2026-03-20 Zedong Wang , Siva Theja Maguluri

Stochastic Gradient Descent (SGD) is widely used in machine learning problems to efficiently perform empirical risk minimization, yet, in practice, SGD is known to stall before reaching the actual minimizer of the empirical risk. SGD…

机器学习 · 统计学 2017-02-09 Vivak Patel

This note presents an operational measure of fat-tailedness for univariate probability distributions, in $[0,1]$ where 0 is maximally thin-tailed (Gaussian) and 1 is maximally fat-tailed. Among others,1) it helps assess the sample size…

统计方法学 · 统计学 2019-04-30 Nassim Nicholas Taleb

Skewness and non-Gaussian behavior are essential features of the distribution of short-scale velocity increments in isotropic turbulent flows. Yet, although the skewness has been generally linked to time-reversal symmetry breaking and…

Distributionally-robust optimization is often studied for a fixed set of distributions rather than time-varying distributions that can drift significantly over time (which is, for instance, the case in finance and sociology due to…

最优化与控制 · 数学 2020-10-01 Iman Shames , Farhad Farokhi

Stochastic gradient descent (SGD) is a powerful optimization technique that is particularly useful in online learning scenarios. Its convergence analysis is relatively well understood under the assumption that the data samples are…

机器学习 · 计算机科学 2024-10-03 Ethan Che , Jing Dong , Xin T. Tong

Unraveling the reasons behind the remarkable success and exceptional generalization capabilities of deep neural networks presents a formidable challenge. Recent insights from random matrix theory, specifically those concerning the spectral…

机器学习 · 统计学 2023-04-10 Xuanzhe Xiao , Zeng Li , Chuanlong Xie , Fengwei Zhou

Preferential attachment is widely used to model power-law behavior of degree distributions in both directed and undirected networks. In a directed preferential attachment model, despite the well-known marginal power-law degree…

概率论 · 数学 2018-08-07 Tiandong Wang , Sidney I. Resnick

We consider the task of heavy-tailed statistical estimation given streaming $p$-dimensional samples. This could also be viewed as stochastic optimization under heavy-tailed distributions, with an additional $O(p)$ space complexity…

机器学习 · 计算机科学 2022-02-28 Che-Ping Tsai , Adarsh Prasad , Sivaraman Balakrishnan , Pradeep Ravikumar

We provide upper bounds on the end-to-end backlog and delay in a network with heavy-tailed and self-similar traffic. The analysis follows a network calculus approach where traffic is characterized by envelope functions and service is…

网络与互联网体系结构 · 计算机科学 2013-12-30 Jorg Liebeherr , Almut Burchard , Florin Ciucu