English
Related papers

Related papers: Scaling Laws are Redundancy Laws

200 papers

Deep learning has non-convex loss landscape and its optimization dynamics is hard to analyze or control. Nevertheless, the dynamics can be empirically convex-like across various tasks, models, optimizers, hyperparameters, etc. In this work,…

Machine Learning · Computer Science 2026-02-10 Zhiqi Bu , Shiyun Xu , Jialin Mao

We present scaling laws that dictate both local and global connectivity properties of bounded wireless networks. These laws are defined with respect to the key system parameters of per-node transmit power and the number of antennas…

Networking and Internet Architecture · Computer Science 2017-06-15 Justin P. Coon , Orestis Georgiou , Carl P. Dettmann

An analytic perturbation theory is suggested in order to find finite-size corrections to the scaling power laws. In the frame of this theory it is shown that the first order finite-size correction to the scaling power laws has following…

Chaotic Dynamics · Physics 2011-11-10 A. Bershadskii

Training large language models (LLMs) is computationally expensive, partly because the loss exhibits slow power-law convergence whose origin remains debatable. Through systematic analysis of toy models and empirical evaluation of LLMs, we…

Machine Learning · Computer Science 2026-02-04 Yizhou Liu , Ziming Liu , Cengiz Pehlevan , Jeff Gore

The ever-growing ecosystem of LLMs has posed a challenge in selecting the most appropriate pre-trained model to fine-tune amidst a sea of options. Given constrained resources, fine-tuning all models and making selections afterward is…

Machine Learning · Computer Science 2024-05-29 Haowei Lin , Baizhou Huang , Haotian Ye , Qinyu Chen , Zihao Wang , Sujian Li , Jianzhu Ma , Xiaojun Wan , James Zou , Yitao Liang

We consider the problem of kernel classification. While worst-case bounds on the decay rate of the prediction error with the number of samples are known for some classifiers, they often fail to accurately describe the learning curves of…

Machine Learning · Statistics 2023-09-07 Hugo Cui , Bruno Loureiro , Florent Krzakala , Lenka Zdeborová

Neural scaling laws are driving the machine learning community toward training ever-larger foundation models across domains, assuring high accuracy and transferable representations for extrapolative tasks. We test this promise in quantum…

Chemical Physics · Physics 2025-10-01 Siwoo Lee , Adji Bousso Dieng

Recent work has shown that, in generative modeling, cross-entropy loss improves smoothly with model size and training compute, following a power law plus constant scaling law. One challenge in extending these results to reinforcement…

Machine Learning · Computer Science 2023-02-21 Jacob Hilton , Jie Tang , John Schulman

The scaling of the time delay near a "bottleneck" of a generic saddle-node bifurcation is well-known to be given by an inverse square-root law. We extend the analysis to several non-generic cases for smooth vector fields. We proceed to…

Dynamical Systems · Mathematics 2012-01-31 Christian Kuehn

We investigate trends in the data-error scaling laws of machine learning (ML) models trained on discrete combinatorial spaces that are prone-to-mutation, such as proteins or organic small molecules. We trained and evaluated kernel ridge…

Chemical Physics · Physics 2025-10-10 Vanni Doffini , O. Anatole von Lilienfeld , Michael A. Nash

Scale independence is a ubiquitous feature of complex systems which implies a highly skewed distribution of resources with no characteristic scale. Research has long focused on why systems as varied as protein networks, evolution and stock…

Physics and Society · Physics 2016-02-08 Laurent Hébert-Dufresne , Antoine Allard , Jean-Gabriel Young , Louis J. Dubé

Power-law probability distributions are widely used to model extreme statistical events in complex systems, with applications to a vast array of natural phenomena ranging from earthquakes to stock market crashes to pandemics. We show that…

Quantum Physics · Physics 2026-04-08 Wai-Keong Mok

We provide large deviations estimates for the upper tail of the number of triangles in scale-free inhomogeneous random graphs where the degrees have power law tails with index $-\alpha, \alpha \in (1,2)$. We show that upper tail…

Probability · Mathematics 2024-03-25 Clara Stegehuis , Bert Zwart

How do neural language models acquire a language's structure when trained for next-token prediction? We address this question by deriving theoretical scaling laws for neural network performance on synthetic datasets generated by the Random…

Machine Learning · Computer Science 2025-05-13 Francesco Cagnetta , Alessandro Favero , Antonio Sclocchi , Matthieu Wyart

Scaling limits, such as infinite-width limits, serve as promising theoretical tools to study large-scale models. However, it is widely believed that existing infinite-width theory does not faithfully explain the behavior of practical…

Machine Learning · Computer Science 2025-10-28 Moritz Haas , Sebastian Bordt , Ulrike von Luxburg , Leena Chennuru Vankadara

Understanding the geometric properties of gradient descent dynamics is a key ingredient in deciphering the recent success of very large machine learning models. A striking observation is that trained over-parameterized models retain some…

Machine Learning · Computer Science 2024-07-11 Sibylle Marcotte , Rémi Gribonval , Gabriel Peyré

Deep ResNets are recognized for achieving state-of-the-art results in complex machine learning tasks. However, the remarkable performance of these architectures relies on a training procedure that needs to be carefully crafted to avoid…

Machine Learning · Computer Science 2025-03-04 Pierre Marion , Adeline Fermanian , Gérard Biau , Jean-Philippe Vert

We consider a general class of preferential attachment schemes evolving by a reinforcement rule with respect to certain sublinear weights. In these schemes, which grow a random network, the sequence of degree distributions is an object of…

Probability · Mathematics 2014-02-19 Jihyeok Choi , Sunder Sethuraman , Shankar C. Venkataramani

We develop a solvable model of neural scaling laws beyond the kernel limit. Theoretical analysis of this model shows how performance scales with model size, training time, and the total amount of available data. We identify three scaling…

Machine Learning · Statistics 2025-04-07 Blake Bordelon , Alexander Atanasov , Cengiz Pehlevan

Data duplication during pretraining can degrade generalization and lead to memorization, motivating aggressive deduplication pipelines. However, at web scale, it is unclear what constitutes a ``duplicate'': beyond surface-form matches,…

Machine Learning · Computer Science 2026-03-10 Joshua Kazdan , Noam Levi , Rylan Schaeffer , Jessica Chudnovsky , Abhay Puri , Bo He , Mehmet Donmez , Sanmi Koyejo , David Donoho
‹ Prev 1 8 9 10 Next ›