English
Related papers

Related papers: Occam's Razor is Only as Sharp as Your ELBO

200 papers

We demonstrate that the principle of maximum relative entropy (ME), used judiciously, can ease the specification of priors in model selection problems. The resulting effect is that models that make sharp predictions are disfavoured,…

Data Analysis, Statistics and Probability · Physics 2009-12-07 Brendon J. Brewer , Matthew J. Francis

Extreme Learning Machines (ELMs) have become a popular tool in the field of Artificial Intelligence due to their very high training speed and generalization capabilities. Another advantage is that they have a single hyper-parameter that…

Machine Learning · Computer Science 2019-12-05 Nicolás Nieto , Francisco Ibarrola , Victoria Peterson , Hugo Rufiner , Ruben Spies

Many crucial problems in deep learning and statistical inference are caused by a variational gap, i.e., a difference between model evidence (log-likelihood) and evidence lower bound (ELBO). In particular, in a classical VAE setting that…

Machine Learning · Computer Science 2025-03-06 Łukasz Struski , Marcin Mazur , Paweł Batorski , Przemysław Spurek , Jacek Tabor

Interpretation of cosmological data to determine the number and values of parameters describing the universe must not rely solely on statistics but involve physical insight. When statistical techniques such as "model selection" or…

Astrophysics · Physics 2011-08-31 Eric V. Linder , Ramon Miquel

Solomonoff's general theory of inference and the Minimum Description Length principle formalize Occam's razor, and hold that a good model of data is a model that is good at losslessly compressing the data, including the cost of describing…

Machine Learning · Computer Science 2019-01-29 Léonard Blier , Yann Ollivier

We provide theoretical and empirical evidence that using tighter evidence lower bounds (ELBOs) can be detrimental to the process of learning an inference network by reducing the signal-to-noise ratio of the gradient estimator. Our results…

Machine Learning · Statistics 2019-03-07 Tom Rainforth , Adam R. Kosiorek , Tuan Anh Le , Chris J. Maddison , Maximilian Igl , Frank Wood , Yee Whye Teh

Bayesian model comparison implements Occam's razor through its sensitivity to the prior. However, prior-dependence makes it important to assess the influence of plausible alternative priors. Such prior sensitivity analyses for the Bayesian…

Methodology · Statistics 2026-01-22 Zixiao Hu , Jason D. McEwen

This paper is dedicated to a cautious learning methodology for predicting preferences between alternatives characterized by binary attributes (formally, each alternative is seen as a subset of attributes). By "cautious", we mean that the…

Artificial Intelligence · Computer Science 2022-06-16 Hugo Gilbert , Mohamed Ouaguenouni , Meltem Ozturk , Olivier Spanjaard

This paper studies the Variational Inference (VI) used for training Bayesian Neural Networks (BNN) in the overparameterized regime, i.e., when the number of neurons tends to infinity. More specifically, we consider overparameterized…

Machine Learning · Statistics 2022-07-11 Tom Huix , Szymon Majewski , Alain Durmus , Eric Moulines , Anna Korba

This paper derives an objective Bayesian "prior" based on considerations of entropy/information. By this means, it produces a quantitative measure of goodness of fit (the "H-statistic") that balances higher likelihood against the number of…

Astrophysics · Physics 2008-11-26 Rafael D. Sorkin

Neural networks' expressiveness comes at the cost of complex, black-box models that often extrapolate poorly beyond the domain of the training dataset, conflicting with the goal of finding compact analytic expressions to describe scientific…

Machine Learning · Computer Science 2023-11-29 Owen Dugan , Rumen Dangovski , Allan Costa , Samuel Kim , Pawan Goyal , Joseph Jacobson , Marin Soljačić

We investigate the evidence/flexibility (i.e., "Occam") paradigm and demonstrate the theoretical and empirical consistency of Bayesian evidence for the task of determining an appropriate generative model for network data. This model…

Methodology · Statistics 2024-05-09 Tianyu Wang , Zachary M. Pisano , Carey E. Priebe

Semi-supervised learning by self-training heavily relies on pseudo-label selection (PLS). The selection often depends on the initial model fit on labeled data. Early overfitting might thus be propagated to the final model by selecting…

Machine Learning · Statistics 2023-06-27 Julian Rodemann , Jann Goschenhofer , Emilio Dorigatti , Thomas Nagler , Thomas Augustin

This paper addresses the problem of model selection in the sequence model $Y=\theta+\varepsilon\xi$, when $\xi$ is sub-Gaussian, for non-euclidian loss-functions. In this model, the Penalized Comparison to Overfitting procedure is studied…

Statistics Theory · Mathematics 2025-04-16 Claire Lacour , Pascal Massart , Vincent Rivoirard

In applications of Bayesian procedures, once a class of priors has been chosen, it may be tempting to fix the prior's hyperparameters from the data, in an empirical Bayes (EB) fashion, usually by their maximum marginal likelihood estimates…

Statistics Theory · Mathematics 2026-04-14 Stefano Rizzelli , Judith Rousseau , Sonia Petrone

A recent line of work has shown that an overparametrized neural network can perfectly fit the training data, an otherwise often intractable nonconvex optimization problem. For (fully-connected) shallow networks, in the best case scenario,…

Machine Learning · Computer Science 2019-10-30 Armin Eftekhari , ChaeHwan Song , Volkan Cevher

Variational mean field approximations tend to struggle with contemporary overparametrized deep neural networks. Where a Bayesian treatment is usually associated with high-quality predictions and uncertainties, the practical reality has been…

This article applies the principle of Occam's Razor to non-parametric model building of statistical data, by finding a model with the minimal number of bits, leading to an exceptionally effective regularization method for probability…

Machine Learning · Statistics 2020-06-18 Peter Kövesarki

Decomposition of the evidence lower bound (ELBO) objective of VAE used for density estimation revealed the deficiency of VAE for representation learning and suggested ways to improve the model. In this paper, we investigate whether we can…

Machine Learning · Computer Science 2022-11-22 Fahim Faisal Niloy , M. Ashraful Amin , AKM Mahbubur Rahman , Amin Ahsan Ali

The importance of Variational Autoencoders reaches far beyond standalone generative models -- the approach is also used for learning latent representations and can be generalized to semi-supervised learning. This requires a thorough…

Machine Learning · Computer Science 2022-04-12 Alexander Shekhovtsov , Dmitrij Schlesinger , Boris Flach