English
Related papers

Related papers: Finite size scaling of the bayesian perceptron

200 papers

We study the first-order transition in the model of a simple perceptron with continuous weights and large, bit finite value of the inputs. Making the analogy with the usual finite-size physical systems, we calculate the shift and the…

Statistical Mechanics · Physics 2007-05-23 Elka Korutcheva , N. Tonchev

Recent studies show that transformer-based architectures emulate gradient descent during a forward pass, contributing to in-context learning capabilities - an ability where the model adapts to new tasks based on a sequence of prompt…

Statistics Theory · Mathematics 2024-05-13 Karthik Duraisamy

We consider a commonly studied supervised classification of a synthetic dataset whose labels are generated by feeding a one-layer neural network with random iid inputs. We study the generalization performances of standard classifiers in the…

Machine Learning · Statistics 2021-02-18 Benjamin Aubin , Florent Krzakala , Yue M. Lu , Lenka Zdeborová

One of the most classical results in high-dimensional learning theory provides a closed-form expression for the generalisation error of binary classification with the single-layer teacher-student perceptron on i.i.d. Gaussian inputs. Both…

We analyze the connection between minimizers with good generalizing properties and high local entropy regions of a threshold-linear classifier in Gaussian mixtures with the mean squared error loss function. We show that there exist…

Machine Learning · Computer Science 2021-02-03 Carlo Baldassi , Enrico M. Malatesta , Matteo Negri , Riccardo Zecchina

Recent theoretical results show that gradient descent on deep neural networks under exponential loss functions locally maximizes classification margin, which is equivalent to minimizing the norm of the weight matrices under margin…

Machine Learning · Computer Science 2021-07-22 Andrzej Banburski , Fernanda De La Torre , Nishka Pant , Ishana Shastri , Tomaso Poggio

We obtain rates of contraction of posterior distributions in inverse problems defined by scales of smoothness classes. We derive abstract results for general priors, with contraction rates determined by Galerkin approximation. The rate…

Statistics Theory · Mathematics 2020-07-15 Shota Gugushvili , Aad van der Vaart , Dong Yan

We study the posterior distribution of the Bayesian multiple change-point regression problem when the number and the locations of the change-points are unknown. While it is relatively easy to apply the general theory to obtain the…

Statistics Theory · Mathematics 2008-08-21 Heng Lian

Bayesian inference is a widely used statistical method. The free energy and generalization loss, which are used to estimate the accuracy of Bayesian inference, are known to be small in singular models that do not have a unique optimal…

Statistics Theory · Mathematics 2020-12-16 Shuya Nagayasu , Sumio Watanabe

Finite size effects for the Ising Model coupled to two dimensional random surfaces are studied by exploiting the exact results from the 2-matrix models. The fixed area partition function is numerically calculated with arbitrary precision by…

High Energy Physics - Theory · Physics 2009-10-28 N. D. Hari Dass , B. E. Hanlon , T. Yukawa

We derive, in the classical framework of Bayesian sensitivity analysis, optimal lower and upper bounds on posterior values obtained from Bayesian models that exactly capture an arbitrarily large number of finite-dimensional marginals of the…

Statistics Theory · Mathematics 2016-05-20 Houman Owhadi , Clint Scovel , Tim Sullivan

In this paper, we study the learning rate of generalized Bayes estimators in a general setting where the hypothesis class can be uncountable and have an irregular shape, the loss function can have heavy tails, and the optimal hypothesis may…

Statistics Theory · Mathematics 2021-11-22 Lam Si Tung Ho , Binh T. Nguyen , Vu Dinh , Duy Nguyen

This paper revisits a fundamental problem in statistical inference from a non-asymptotic theoretical viewpoint $\unicode{x2013}$ the construction of confidence sets. We establish a finite-sample bound for the estimator, characterizing its…

Statistics Theory · Mathematics 2023-01-03 Lang Liu , Zaid Harchaoui

We calculate universal finite-size scaling functions for systems with an n-component order parameter and algebraically decaying interactions. Just as previously has been found for short-range interactions, this leads to a singular…

Statistical Mechanics · Physics 2009-10-31 Erik Luijten

Modern machine learning classifiers often exhibit vanishing classification error on the training set. They achieve this by learning nonlinear representations of the inputs that maps the data into linearly separable classes. Motivated by…

Statistics Theory · Mathematics 2023-03-23 Andrea Montanari , Feng Ruan , Youngtak Sohn , Jun Yan

According to the concept of typicality, an ensemble average can be accurately approximated by an expectation value with respect to a single pure state drawn at random from a high-dimensional Hilbert space. This random-vector approximation,…

Statistical Mechanics · Physics 2020-05-22 J. Schnack , J. Richter , T. Heitmann , J. Richter , R. Steinigeweg

Bayesian methods for low-rank matrix completion with noise have been shown to be very efficient computationally. While the behaviour of penalized minimization methods is well understood both from the theoretical and computational points of…

Statistics Theory · Mathematics 2015-04-08 The Tien Mai , Pierre Alquier

We consider Bayesian inference of signals with vector-valued entries. Extending concentration techniques from the mathematical physics of spin glasses, we show that the matrix-valued minimum mean-square error concentrates when the size of…

Information Theory · Computer Science 2019-07-17 Jean Barbier

Bayesian and frequentist criteria fundamentally differ, but often posterior and sampling distributions agree asymptotically (e.g., Gaussian with same covariance). For the corresponding single-draw experiment, we characterize the frequentist…

Statistics Theory · Mathematics 2024-07-04 David M. Kaplan , Longhao Zhuo

We study full Bayesian procedures for high-dimensional linear regression under sparsity constraints. The prior is a mixture of point masses at zero and continuous distributions. Under compatibility conditions on the design matrix, the…

Statistics Theory · Mathematics 2015-10-15 Ismaël Castillo , Johannes Schmidt-Hieber , Aad van der Vaart
‹ Prev 1 2 3 10 Next ›