English
Related papers

Related papers: Data-driven priors and their posterior concentrati…

200 papers

We study full Bayesian procedures for high-dimensional linear regression. We adopt data-dependent empirical priors introduced in [1]. In their paper, these priors have nice posterior contraction properties and are easy to compute. Our paper…

Statistics Theory · Mathematics 2022-02-14 Xiao Fang , Malay Ghosh

What is the best way to exploit extra data -- be it unlabeled data from the same task, or labeled data from a related task -- to learn a given task? This paper formalizes the question using the theory of reference priors. Reference priors…

Machine Learning · Statistics 2022-06-17 Yansong Gao , Rahul Ramesh , Pratik Chaudhari

We study full Bayesian procedures for high-dimensional linear regression under sparsity constraints. The prior is a mixture of point masses at zero and continuous distributions. Under compatibility conditions on the design matrix, the…

Statistics Theory · Mathematics 2015-10-15 Ismaël Castillo , Johannes Schmidt-Hieber , Aad van der Vaart

In multi-parameter models, reference priors typically depend on the parameter or quantity of interest, and it is well known that this is necessary to produce objective posterior distributions with optimal properties. There are, however,…

Statistics Theory · Mathematics 2015-04-13 James O. Berger , Jose M. Bernardo , Dongchu Sun

Bayesian computational strategies for inference can be inefficient in approximating the posterior distribution in models that exhibit some form of periodicity. This is because the probability mass of the marginal posterior distribution of…

Machine Learning · Statistics 2025-12-01 Javier Lopez-Santiago , Luca Martino , Joaquin Miguez , Gonzalo Vazquez-Vilar

In this paper we adopt the familiar sparse, high-dimensional linear regression model and focus on the important but often overlooked task of prediction. In particular, we consider a new empirical Bayes framework that incorporates data in…

Statistics Theory · Mathematics 2020-07-28 Ryan Martin , Yiqi Tang

Foundation models, and in particular large language models, can generate highly informative responses, prompting growing interest in using these ''synthetic'' outputs as data in empirical research and decision-making. This paper introduces…

Artificial Intelligence · Computer Science 2025-12-02 Sanjog Misra

We study the behavior of the posterior distribution in high-dimensional Bayesian Gaussian linear regression models having $p\gg n$, with $p$ the number of predictors and $n$ the sample size. Our focus is on obtaining quantitative finite…

Statistics Theory · Mathematics 2014-01-06 Nate Strawn , Artin Armagan , Rayan Saab , Lawrence Carin , David Dunson

We establish concentration rates for estimation of treatment effects in experiments that incorporate prior sources of information -- such as past pilots, related studies, or expert assessments -- whose external validity is uncertain. Each…

Econometrics · Economics 2026-03-24 Frederico Finan , Demian Pouzo

A key sticking point of Bayesian analysis is the choice of prior distribution, and there is a vast literature on potential defaults including uniform priors, Jeffreys' priors, reference priors, maximum entropy priors, and weakly informative…

Methodology · Statistics 2017-11-22 Andrew Gelman , Daniel Simpson , Michael Betancourt

Modern applications routinely collect high-dimensional data, leading to statistical models having more parameters than there are samples available. A common solution is to impose sparsity in parameter estimation, often using penalized…

Methodology · Statistics 2025-07-08 Paolo Onorati , David B. Dunson , Antonio Canale

Posterior sampling allows exploitation of prior knowledge on the environment's transition dynamics to improve the sample efficiency of reinforcement learning. The prior is typically specified as a class of parametric distributions, the…

Machine Learning · Computer Science 2024-04-09 Mirco Mutti , Riccardo De Santi , Marcello Restelli , Alexander Marx , Giorgia Ramponi

Bayesian model comparison is often based on the posterior distribution over the set of compared models. This distribution is often observed to concentrate on a single model even when other measures of model fit or forecasting ability…

Statistics Theory · Mathematics 2020-03-10 Oscar Oelrich , Shutong Ding , Måns Magnusson , Aki Vehtari , Mattias Villani

When dealing with Bayesian inference the choice of the prior often remains a debatable question. Empirical Bayes methods offer a data-driven solution to this problem by estimating the prior itself from an ensemble of data. In the…

Methodology · Statistics 2020-05-13 Ilja Klebanov , Alexander Sikorski , Christof Schütte , Susanna Röblitz

How should one leverage historical data when past observations are not perfectly indicative of the future, e.g., due to the presence of unobserved confounders which one cannot "correct" for? Motivated by this question, we study a…

Machine Learning · Computer Science 2025-01-03 Omar Besbes , Will Ma , Omar Mouchtaki

The integration of data and knowledge from several sources is known as data fusion. When data is only available in a distributed fashion or when different sensors are used to infer a quantity of interest, data fusion becomes essential. In…

Machine Learning · Computer Science 2023-12-11 Peng Wu , Tales Imbiriba , Victor Elvira , Pau Closas

We present a new approach to semiparametric inference using corrected posterior distributions. The method allows us to leverage the adaptivity, regularization and predictive power of nonparametric Bayesian procedures to estimate…

Methodology · Statistics 2023-06-21 Andrew Yiu , Edwin Fong , Chris Holmes , Judith Rousseau

Estimation of parameters that obey specific constraints is crucial in statistics and machine learning; for example, when parameters are required to satisfy boundedness, monotonicity, or linear inequalities. Traditional approaches impose…

Methodology · Statistics 2026-04-03 Lachlan Astfalck , Deborshee Sen , Sayan Patra , Edward Cripps , David Dunson

Prior distributions for high-dimensional linear regression require specifying a joint distribution for the unobserved regression coefficients, which is inherently difficult. We instead propose a new class of shrinkage priors for linear…

Methodology · Statistics 2020-07-09 Yan Dora Zhang , Brian P. Naughton , Howard D. Bondell , Brian J. Reich

High dimensional statistics deals with the challenge of extracting structured information from complex model settings. Compared with the growing number of frequentist methodologies, there are rather few theoretically optimal Bayes methods…

Statistics Theory · Mathematics 2018-08-21 Chao Gao , Aad W. van der Vaart , Harrison H. Zhou