English
Related papers

Related papers: Sufficient digits and density estimation: A Bayesi…

200 papers

We consider inference from non-random samples in data-rich settings where high-dimensional auxiliary information is available both in the sample and the target population, with survey inference being a special case. We propose a regularized…

Methodology · Statistics 2021-04-13 Yutao Liu , Andrew Gelman , Qixuan Chen

The Newcomb-Benford law, also known as the first-digit law, gives the probability distribution associated with the first digit of a dataset, so that, for example, the first significant digit has a probability of $30.1$ % of being $1$ and…

Popular Physics · Physics 2021-08-25 Andrea Burgos , Andrés Santos

The ratio of two densities provides a direct characterization of their differences. We consider the two-sample comparison problem by estimating this ratio given i.i.d. observations from two distributions. To this end, we propose additive…

Methodology · Statistics 2026-04-23 Naoki Awaya , Yuliang Xu , Li Ma

We consider the estimation of densities in multiple subpopulations, where the available sample size in each subpopulation greatly varies. This problem occurs in epidemiology, for example, where different diseases may share similar…

Methodology · Statistics 2021-09-15 Jiaming Qiu , Xiongtao Dai , Zhengyuan Zhu

A Bayesian nonparametric method for unimodal densities on the real line is provided by considering a class of species sampling mixture models containing random densities that are unimodal and not necessarily symmetric. This class of…

Statistics Theory · Mathematics 2007-06-13 Man-Wai Ho

We wish to estimate the total number of classes in a population based on sample counts, especially in the presence of high latent diversity. Drawing on probability theory that characterizes distributions on the integers by ratios of…

Methodology · Statistics 2014-12-10 A. Willis , J. Bunge

We derive a necessary and sufficient condition for the sum of M independent continuous random variables modulo 1 to converge to the uniform distribution in L^1([0,1]), and discuss generalizations to discrete random variables. A consequence…

Probability · Mathematics 2010-09-15 Steven J. Miller , Mark J. Nigrini

Robust statistical data modelling under potential model mis-specification often requires leaving the parametric world for the nonparametric. In the latter, parameters are infinite dimensional objects such as functions, probability…

Consider a density $f$ on $[0,1]$ that must be estimated from an i.i.d. sample $X_1,...,X_n$ drawn from $f$. In this note, we study binary-tree-based histogram estimates that use recursive splitting of intervals. If the decision to split an…

Statistics Theory · Mathematics 2025-04-24 Luc Devroye , Jad Hamdan

Data sets for statistical analysis become extremely large even with some difficulty of being stored on one single machine. Even when the data can be stored in one machine, the computational cost would still be intimidating. We propose a…

Methodology · Statistics 2020-02-18 Ya Su

We describe, analyze, and evaluate experimentally a new probabilistic model for word-sequence prediction in natural language based on prediction suffix trees (PSTs). By using efficient data structures, we extend the notion of PST to…

cmp-lg · Computer Science 2008-02-03 Fernando C. N. Pereira , Yoram Singer , Naftali Tishby

P\'olya trees are rooted, unlabeled trees on $n$ vertices. This paper gives an efficient, new way to generate P\'olya trees. This allows comparing typical unlabeled and labeled tree statistics and comparing asymptotic theorems with…

Combinatorics · Mathematics 2024-11-27 Laurent Bartholdi , Persi Diaconis

Our focus is on constructing a multiscale nonparametric prior for densities. The Bayes density estimation literature is dominated by single scale methods, with the exception of Polya trees, which favor overly-spiky densities even when the…

Methodology · Statistics 2014-10-06 Antonio Canale , David B. Dunson

We describe a new and computationally efficient Bayesian methodology for inferring species trees and demographics from unlinked binary markers. Likelihood calculations are carried out using diffusion models of allele frequency dynamics…

Populations and Evolution · Quantitative Biology 2019-09-18 Marnus Stoltz , Boris Bauemer , Remco Bouckaert , Colin Fox , Gordon Hiscott , David Bryant

Prime numbers seem to distribute among the natural numbers with no other law than that of chance, however its global distribution presents a quite remarkable smoothness. Such interplay between randomness and regularity has motivated sci-…

Number Theory · Mathematics 2008-11-21 Bartolo Luque , Lucas Lacasa

In Bayesian nonparametric inference, random discrete probability measures are commonly used as priors within hierarchical mixture models for density estimation and for inference on the clustering of the data. Recently, it has been shown…

Statistics Theory · Mathematics 2012-11-26 Stefano Favaro , Antonio Lijoi , Igor Prünster

Projected distributions have proved to be useful in the study of circular and directional data. Although any multivariate distribution can be used to produce a projected model, these distributions are typically parametric. In this article…

Methodology · Statistics 2023-10-12 Luis E. Nieto-Barajas

Flexibly modeling how an entire density changes with covariates is an important but challenging generalization of mean and quantile regression. While existing methods for density regression primarily consist of covariate-dependent discrete…

Methodology · Statistics 2021-12-24 Vittorio Orlandi , Jared Murray , Antonio Linero , Alexander Volfovsky

Big Data often presents as massive non-probability samples. Not only is the selection mechanism often unknown, but larger data volume amplifies the relative contribution of selection bias to total error. Existing bias adjustment approaches…

Methodology · Statistics 2022-03-29 Ali Rafei , Carol A. C. Flannagan , Brady T. West , Michael R. Elliott

Random forests is a common non-parametric regression technique which performs well for mixed-type unordered data and irrelevant features, while being robust to monotonic variable transformations. Standard random forests, however, do not…

Computation · Statistics 2019-06-19 Taylor Pospisil , Ann B. Lee