English
Related papers

Related papers: Minimum sample size for detection of Gutenberg-Ric…

200 papers

A family of maximum mean discrepancy (MMD) kernel two-sample tests is introduced. Members of the test family are called Block-tests or B-tests, since the test statistic is an average over MMDs computed on subsets of the samples. The choice…

Machine Learning · Computer Science 2014-02-11 Wojciech Zaremba , Arthur Gretton , Matthew Blaschko

In nonparametric statistical problems, we wish to find an estimator of an unknown function f. We can split its error into bias and variance terms; Smirnov, Bickel and Rosenblatt have shown that, for a histogram or kernel estimate, the…

Statistics Theory · Mathematics 2013-02-19 Adam D. Bull

This paper considers properties of an optimization based sampler for targeting the posterior distribution when the likelihood is intractable and auxiliary statistics are used to summarize information in the data. Our reverse sampler…

Methodology · Statistics 2015-12-02 Jean-Jacques Forneron , Serena Ng

False discovery rates (FDR) are typically estimated from a mixture of a null and an alternative distribution. Here, we study a complementary approach proposed by Rice and Spiegelhalter (2008) that uses as primary quantities the null model…

Methodology · Statistics 2011-08-03 Bernd Klaus , Korbinian Strimmer

We present a general M-estimation framework for inference on the wavelet variance. This framework generalizes the results on the scale-wise properties of the standard estimator and extends them to deliver the joint asymptotic properties of…

Methodology · Statistics 2016-07-21 Stéphane Guerrier , Roberto Molinari

We present the first efficient averaging sampler that achieves asymptotically optimal randomness complexity and near-optimal sample complexity. For any $\delta < \varepsilon$ and any constant $\alpha > 0$, our sampler uses $m + O(\log (1 /…

Computational Complexity · Computer Science 2025-08-18 Zhiyang Xun , David Zuckerman

Approximate Markov chain Monte Carlo (MCMC) offers the promise of more rapid sampling at the cost of more biased inference. Since standard MCMC diagnostics fail to detect these biases, researchers have developed computable Stein discrepancy…

Machine Learning · Statistics 2020-10-16 Jackson Gorham , Lester Mackey

The state-of-the-art methods for estimating high-dimensional covariance matrices all shrink the eigenvalues of the sample covariance matrix towards a data-insensitive shrinkage target. The underlying shrinkage transformation is either…

Machine Learning · Statistics 2025-11-25 Man-Chung Yue , Yves Rychener , Daniel Kuhn , Viet Anh Nguyen

Good robust estimators can be tuned to combine a high breakdown point and a specified asymptotic efficiency at a central model. This happens in regression with MM- and tau-estimators among others. However, the finite-sample efficiency of…

Statistics Theory · Mathematics 2013-11-21 Ricardo Maronna , Víctor Yohai

We propose a method to optimize the representation and distinguishability of samples from two probability distributions, by maximizing the estimated power of a statistical test based on the maximum mean discrepancy (MMD). This optimized MMD…

We propose a novel coupling inequality of the min-max type for two random matrices with finite absolute third moments, which generalizes the quantitative versions of the well-known inequalities by Gordon. Previous results have calculated…

Probability · Mathematics 2024-11-14 Zijun Chen , Yiming Chen , Chengfu Wei

Many methods for machine learning rely on approximate inference from intractable probability distributions. Variational inference approximates such distributions by tractable models that can be subsequently used for approximate inference.…

Machine Learning · Computer Science 2020-10-08 Oleg Arenz , Mingjun Zhong , Gerhard Neumann

Let $\theta$ be a finitely supported probability measure on $\mathrm{SL}(2,\mathbb{C})$, and suppose that the semigroup generated by $\mathcal{G}:=\mathrm{supp}(\theta)$ is strongly irreducible and proximal. Let $\mu$ denote the Furstenberg…

Dynamical Systems · Mathematics 2025-11-12 Ariel Rapaport , Haojie Ren

Estimating the underlying distribution from \textit{iid} samples is a classical and important problem in statistics. When the alphabet size is large compared to number of samples, a portion of the distribution is highly likely to be…

Statistics Theory · Mathematics 2023-05-30 Prafulla Chandra , Andrew Thangaraj

Learning a Gaussian mixture model (GMM) is a fundamental problem in machine learning, learning theory, and statistics. One notion of learning a GMM is proper learning: here, the goal is to find a mixture of $k$ Gaussians $\mathcal{M}$ that…

Data Structures and Algorithms · Computer Science 2015-06-04 Jerry Li , Ludwig Schmidt

Finite mixture models have long been used across a variety of fields in engineering and sciences. Recently there has been a great deal of interest in quantifying the convergence behavior of the \emph{mixing measure}, a fundamental object…

Statistics Theory · Mathematics 2025-09-05 Yun Wei , Sayan Mukherjee , XuanLong Nguyen

Score-based diffusion models, while achieving minimax optimality for sampling, are often hampered by slow sampling speeds due to the high computational burden of score function evaluations. Despite the recent remarkable empirical advances…

Machine Learning · Computer Science 2025-02-27 Gen Li , Changxiao Cai

The generalized linear models (GLM) have been widely used in practice to model non-Gaussian response variables. When the number of explanatory features is relatively large, scientific researchers are of interest to perform controlled…

Methodology · Statistics 2020-07-03 Chenguang Dai , Buyu Lin , Xin Xing , Jun S. Liu

A fundamental problem in statistics is estimating the shape matrix of an Elliptical distribution. This generalizes the familiar problem of Gaussian covariance estimation, for which the sample covariance achieves optimal estimation error.…

Statistics Theory · Mathematics 2025-10-16 Lap Chi Lau , Akshay Ramachandran

We have measured the dissimilarities among several printed characters of a single page in the Gutenberg 42-line bible and we prove statistically the existence of several different matrices from which the metal types where constructed. This…

Machine Learning · Statistics 2010-02-04 Aureli Alabert , Luz Ma. Rangel