English
Related papers

Related papers: Deficiency bounds for the multivariate inverse hyp…

200 papers

Imputation of missing values is a strategy for handling non-responses in surveys or data loss in measurement processes, which may be more effective than ignoring them. When the variable represents a count, the literature dealing with this…

Applications · Statistics 2020-07-31 Gilma Hernández-Herrera , Albert Navarro , David Moriña

The most fundamental problem in statistics is the inference of an unknown probability distribution from a finite number of samples. For a specific observed data set, answers to the following questions would be desirable: (1) Estimation:…

Statistics Theory · Mathematics 2013-01-23 Ali Kinkhabwala

Missing data is a ubiquitous problem. It is especially challenging in medical settings because many streams of measurements are collected at different - and often irregular - times. Accurate estimation of those missing measurements is…

Machine Learning · Computer Science 2017-11-27 Jinsung Yoon , William R. Zame , Mihaela van der Schaar

Random probabilities are a key component to many nonparametric methods in Statistics and Machine Learning. To quantify comparisons between different laws of random probabilities several works are starting to use the elegant Wasserstein over…

Statistics Theory · Mathematics 2024-05-27 Marta Catalano , Hugo Lavenant

General hypergeometric distribution (GHGD) definition: from a finite space $N$ containing $n$ elements, randomly select totally $T$ subsets $M_i$ (each contains $m_i$ elements, $1 \geq i \geq T$), what is the probability that exactly $x$…

Probability · Mathematics 2022-09-01 Xing-gang Mao , Xiao-yan Xue

Item nonresponse is frequently encountered in practice. Ignoring missing data can lose efficiency and lead to misleading inference. Fractional imputation is a frequentist approach of imputation for handling missing data. However, the…

Methodology · Statistics 2018-09-18 Hejian Sang , Jae Kwang Kim

We study the problem of mismatched binary hypothesis testing between i.i.d. distributions. We analyze the tradeoff between the pairwise error probability exponents when the actual distributions generating the observation are different from…

Information Theory · Computer Science 2022-04-28 Parham Boroumand , Albert Guillén i Fàbregas

Pseudo-marginal Metropolis-Hastings (pmMH) is a versatile algorithm for sampling from target distributions which are not easy to evaluate point-wise. However, pmMH requires good proposal distributions to sample efficiently from the target,…

Computation · Statistics 2018-07-30 Johan Dahlin , Adrian Wills , Brett Ninness

We study mixture of linear regression (random coefficient) models, which capture population heterogeneity by allowing the regression coefficients to follow an unknown distribution $G^*$. In contrast to common parametric methods that fix the…

Methodology · Statistics 2025-07-01 Hansheng Jiang , Adityanand Guntuboyina

When one applies a type of non-Hermitian effect, constant imaginary vector potential, to disordered systems, delocalization is induced even in two or lower dimension. By using the non-Hermitian induced transition as a probe, We propose a…

Disordered Systems and Neural Networks · Physics 2017-10-23 Tsunenao Kuwae , Nobuhiko Taniguchi

Random feature neural network approximations of the potential in Hamiltonian systems yield approximations of molecular dynamics correlation observables that have the expected error $\mathcal{O}\big((K^{-1}+J^{-1/2})^{\frac{1}{2}}\big)$, for…

Numerical Analysis · Mathematics 2024-06-24 Xin Huang , Petr Plechac , Mattias Sandberg , Anders Szepessy

Instance embeddings are an efficient and versatile image representation that facilitates applications like recognition, verification, retrieval, and clustering. Many metric learning methods represent the input as a single point in the…

Machine Learning · Computer Science 2019-08-28 Seong Joon Oh , Kevin Murphy , Jiyan Pan , Joseph Roth , Florian Schroff , Andrew Gallagher

Existing two-sample testing techniques, particularly those based on choosing a kernel for the Maximum Mean Discrepancy (MMD), often assume equal sample sizes from the two distributions. Applying these methods in practice can require…

Machine Learning · Statistics 2025-12-17 Aaron Wei , Milad Jalali , Danica J. Sutherland

We address the multiple testing problem under the assumption that the true/false hypotheses are driven by a Hidden Markov Model (HMM), which is recognized as a fundamental setting to model multiple testing under dependence since the seminal…

Methodology · Statistics 2021-05-04 Marie Perrot-Dockès , Gilles Blanchard , Pierre Neuvial , Etienne Roquain

This paper is an attempt to set a justification for making use of some dicrepancy indexes, starting from the classical Maximum Likelihood definition, and adapting the corresponding basic principle of inference to situations where…

Statistics Theory · Mathematics 2021-02-24 Michel Broniatowski

Test data measured by medical instruments often carry imprecise ranges that include the true values. The latter are not obtainable in virtually all cases. Most learning algorithms, however, carry out arithmetical calculations that are…

Machine Learning · Computer Science 2020-07-27 Mei Wang , Jianwen Su , Haiqin Lu

The distributed Hill estimator is a divide-and-conquer algorithm for estimating the extreme value index when data are stored in multiple machines. In applications, estimates based on the distributed Hill estimator can be sensitive to the…

Methodology · Statistics 2021-12-21 Liujun Chen , Deyuan Li , Chen Zhou

An information-theoretic upper bound on the generalization error of supervised learning algorithms is derived. The bound is constructed in terms of the mutual information between each individual training sample and the output of the…

Machine Learning · Computer Science 2020-08-06 Yuheng Bu , Shaofeng Zou , Venugopal V. Veeravalli

Mutual information is widely used, in a descriptive way, to measure the stochastic dependence of categorical random variables. In order to address questions such as the reliability of the descriptive value, one must consider…

Machine Learning · Computer Science 2007-07-13 Marcus Hutter , Marco Zaffalon

Deep hashing has shown promising performance in large-scale image retrieval. However, latent codes extracted by Deep Neural Networks (DNNs) will inevitably lose semantic information during the binarization process, which damages the…

Computer Vision and Pattern Recognition · Computer Science 2022-01-13 Chengyin Xu , Zenghao Chai , Zhengzhuo Xu , Hongjia Li , Qiruyi Zuo , Lingyu Yang , Chun Yuan