English
Related papers

Related papers: Approximating $f$-Divergences with Rank Statistics

200 papers

This paper concerns the construction of tests for universal hypothesis testing problems, in which the alternate hypothesis is poorly modeled and the observation space is large. The mismatched universal test is a feature-based technique for…

Information Theory · Computer Science 2016-04-18 Dayu Huang , Sean Meyn

The family of rank estimators, including Han's maximum rank correlation (Han, 1987) as a notable example, has been widely exploited in studying regression problems. For these estimators, although the linear index is introduced for…

Statistics Theory · Mathematics 2019-08-15 Yanqin Fan , Fang Han , Wei Li , Xiao-Hua Zhou

We present a local density estimator based on first order statistics. To estimate the density at a point, $x$, the original sample is divided into subsets and the average minimum sample distance to $x$ over all such subsets is used to…

Methodology · Statistics 2014-12-10 Vikram V. Garg , Luis Tenorio , Karen Willcox

Rank-based statistical metrics, such as the invariant statistical loss (ISL), have recently emerged as robust and practically effective tools for training implicit generative models. In this work, we introduce dual-ISL, a novel…

Machine Learning · Computer Science 2025-11-07 José Manuel de Frutos , Manuel A. Vázquez , Pablo M. Olmos , Joaquín Míguez

We discuss a graph-based approach for testing spatial point patterns. This approach falls under the category of data-random graphs, which have been introduced and used for statistical pattern recognition in recent years. Our goal is to test…

Methodology · Statistics 2008-02-06 E. Ceyhan , C. E. Priebe , D. J. Marchette

Rank-based inference methods are applied in various disciplines, typically when procedures relying on standard normal theory are not justifiable, for example when data are not symmetrically distributed, contain outliers, or responses are…

Statistics Theory · Mathematics 2018-02-16 Edgar Brunner , Frank Konietschke , Arne C. Bathke , Markus Pauly

This study introduces a novel model that effectively captures asymmetric structures in multivariate contingency tables with ordinal categories. Leveraging the principle of maximum entropy, our approach employs f-divergence to provide a…

Methodology · Statistics 2025-12-22 Hisaya Okahara , Kouji Tahata

High-dimensional penalized rank regression is a powerful tool for modeling high-dimensional data due to its robustness and estimation efficiency. However, the non-smoothness of the rank loss brings great challenges to the computation. To…

Methodology · Statistics 2025-02-20 Leheng Cai , Xu Guo , Heng Lian , Liping Zhu

We study the approximation of arbitrary distributions $P$ on $d$-dimensional space by distributions with log-concave density. Approximation means minimizing a Kullback--Leibler-type functional. We show that such an approximation exists if…

Statistics Theory · Mathematics 2011-10-17 Lutz Duembgen , Richard Samworth , Dominic Schuhmacher

arXiv:2206.10812v1 [stat.ME] proposes a useful algorithm, named generalized Diversity Subsampling (g-DS) algorithm, to select a subsample following some target probability distribution from a finite data set and demonstrates its…

Methodology · Statistics 2023-09-06 Boyang Shang

We propose a non-parametric anomaly detection algorithm for high dimensional data. We first rank scores derived from nearest neighbor graphs on $n$-point nominal training data. We then train limited complexity models to imitate these scores…

Machine Learning · Statistics 2016-01-25 Jonathan Root , Venkatesh Saligrama , Jing Qian

Kernel density estimation is a popular method for estimating unseen probability distributions. However, the convergence of these classical estimators to the true density slows down in high dimensions. Moreover, they do not define meaningful…

Statistics Theory · Mathematics 2025-05-30 Jack Kendrick

We derive high-probability finite-sample uniform rates of consistency for $k$-NN regression that are optimal up to logarithmic factors under mild assumptions. We moreover show that $k$-NN regression adapts to an unknown lower intrinsic…

Machine Learning · Statistics 2018-11-06 Heinrich Jiang

Let $A$ be an $n\times n$ random symmetric matrix with independent identically distributed subgaussian entries of unit variance. We prove the following large deviation inequality for the rank of $A$: for all $1\leq k\leq c\sqrt{n}$,…

Probability · Mathematics 2026-05-08 Yi Han

$f$-divergences, which quantify discrepancy between probability distributions, are ubiquitous in information theory, machine learning, and statistics. While there are numerous methods for estimating $f$-divergences from data, a limit…

Statistics Theory · Mathematics 2023-10-13 Sreejith Sreekumar , Ziv Goldfeld , Kengo Kato

We study a dimensionality reduction technique for finite mixtures of high-dimensional multivariate response regression models. Both the dimension of the response and the number of predictors are allowed to exceed the sample size. We…

Statistics Theory · Mathematics 2017-02-17 Emilie Devijver

We study a rank based univariate two-sample distribution-free test. The test statistic is the difference between the average of between-group rank distances and the average of within-group rank distances. This test statistic is closely…

Methodology · Statistics 2018-02-28 Jamye Curry , Xin Dang , Hailin Sang

The framework of optimal transport has been leveraged to extend the notion of rank to the multivariate setting while preserving desirable properties of the resulting goodness-of-fit (GoF) statistics. In particular, the rank energy (RE) and…

Machine Learning · Statistics 2022-11-29 Shoaib Bin Masud , Matthew Werenski , James M. Murphy , Shuchin Aeron

So-called linear rank statistics provide a means for distribution-free (even in finite samples), yet highly flexible, two-sample testing in the setting of univariate random variables. Their flexibility derives from a choice of weights that…

Methodology · Statistics 2023-10-03 Dan D. Erdmann-Pham

This work establishes computable bounds between f-divergences for probability measures within a generalized quasi-$\varepsilon_{(M,m)}$-neighborhood framework. We make the following key contributions. (1) a unified characterization of local…

Information Theory · Computer Science 2025-08-12 Xinchun Yu , Shuangqing Wei , Xiao-Ping Zhang