English
Related papers

Related papers: A Simple Bias Reduction for Chatterjee's Correlati…

200 papers

Covariance regression analysis is an approach to linking the covariance of responses to a set of explanatory variables $X$, where $X$ can be a vector, matrix, or tensor. Most of the literature on this topic focuses on the "Fixed-$X$"…

Statistics Theory · Mathematics 2025-01-08 Tao Zou , Wei Lan , Runze Li , Chih-Ling Tsai

Many statistical applications require the quantification of joint dependence among more than two random vectors. In this work, we generalize the notion of distance covariance to quantify joint dependence among d >= 2 random vectors. We…

Methodology · Statistics 2018-06-18 Shubhadeep Chakraborty , Xianyang Zhang

We consider the problem of testing whether a correlation matrix of a multivariate normal population is the identity matrix. We focus on sparse classes of alternatives where only a few entries are nonzero and, in fact, positive. We derive a…

Statistics Theory · Mathematics 2015-04-15 Ery Arias-Castro , Sébastien Bubeck , Gábor Lugosi

Recent advances in molecular simulations allow the evaluation of previously unattainable observables, such as rate constants for protein folding. However, these calculations are usually computationally expensive and even significant…

Applications · Statistics 2019-03-27 Barmak Mostofian , Daniel M. Zuckerman

We propose a new statistical estimation framework for a large family of global sensitivity analysis indices. Our approach is based on rank statistics and uses an empirical correlation coefficient recently introduced by Chatterjee [9]. We…

Methodology · Statistics 2026-05-25 Fabrice Gamboa , Pierre Gremaud , Thierry Klein , Agnès Lagnoux

Understanding the correlation between two different scores for the same set of items is a common problem in information retrieval, and the most commonly used statistics that quantifies this correlation is Kendall's $\tau$. However, the…

Social and Information Networks · Computer Science 2014-11-03 Sebastiano Vigna

A reasonable confidence interval should have a confidence coefficient no less than the given nominal level and a small expected length to reliably and accurately estimate the parameter of interest, and the bootstrap interval is considered…

Statistics Theory · Mathematics 2024-02-15 Weizhen Wang , Chongxiu Yu , Zhongzhan Zhang

In variable selection, most existing screening methods focus on marginal effects and ignore dependence between covariates. To improve the performance of selection, we incorporate pairwise effects in covariates for screening and…

Methodology · Statistics 2019-02-12 Siliang Gong , Kai Zhang , Yufeng Liu

Real-life statistical samples are often plagued by selection bias, which complicates drawing conclusions about the general population. When learning causal relationships between the variables is of interest, the sample may be assumed to be…

Statistics Theory · Mathematics 2018-11-15 Angelos P. Armen , Robin J. Evans

Bootstrap is a widely used technique that allows estimating the properties of a given estimator, such as its bias and standard error. In this paper, we evaluate and compare five bootstrap-based methods for making confidence intervals: two…

We study the problem of distributed mean estimation and optimization under communication constraints. We propose a correlated quantization protocol whose leading term in the error guarantee depends on the mean deviation of data points…

Machine Learning · Computer Science 2022-07-12 Ananda Theertha Suresh , Ziteng Sun , Jae Hun Ro , Felix Yu

This paper studies parametric bootstrap methods for network data, with the goal of quantifying the uncertainty of network statistics of interest. While existing network resampling methods primarily focus on count statistics under…

Methodology · Statistics 2026-05-29 Zhixuan Shao , Can M. Le

Detecting the components common or correlated across multiple data sets is challenging due to a large number of possible correlation structures among the components. Even more challenging is to determine the precise structure of these…

Information Theory · Computer Science 2019-02-01 Tanuj Hasija , Christian Lameiro , Timothy Marrinan , Peter J. Schreier

Overparametrization often helps improve the generalization performance. This paper presents a dual view of overparametrization suggesting that downsampling may also help generalize. Focusing on the proportional regime $m\asymp n \asymp p$,…

Statistics Theory · Mathematics 2023-10-17 Xin Chen , Yicheng Zeng , Siyue Yang , Qiang Sun

Let X, X_1,X_2,... be a sequence of i.i.d. random variables with mean $\mu=E X$. Let ${v_1^{(n)},...,v_n^{(n)}}_{n=1}^\infty$ be vectors of non-negative random variables (weights), independent of the data sequence…

Statistics Theory · Mathematics 2013-05-28 Miklos Csorgo , Yuliya Martsynyuk , Masoud Nasari

We consider learning methods based on the regularization of a convex empirical risk by a squared Hilbertian norm, a setting that includes linear predictors and non-linear predictors through positive-definite kernels. In order to go beyond…

Machine Learning · Computer Science 2019-06-19 Ulysse Marteau-Ferey , Dmitrii Ostrovskii , Francis Bach , Alessandro Rudi

Bootstrapping can produce confidence levels for hypotheses about quadratic regression models - such as whether the U-shape is inverted, and the location of optima. The method has several advantages over conventional methods: it provides…

Methodology · Statistics 2012-07-09 Michael Wood

Covariate balancing is a popular technique for controlling confounding in observational studies. It finds weights for the treatment group which are close to uniform, but make the group's covariate means (approximately) equal to those of the…

Methodology · Statistics 2025-03-07 Shiva Kaul , Min-Gyu Kim

Cross-validation is a widely used technique for evaluating the performance of prediction models, ranging from simple binary classification to complex precision medicine strategies. It helps correct for optimism bias in error estimates,…

The total correlation(TC) is a crucial index to measure the correlation between marginal distribution in multidimensional random variables, and it is frequently applied as an inductive bias in representation learning. Previous research has…

Methodology · Statistics 2023-05-01 Zihao Chen