English
Related papers

Related papers: Efficient Calculation of P-value and Power for Qua…

200 papers

In this article, we consider the complete independence test of high-dimensional data. Based on Chatterjee coefficient, we pioneer the development of quadratic test and extreme value test which possess good testing performance for…

Statistics Theory · Mathematics 2024-09-17 Liqi Xia , Ruiyuan Cao , Jiang Du , Jun Dai

We study the fundamental problem of Principal Component Analysis in a statistical distributed setting in which each machine out of $m$ stores a sample of $n$ points sampled i.i.d. from a single unknown distribution. We study algorithms for…

Machine Learning · Computer Science 2017-02-28 Dan Garber , Ohad Shamir , Nathan Srebro

We study in detail the prediction for the semileptonic decays $\bar{B} \to D (D^*) \pi \ell \bar{\nu}$ by heavy quark and chiral symmetry. The branching ratio for $\bar{B} \to D \pi \ell \bar{\nu}$ is quite significant, as big as…

High Energy Physics - Phenomenology · Physics 2009-10-22 H-Y Cheng , C-Y Cheung , W. Dimm , G-L Lin , Y. C. Lin , T-M Yan , H-L Yu

For testing goodness of fit it is very popular to use either the chi square statistic or G statistics (information divergence). Asymptotically both are chi square distributed so an obvious question is which of the two statistics that has a…

Statistics Theory · Mathematics 2012-06-19 Peter Harremoës , Gábor Tusnády

Pearson's chi-squared test is widely used to assess the uniformity of discrete histograms, typically relying on a continuous chi-squared distribution to approximate the test statistic, since computing the exact distribution is…

Methodology · Statistics 2025-07-01 Nikola Banić , Neven Elezović

We introduce a powerful analytic method to study the statistics of the number $\mathcal{N}_{\textbf{A}}(\gamma)$ of eigenvalues inside any contour $\gamma \in \mathbb{C}$ for infinitely large non-Hermitian random matrices ${\textbf A}$. Our…

Disordered Systems and Neural Networks · Physics 2021-06-09 Antonio Tonatiúh Ramos Sánchez , Edgar Guzmán-González , Isaac Pérez Castillo , Fernando L. Metz

We study inference with a small labeled sample, a large unlabeled sample, and high-quality predictions from an external model. We link prediction-powered inference with empirical likelihood by stacking supervised estimating equations based…

Methodology · Statistics 2025-12-19 Guanghui Wang , Mengtao Wen , Changliang Zou

We have investigated a weighted chi-square distribution of the variable $\xi$ which is a weighted sum of squared normally distributed independent variables whose weights are cosines of angles $\phi_k=2\pi k/N$, where $k \in \{0,1,...,N-1\}$…

Disordered Systems and Neural Networks · Physics 2024-12-24 Vladislav Egorov , Boris Kryzhanovsky

In light of the rapidly growing large-scale data in federated ecosystems, the traditional principal component analysis (PCA) is often not applicable due to privacy protection considerations and large computational burden. Algorithms were…

Methodology · Statistics 2025-08-27 Shuting Shen , Junwei Lu , Xihong Lin

In this paper, we observe a fixed number of unknown $2\pi$-periodic functions differing from each other by both phases and amplitude. This semiparametric model appears in literature under the name "shape invariant model." While the common…

Statistics Theory · Mathematics 2010-10-06 Myriam Vimond

A (p-1)-variate integral representation is given for the cumulative distribution function of the general p-variate non-central gamma distribution with a non-centrality matrix of any admissible rank. The real part of products of well known…

Statistics Theory · Mathematics 2016-07-06 Thomas Royen

This paper introduces chi-square goodness-of-fit tests to check for conditional distribution model specification. The data is cross-classified according to the Rosenblatt transform of the dependent variable and the explanatory variables,…

Econometrics · Economics 2023-09-25 Miguel A. Delgado , Julius Vainora

Covariance matrix estimation is an important problem in multivariate data analysis, both from theoretical as well as applied points of view. Many simple and popular covariance matrix estimators are known to be severely affected by model…

Methodology · Statistics 2025-11-21 Soumya Chakraborty , Ayanendranath Basu , Abhik Ghosh

A new test statistic based on success runs of weighted deviations is introduced. Its use for observations sampled from independent normal distributions is worked out in detail. It supplements the classic $\chi^{2}$ test which ignores the…

Statistics Theory · Mathematics 2017-04-10 Frederik Beaujean , Allen Caldwell

In the context of supervised parametric models, we introduce the concept of e-values. An e-value is a scalar quantity that represents the proximity of the sampling distribution of parameter estimates in a model trained on a subset of…

Machine Learning · Statistics 2022-07-19 Subhabrata Majumdar , Snigdhansu Chatterjee

Most of the statistical tests currently used to detect differentially expressed genes are based on asymptotic results, and perform poorly for low expression tags. Another problem is the common use of a single canonical cutoff for the…

Genomics · Quantitative Biology 2008-08-04 Leonardo Varuzza , Arthur Gruber , Carlos A. de B. Pereira

We propose generalized additive partial linear models for complex data which allow one to capture nonlinear patterns of some covariates, in the presence of linear components. The proposed method improves estimation efficiency and increases…

Statistics Theory · Mathematics 2014-05-26 Li Wang , Lan Xue , Annie Qu , Hua Liang

Statistical depth, which measures the center-outward rank of a given sample with respect to its underlying distribution, has become a popular and powerful tool in nonparametric inference. In this paper, we investigate the use of statistical…

Methodology · Statistics 2025-11-25 Chifeng Shen , Yuejiao Fu , Michael Chen , Xiaoping Shi

The large-scale multiple testing inherent to high throughput biological data necessitates very high statistical stringency and thus true effects in data are difficult to detect unless they have high effect sizes. One solution to this…

Methodology · Statistics 2017-12-21 Mohamad S. Hasan

This paper investigates the utilization of maximum and average distance correlations for multivariate independence testing. We characterize their consistency properties in high-dimensional settings with respect to the number of marginally…

Machine Learning · Statistics 2025-06-11 Cencheng Shen , Yuexiao Dong