English
Related papers

Related papers: A Tracy-Widom Empirical Estimator For Valid P-valu…

200 papers

Let the sample correlation matrix be $W=YY^T$, where $Y=(y_{ij})_{p,n}$ with $y_{ij}=x_{ij}/\sqrt{\sum_{j=1}^nx_{ij}^2}$. We assume $\{x_{ij}: 1\leq i\leq p, 1\leq j\leq n\}$ to be a collection of independent symmetric distributed random…

Statistics Theory · Mathematics 2011-11-01 Zhigang Bao , Guangming Pan , Wang Zhou

Estimation of genewise variance arises from two important applications in microarray data analysis: selecting significantly differentially expressed genes and validation tests for normalization of microarray data. We approach the problem by…

Statistics Theory · Mathematics 2010-11-11 Jianqing Fan , Yang Feng , Yue S. Niu

Modern statistical analyses often encounter datasets with massive sizes and heavy-tailed distributions. For datasets with massive sizes, traditional estimation methods can hardly be used to estimate the extreme value index directly. To…

Methodology · Statistics 2022-07-26 Yongxin Li , Liujun Chen , Deyuan Li , Hansheng Wang

$\textbf{Motivation:}$ Small $p$-values are often required to be accurately estimated in large-scale genomic studies for the adjustment of multiple hypothesis tests and the ranking of genomic features based on their statistical…

Applications · Statistics 2023-08-29 Yang Shi , Mengqiao Wang , Weiping Shi , Ji-Hyun Lee , Huining Kang , Hui Jiang

In big data analysis, a simple task such as linear regression can become very challenging as the variable dimension $p$ grows. As a result, variable screening is inevitable in many scientific studies. In recent years, randomized algorithms…

Methodology · Statistics 2019-02-13 Yu-Hsiang Cheng , Tzee-Ming Huang , Su-Yun Huang

So-called linear rank statistics provide a means for distribution-free (even in finite samples), yet highly flexible, two-sample testing in the setting of univariate random variables. Their flexibility derives from a choice of weights that…

Methodology · Statistics 2023-10-03 Dan D. Erdmann-Pham

Consider semiparametric estimation where a doubly robust estimating function for a low-dimensional parameter is available, depending on two working models. With high-dimensional data, we develop regularized calibrated estimation as a…

Methodology · Statistics 2020-09-28 Satyajit Ghosh , Zhiqiang Tan

Distributionally robust optimization (DRO) has become a powerful framework for estimation under uncertainty, offering strong out-of-sample performance and principled regularization. In this paper, we propose a DRO-based method for linear…

Machine Learning · Statistics 2025-05-06 Liviu Aolaritei , Soroosh Shafiee , Florian Dörfler

Variable selection in high-dimensional scenarios is of great interested in statistics. One application involves identifying differentially expressed genes in genomic analysis. Existing methods for addressing this problem have some limits or…

Methodology · Statistics 2018-06-19 Liuhua Peng , Long Qu , Dan Nettleton

Class prediction is an important application of microarray gene expression data analysis. The high-dimensionality of microarray data, where number of genes (variables) is very large compared to the number of samples (obser- vations), makes…

Artificial Intelligence · Computer Science 2018-04-03 Andrej Kastrin , Borut Peterlin

G-computation has become a widely used robust method for estimating unconditional (marginal) treatment effects with covariate adjustment in the analysis of randomized clinical trials. Statistical inference in this context typically relies…

Methodology · Statistics 2025-03-18 Xin Zhang , Haitao Chu , Lin Liu , Satrajit Roychoudhury

The doubly-robust (DR) estimator is popular for evaluating causal effects in observational studies and is often perceived as more desirable than inverse probability weighting (IPW) or outcome modeling alone because it provides extra…

Methodology · Statistics 2026-02-03 Chengxin Yang , Laine E. Thomas , Fan Li

Large amount of multidimensional data represented by multiway arrays or tensors are prevalent in modern applications across various fields such as chemometrics, genomics, physics, psychology, and signal processing. The structural complexity…

Statistics Theory · Mathematics 2024-05-29 Arnab Auddy , Dong Xia , Ming Yuan

We establish two theorems for assessing the accuracy in total variation of multivariate discrete normal approximation to the distribution of an integer valued random vector $W$. The first is for sums of random vectors whose dependence…

Probability · Mathematics 2018-07-19 A. D. Barbour , A. Xia

Predictive analytics is increasingly used to guide decision-making in many applications. However, in practice, we often have limited data on the true predictive task of interest, and must instead rely on more abundant data on a…

Machine Learning · Statistics 2020-05-07 Hamsa Bastani

High-dimensional group inference is an essential part of statistical methods for analysing complex data sets, including hierarchical testing, tests of interaction, detection of heterogeneous treatment effects and inference for local…

Methodology · Statistics 2020-12-01 Zijian Guo , Claude Renaux , Peter Bühlmann , T. Tony Cai

In this article, we introduce a two-way factor model for a high-dimensional data matrix and study the properties of the maximum likelihood estimation (MLE). The proposed model assumes separable effects of row and column attributes and…

Methodology · Statistics 2021-03-17 Gao Zhigen , Yuan Chaofeng , Jing Bingyi , Huang Wei , Guo Jianhua

While generalized linear mixed models are a fundamental tool in applied statistics, many specifications, such as those involving categorical factors with many levels or interaction terms, can be computationally challenging to estimate due…

Methodology · Statistics 2024-12-03 Max Goplerud , Omiros Papaspiliopoulos , Giacomo Zanella

With origins in game theory, probabilistic values like Shapley values, Banzhaf values, and semi-values have emerged as a central tool in explainable AI. They are used for feature attribution, data attribution, data valuation, and more.…

Machine Learning · Computer Science 2026-01-14 R. Teal Witter , Yurong Liu , Christopher Musco

We consider the problem of estimating the counterfactual joint distribution of multiple quantities of interests (e.g., outcomes) in a multivariate causal model extended from the classical difference-in-difference design. Existing methods…

Machine Learning · Statistics 2023-11-03 Thong Pham , Shohei Shimizu , Hideitsu Hino , Tam Le