English
Related papers

Related papers: Improved Methods for Making Inferences About Multi…

200 papers

Correlation matrix visualization is essential for understanding the relationships between variables in a dataset, but missing data can pose a significant challenge in estimating correlation coefficients. In this paper, we compare the…

Machine Learning · Computer Science 2023-09-06 Nhat-Hao Pham , Khanh-Linh Vo , Mai Anh Vu , Thu Nguyen , Michael A. Riegler , Pål Halvorsen , Binh T. Nguyen

Multivariate correlation analysis plays a key role in various fields such as statistics and big data analytics. In this paper, it is presented a new non-parametric association measure between more than two variables based on the concept of…

Methodology · Statistics 2024-02-13 Eloi Martinez-Rabert

The problem of identifying the most discriminating features when performing supervised learning has been extensively investigated. In particular, several methods for variable selection in model-based classification have been proposed.…

Applications · Statistics 2020-12-16 Andrea Cappozzo , Francesca Greselin , Thomas Brendan Murphy

We develop a new rank-based approach for univariate two-sample testing in the presence of missing data which makes no assumptions about the missingness mechanism. This approach is a theoretical extension of the Wilcoxon-Mann-Whitney test…

Methodology · Statistics 2024-03-25 Yijin Zeng , Niall M. Adams , Dean A. Bodenham

Inferring linear relationships lies at the heart of many empirical investigations. A measure of linear dependence should correctly evaluate the strength of the relationship as well as qualify whether it is meaningful for the population.…

Methodology · Statistics 2022-08-16 Kaustubh R. Patil , Simon B. Eickhoff , Robert Langner

In this paper, we address the problem of two-sample testing in the presence of missing data under a variety of missingness mechanisms. Our focus is on the well-known energy distance-based two-sample test. In addition to the standard…

Methodology · Statistics 2025-08-18 Danijel G. Aleksić , Bojana Milošević

We propose a method to distinguish causal influence from hidden confounding in the following scenario: given a target variable Y, potential causal drivers X, and a large number of background features, we propose a novel criterion for…

Machine Learning · Statistics 2022-02-07 You-Lin Chen , Lenon Minorics , Dominik Janzing

We consider the hypothesis testing problem of deciding whether an observed high-dimensional vector has independent normal components or, alternatively, if it has a small subset of correlated components. The correlated components may have a…

Statistics Theory · Mathematics 2012-06-04 Ery Arias-Castro , Sébastien Bubeck , Gábor Lugosi

In high dimensional analysis, effects of explanatory variables on responses sometimes rely on certain exposure variables, such as time or environmental factors. In this paper, to characterize the importance of each predictor, we utilize its…

Methodology · Statistics 2018-04-11 Yeqing Zhou , Jingyuan Liu , Zhihui Hao , Liping Zhu

For the purpose of explaining multivariate outlyingness, it is shown that the squared Mahalanobis distance of an observation can be decomposed into outlyingness contributions originating from single variables. The decomposition is obtained…

Methodology · Statistics 2025-03-17 Marcus Mayrhofer , Peter Filzmoser

Hidden variable graphical models can sometimes imply constraints on the observable distribution that are more complex than simple conditional independence relations. These observable constraints can falsify assumptions of the model that…

Methodology · Statistics 2026-05-12 Michael C. Sachs , Erin E. Gabriel , Robin J. Evans , Arvid Sjölander

Conformalized multiple testing offers a model-free way to control predictive uncertainty in decision-making. Existing methods typically use only part of the available data to build score functions tailored to specific settings. We propose a…

Methodology · Statistics 2026-05-22 Yuyang Huo , Xiaoyang Wu , Changliang Zou , Haojie Ren

Recent work has shown that deep learning models in NLP are highly sensitive to low-level correlations between simple features and specific output labels, leading to overfitting and lack of generalization. To mitigate this problem, a common…

Computation and Language · Computer Science 2022-04-28 Roy Schwartz , Gabriel Stanovsky

Multiple imputation has become one of the standard methods in drawing inferences in many incomplete data applications. Applications of multiple imputation in relatively more complex settings, such as high-dimensional clustered data, require…

Methodology · Statistics 2025-04-08 Qiushuang Li , Recai Yucel

Recently, there has been significant interest in linear regression in the situation where predictors and responses are not observed in matching pairs corresponding to the same statistical unit as a consequence of separate data collection…

Methodology · Statistics 2019-10-04 Martin Slawski , Guoqing Diao , Emanuel Ben-David

While data science is battling to extract information from the enormous explosion of data, many estimators and algorithms are being developed for better prediction. Researchers and data scientists often introduce new methods and evaluate…

Applications · Statistics 2019-05-22 Raju Rimal , Trygve Almøy , Solve Sæbø

The interplay between missing data and model uncertainty -- two classic statistical problems -- leads to primary questions that we formally address from an objective Bayesian perspective. For the general regression problem, we discuss the…

Confounding matters in almost all observational studies that focus on causality. In order to eliminate bias caused by connfounders, oftentimes a substantial number of features need to be collected in the analysis. In this case, large p…

Statistics Theory · Mathematics 2019-12-30 Shinyuu Lee , Yuru Zhu

Meta-analysis of genome-wide association studies is increasingly popular and many meta-analytic methods have been recently proposed. A majority of meta-analytic methods combine information from multiple studies by assuming that studies are…

Quantitative Methods · Quantitative Biology 2014-01-20 Buhm Han , Jae Hoon Sul , Eleazar Eskin , Paul I. W. de Bakker , Soumya Raychaudhuri

This paper provides partial identification of various binary choice models with misreported dependent variables. We propose two distinct approaches by exploiting different instrumental variables respectively. In the first approach, the…

Econometrics · Economics 2024-01-31 Orville Mondal , Rui Wang