English
Related papers

Related papers: Distance-based regression analysis for measuring a…

200 papers

Many statistical methodologies for high-dimensional data assume the population is normal. Although a few multivariate normality tests have been proposed, to the best of our knowledge, none of them can properly control the type I error when…

Methodology · Statistics 2021-05-04 Hao Chen , Yin Xia

Standard logistic regression analysis of case-control data has low power to detect gene-environment interactions, but until recently it was the only method that could be used on complex polygenic data for which parametric distributional…

Methodology · Statistics 2020-10-13 Tianying Wang , Alex Asher

This paper derives the rate of convergence and asymptotic distribution for a class of Kolmogorov-Smirnov style test statistics for conditional moment inequality models for parameters on the boundary of the identified set under general…

Applications · Statistics 2011-12-06 Timothy B. Armstrong

In genetic association studies, detecting phenotype-genotype association is a primary goal. We assume that the relationship between the data -phenotype, genetic markers and environmental covariates - can be modelled by a generalized linear…

Methodology · Statistics 2020-04-13 K. K. Halle , Ø. Bakke , S. Djurovic , A. Bye , E. Ryeng , U. Wisløff , O. A. Andreassen , M. Langaas

The problem of testing for the parametric form of the conditional variance is considered in a fully nonparametric regression model. A test statistic based on a weighted $L_2$-distance between the empirical characteristic functions of…

Methodology · Statistics 2018-07-24 Juan Carlos Pardo-Fernandez , M. Dolores Jimenez-Gamero

Generative artificial intelligence (AI) models in smart grids have advanced significantly in recent years due to their ability to generate large amounts of synthetic data, which would otherwise be difficult to obtain in the real world due…

Machine Learning · Computer Science 2025-10-27 Yuting Cai , Shaohuai Liu , Chao Tian , Le Xie

Many high-dimensional hypothesis tests aim to globally examine marginal or low-dimensional features of a high-dimensional joint distribution, such as testing of mean vectors, covariance matrices and regression coefficients. This paper…

Statistics Theory · Mathematics 2020-02-04 Yinqiu He , Gongjun Xu , Chong Wu , Wei Pan

Composite likelihood inference has gained much popularity thanks to its computational manageability and its theoretical properties. Unfortunately, performing composite likelihood ratio tests is inconvenient because of their awkward…

Computation · Statistics 2014-08-01 Manuela Cattelan , Nicola Sartori

Advancements in data collection have led to increasingly common repeated observations with complex structures in biomedical studies. Treating these observations as random objects, rather than summarizing features as vectors, avoids feature…

Methodology · Statistics 2025-03-04 Jingru Zhang , Shengjie Zhang , Christopher W Jones , Mathias Basner , Haochang Shou

Let $\mathbf{X} = (X_i)_{1\leq i \leq n}$ be an i.i.d. sample of square-integrable variables in $\mathbb{R}^d$, \GB{with common expectation $\mu$ and covariance matrix $\Sigma$, both unknown.} We consider the problem of testing if $\mu$ is…

Machine Learning · Computer Science 2021-10-11 Gilles Blanchard , Jean-Baptiste Fermanian

A non parametric method based on the empirical likelihood is proposed for detecting the change in the coefficients of high-dimensional linear model where the number of model variables may increase as the sample size increases. This amounts…

Statistics Theory · Mathematics 2015-06-22 Gabriela Ciuperca , Zahraa Salloum

Comparing counterfactual distributions can provide more nuanced and valuable measures for causal effects, going beyond typical summary statistics such as averages. In this work, we consider characterizing causal effects via distributional…

Machine Learning · Statistics 2024-11-05 Kwangho Kim , Jisu Kim , Edward H. Kennedy

We propose a family of near-metrics based on local graph diffusion to capture similarity for a wide class of data sets. These quasi-metametrics, as their names suggest, dispense with one or two standard axioms of metric spaces, specifically…

Machine Learning · Statistics 2017-10-18 Chu Wang , Iraj Saniee , William S. Kennedy , Chris A. White

Distance covariance is a popular measure of dependence between random variables. It has some robustness properties, but not all. We prove that the influence function of the usual distance covariance is bounded, but that its breakdown value…

Methodology · Statistics 2025-08-26 Sarah Leyder , Jakob Raymaekers , Peter J. Rousseeuw

The statistics and machine learning communities have recently seen a growing interest in classification-based approaches to two-sample testing. The outcome of a classification-based two-sample test remains a rejection decision, which is not…

Statistics Theory · Mathematics 2022-11-15 Loris Michel , Jeffrey Näf , Nicolai Meinshausen

This paper is concerned with the problem of comparing the population means of two groups of independent observations. An approximate randomization test procedure based on the test statistic of Chen and Qin (2010) is proposed. The asymptotic…

Statistics Theory · Mathematics 2022-08-23 Rui Wang , Wangli Xu

Data-dependent metrics are powerful tools for learning the underlying structure of high-dimensional data. This article develops and analyzes a data-dependent metric known as diffusion state distance (DSD), which compares points using a…

Machine Learning · Statistics 2020-03-10 Lenore Cowen , Kapil Devkota , Xiaozhe Hu , James M. Murphy , Kaiyi Wu

The study of mixture models constitutes a large domain of research in statistics. In the first part of this work, we present phi-divergences and the existing methods which produce robust estimators. We are more particularly interested in…

Methodology · Statistics 2016-11-28 Diaa Al Mohamad

Determining whether two sets of images belong to the same or different distributions or domains is a crucial task in modern medical image analysis and deep learning; for example, to evaluate the output quality of image generative models.…

Randomly censored survival data are frequently encountered in applied sciences including biomedical or reliability applications and clinical trial analyses. Testing the significance of statistical hypotheses is crucial in such analyses to…

Methodology · Statistics 2019-01-08 Abhik Ghosh , Ayanendranath Basu , Leandro Pardo