English
Related papers

Related papers: Testing High Dimensional Covariance Matrices, with…

200 papers

Coherence is a widely used measure to assess linear relationships between time series. However, it fails to capture nonlinear dependencies. To overcome this limitation, this paper introduces the notion of residual spectral density as a…

Statistics Theory · Mathematics 2024-05-21 Yuichi Goto , Xuze Zhang , Benjamin Kedem , Shuo Chen

The complexity of high-dimensional datasets presents significant challenges for machine learning models, including overfitting, computational complexity, and difficulties in interpreting results. To address these challenges, it is essential…

Machine Learning · Computer Science 2023-08-01 Gaurav Srivastava , Mahesh Jangid

In this paper we consider the uniformity testing problem for high-dimensional discrete distributions (multinomials) under sparse alternatives. More precisely, we derive sharp detection thresholds for testing, based on $n$ samples, whether a…

Statistics Theory · Mathematics 2022-02-17 Bhaswar B. Bhattacharya , Rajarshi Mukherjee

We study the problem of testing $H_0: \xi^\top\beta=t_0$ in high-dimensional sparse linear regression with Gaussian random design and unknown design covariance. The loading vector $\xi$ is arbitrary, and the exact sparsity level $k$ is…

Statistics Theory · Mathematics 2026-05-21 Jie Xie , Dongming Huang

In this paper, we propose a flexible model for survival analysis using neural networks along with scalable optimization algorithms. One key technical challenge for directly applying maximum likelihood estimation (MLE) to censored data is…

Machine Learning · Statistics 2021-12-07 Weijing Tang , Jiaqi Ma , Qiaozhu Mei , Ji Zhu

Advancement in sequencing technology enables the study of association between complex disorders and rare variants with low minor allele frequencies. One of the major challenges in rare variant testing is lack of statistical power of…

Quantitative Methods · Quantitative Biology 2016-07-27 Rui Sun , Haoyi Weng , Inchi Hu , Junfeng Guo , William K. K. Wu , Benny Chung-Ying Zee , Maggie Haitian Wang

Weakly supervised person search aims to jointly detect and match persons with only bounding box annotations. Existing approaches typically focus on improving the features by exploring relations of persons. However, scale variation problem…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Benzhi Wang , Yang Yang , Jinlin Wu , Guo-jun Qi , Zhen Lei

Distance-based regression model, as a nonparametric multivariate method, has been widely used to detect the association between variations in a distance or dissimilarity matrix for outcomes and predictor variables of interest in genetic…

Statistics Theory · Mathematics 2022-03-14 Yuke Shi , Wei Zhang , Aiyi Liu , Qizhai Li

We propose a novel method to cluster gene networks. Based on a dissimilarity built using correlation structures, we consider networks that connect all the genes based on the strength of their dissimilarity. The large number of genes require…

Statistics Theory · Mathematics 2016-07-07 A-C Brunet , J-M Azais , J-M Loubes , J Amar , R Burcelin

It is of great interest to quantify the contributions of genetic variation to brain structure and function, which are usually measured by high-dimensional imaging data (e.g., magnetic resonance imaging). In addition to the variance, the…

Applications · Statistics 2020-05-05 Benjamin B. Risk , Hongtu Zhu

In recent years, a considerable amount of work has been devoted to generalizing linear discriminant analysis to overcome its incompetence for high-dimensional classification (Witten & Tibshirani 2011, Cai & Liu 2011, Mai et al. 2012, Fan et…

Methodology · Statistics 2015-01-15 Qing Mai , Hui Zou

This paper addresses hypothesis testing for the mean of matrix-valued data in high-dimensional settings. We investigate the minimum discrepancy test, originally proposed by Cragg (1997), which serves as a rank test for lower-dimensional…

Methodology · Statistics 2024-12-12 Shijie Cui , Danning Li , Runze Li , Lingzhou Xue

Next-generation sequencing technologies now constitute a method of choice to measure gene expression. Data to analyze are read counts, commonly modeled using Negative Binomial distributions. A relevant issue associated with this…

Methodology · Statistics 2014-11-10 Elisabetta Bonafede , Franck Picard , Stéphane Robin , Cinzia Viroli

We study the problem of testing whether the missing values of a potentially high-dimensional dataset are Missing Completely at Random (MCAR). We relax the problem of testing MCAR to the problem of testing the compatibility of a collection…

Statistics Theory · Mathematics 2024-12-13 Alberto Bordino , Thomas B. Berrett

Even though there is a plethora of research in Microarray gene expression data analysis, still, it poses challenges for researchers to effectively and efficiently analyze the large yet complex expression of genes. The feature (gene)…

Neural and Evolutionary Computing · Computer Science 2023-11-13 Mrutyunjaya Panda

Covariance regression offers an effective way to model the large covariance matrix with the auxiliary similarity matrices. In this work, we propose a sparse covariance regression (SCR) approach to handle the potentially high-dimensional…

Methodology · Statistics 2024-10-17 Yuan Gao , Zhiyuan Zhang , Zhanrui Cai , Xuening Zhu , Tao Zou , Hansheng Wang

Methods to effectively detect multi-locus genetic association are becoming increasingly relevant in the genetic dissection of complex trait in humans. Current approaches typically consider a limited number of hypotheses, most of which are…

Genomics · Quantitative Biology 2007-05-23 Zhong Li , Aris Floratos , David Wang , Andrea Califano

The matrix-variate normal distribution is a popular model for high-dimensional transposable data because it decomposes the dependence structure of the random matrix into the Kronecker product of two covariance matrices: one for each of the…

Methodology · Statistics 2014-11-11 Anestis Touloumis , John Marioni , Simon Tavaré

Sparse principal component analysis (sparse PCA) is a widely used technique for dimensionality reduction in multivariate analysis, addressing two key limitations of standard PCA. First, sparse PCA can be implemented in high-dimensional low…

Methodology · Statistics 2025-10-07 Jan O. Bauer

Bipolar Disorder (BD) is a complex disease. It is heterogeneous, both at the phenotypic and genetic level, although the extent and impact of this heterogeneity is not fully understood. In this paper, we leverage recent advances in…