English
Related papers

Related papers: Testing High Dimensional Covariance Matrices, with…

200 papers

Recent benchmarks reveal that models for single-cell perturbation response are often outperformed by simply predicting the dataset mean. We trace this anomaly to a metric artifact: control-referenced deltas and unweighted error metrics…

This chapter describes gene expression analysis by Singular Value Decomposition (SVD), emphasizing initial characterization of the data. We describe SVD methods for visualization of gene expression data, representation of the data using a…

Biological Physics · Physics 2007-05-23 Michael E. Wall , Andreas Rechtsteiner , Luis M. Rocha

Linkage disequilibrium score regression (LDSC) has emerged as an essential tool for genetic and genomic analyses of complex traits, utilizing high-dimensional data derived from genome-wide association studies (GWAS). LDSC computes the…

Methodology · Statistics 2025-04-16 Fei Xue , Bingxin Zhao

We consider two-sample tests for high-dimensional data under two disjoint models: the strongly spiked eigenvalue (SSE) model and the non-SSE (NSSE) model. We provide a general test statistic as a function of a positive-semidefinite matrix.…

Statistics Theory · Mathematics 2016-11-28 Makoto Aoshima , Kazuyoshi Yata

High-dimensional vector autoregression with measurement error is frequently encountered in a large variety of scientific and business applications. In this article, we study statistical inference of the transition matrix under this model.…

Methodology · Statistics 2020-09-18 Xiang Lyu , Jian Kang , Lexin Li

Gene regulatory network inference is crucial for understanding the complex molecular interactions in various genetic and environmental conditions. The rapid development of single-cell RNA sequencing (scRNA-seq) technologies unprecedentedly…

Methodology · Statistics 2021-11-09 Feiyi Xiao , Junjie Tang , Huaying Fang , Ruibin Xi

Regularized linear discriminant analysis (RLDA) is a widely used tool for classification and dimensionality reduction, but its performance in high-dimensional scenarios is inconsistent. Existing theoretical analyses of RLDA often lack clear…

Machine Learning · Statistics 2025-07-23 Yonghan Zhang , Zhangni Pu , Lu Yan , Jiang Hu

Gene expression levels in a population vary extensively across tissues. Such heterogeneity is caused by genetic variability and environmental factors, and is expected to be linked to disease development. The abundance of experimental data…

Machine Learning · Statistics 2015-06-26 Zi Wang , Wei Yuan , Giovanni Montana

Diffusion models have shown promise in text generation, but often struggle with generating long, coherent, and contextually accurate text. Token-level diffusion doesn't model word-order dependencies explicitly and operates on short, fixed…

Computation and Language · Computer Science 2025-05-27 Xiaochen Zhu , Georgi Karadzhov , Chenxi Whitehouse , Andreas Vlachos

In the early days of gene expression data, researchers have focused on gene-level analysis, and particularly on finding differentially expressed genes. This usually involved making a simplifying assumption that genes are independent, which…

Applications · Statistics 2021-06-29 Haim Bar , Seojin Bang

In this paper, we study the asymptotic behavior of the extreme eigenvalues and eigenvectors of the high dimensional spiked sample covariance matrices, in the supercritical case when a reliable detection of spikes is possible. Especially, we…

Statistics Theory · Mathematics 2020-09-04 Zhigang Bao , Xiucai Ding , Jingming Wang , Ke Wang

We present a method for individual and integrative analysis of high dimension, low sample size data that capitalizes on the recurring theme in multivariate analysis of projecting higher dimensional data onto a few meaningful directions that…

Methodology · Statistics 2016-11-04 Sandra E. Safo , Jeongyoun Ahn , Yongho Jeon , Sungkyu Jung

High dimensional classification has been highlighted for last two decades and much research has been conducted in order to circumvent challenges encountered in high dimensions. While existing methods have focused mainly on developing…

Methodology · Statistics 2022-11-16 Seungchul Baek

Modern high-throughput biomedical devices routinely produce data on a large scale, and the analysis of high-dimensional datasets has become commonplace in biomedical studies. However, given thousands or tens of thousands of measured…

Methodology · Statistics 2022-02-28 Vladimir Vutov , Thorsten Dickhaus

The assumption of independent subvectors arises in many aspects of multivariate analysis. In most real-world applications, however, we lack prior knowledge about the number of subvectors and the specific variables within each subvector.…

Methodology · Statistics 2024-01-23 Jan O. Bauer

In genetic association studies, rare variants with extremely small allele frequency play a crucial role in complex traits, and the set-based testing methods that jointly assess the effects of groups of single nucleotide polymorphisms (SNPs)…

Methodology · Statistics 2020-03-13 Shonosuke Sugasawa , Hisashi Noma

Deep Learning (DL) is increasingly used in safety-critical applications, raising concerns about its reliability. DL suffers from a well-known problem of lacking robustness, especially when faced with adversarial perturbations known as…

Software Engineering · Computer Science 2023-09-06 Wei Huang , Xingyu Zhao , Alec Banks , Victoria Cox , Xiaowei Huang

Machine learning models have been successfully employed in the diagnosis of Schizophrenia disease. The impact of classification models and the feature selection techniques on the diagnosis of Schizophrenia have not been evaluated. Here, we…

Machine Learning · Computer Science 2022-07-05 M. Tanveer , Jatin Jangir , M. A. Ganaie , Iman Beheshti , M. Tabish , Nikunj Chhabra

In many applications, data can be heterogeneous in the sense of spanning latent groups with different underlying distributions. When predictive models are applied to such data the heterogeneity can affect both predictive performance and…

Machine Learning · Statistics 2022-05-04 Thomas Lartigue , Sach Mukherjee

The differential network (DN) analysis identifies changes in measures of association among genes under two or more experimental conditions. In this article, we introduce a Pseudo-value Regression Approach for Network Analysis (PRANA). This…

Methodology · Statistics 2023-03-27 Seungjun Ahn , Tyler Grimes , Somnath Datta