English
Related papers

Related papers: Compressed spectral screening for large-scale diff…

200 papers

High-dimensional classification and feature selection tasks are ubiquitous with the recent advancement in data acquisition technology. In several application areas such as biology, genomics and proteomics, the data are often functional in…

Machine Learning · Statistics 2021-09-30 W Yu , S Wade , H D Bondell , L Azizi

As an alternative to the traditional sampling theory, compressed sensing allows acquiring much smaller amount of data, still estimating the spectra of frequency-sparse signals accurately. However, compressed sensing usually requires random…

Information Theory · Computer Science 2016-07-22 Shan Huang , Hong Sun , Haijian Zhang , Lei Yu

Directed networks are broadly used to represent asymmetric relationships among units. Co-clustering aims to cluster the senders and receivers of directed networks simultaneously. In particular, the well-known spectral clustering algorithm…

Machine Learning · Statistics 2022-04-12 Xiao Guo , Yixuan Qiu , Hai Zhang , Xiangyu Chang

In cancer research, the comparison of gene expression or DNA methylation networks inferred from healthy controls and patients can lead to the discovery of biological pathways associated to the disease. As a cancer progresses, its signalling…

Methodology · Statistics 2015-06-17 Da Ruan , Alastair Young , Giovanni Montana

Spectral clustering is one of the most popular unsupervised machine learning methods. Constructing similarity matrix is crucial to this type of method. In most existing works, the similarity matrix is computed once for all or is updated…

Machine Learning · Computer Science 2023-06-30 Yongyan Guo , Gang Wu

Compressed sensing (CS) provides an elegant framework for recovering sparse signals from compressed measurements. For example, CS can exploit the structure of natural images and recover an image from only a few random measurements. CS is…

Machine Learning · Computer Science 2019-05-21 Yan Wu , Mihaela Rosca , Timothy Lillicrap

Consider a linear regression model where the design matrix X has n rows and p columns. We assume (a) p is much large than n, (b) the coefficient vector beta is sparse in the sense that only a small fraction of its coordinates is nonzero,…

Statistics Theory · Mathematics 2014-06-16 Jiashun Jin , Cun-Hui Zhang , Qi Zhang

In this paper, we study the problem of learning multi-dimensional Gaussian Mixture Models (GMMs), with a specific focus on model order selection and efficient mixing distribution estimation. We first establish an information-theoretic lower…

Machine Learning · Statistics 2026-03-23 Xinyu Liu , Hai Zhang

Selection of covariates is crucial in the estimation of average treatment effects given observational data with high or even ultra-high dimensional pretreatment variables. Existing methods for this problem typically assume sparse linear…

Methodology · Statistics 2023-03-20 Juan Chen , Yingchun Zhou

Cancer is the second leading cause of death, with chemotherapy as one of the primary forms of treatment. As a result, researchers are turning to drug combination therapy to decrease drug resistance and increase efficacy. Current methods of…

Quantitative Methods · Quantitative Biology 2024-11-08 Zachary Schwehr

In The Cancer Genome Atlas (TCGA) data set, there are many interesting nonlinear dependencies between pairs of genes that reveal important relationships and subtypes of cancer. Such genomic data analysis requires a rapid, powerful and…

Applications · Statistics 2022-11-30 Siqi Xiang , Wan Zhang , Siyao Liu , Katherine A. Hoadley , Charles M. Perou , Kai Zhang , J. S. Marron

In high-throughput data, dynamic correlation between genes, i.e. changing correlation patterns under different biological conditions, can reveal important regulatory mechanisms. Given the complex nature of dynamic correlation, and the…

Applications · Statistics 2017-05-09 Tianwei Yu

In the genomic era, the identification of gene signatures associated with disease is of significant interest. Such signatures are often used to predict clinical outcomes in new patients and aid clinical decision-making. However, recent…

Methodology · Statistics 2019-03-27 Naim U. Rashid , Quefeng Li , Jen Jen Yeh , Joseph G. Ibrahim

Latent Gaussian copula models provide a powerful means to perform multi-view data integration since these models can seamlessly express dependencies between mixed variable types (binary, continuous, zero-inflated) via latent Gaussian…

Computation · Statistics 2022-04-22 Grace Yoon , Christian L. Müller , Irina Gaynanova

With the recent advent of high-throughput genotyping techniques, genetic data for genome-wide association studies (GWAS) have become increasingly available, which entails the development of efficient and effective statistical approaches.…

Applications · Statistics 2015-02-04 Jiahan Li , Wei Zhong , Runze Li , Rongling Wu

Discriminative Canonical Correlation Analysis (DCCA) is a powerful supervised feature extraction technique for two sets of multivariate data, which has wide applications in pattern recognition. DCCA consists of two parts: (i) mean-centering…

Quantum Physics · Physics 2022-06-14 Yong-Mei Li , Hai-Ling Liu , Shi-Jie Pan , Su-Juan Qin , Fei Gao , Qiao-Yan Wen

Motivation: Tumor classification using Imaging Mass Spectrometry (IMS) data has a high potential for future applications in pathology. Due to the complexity and size of the data, automated feature extraction and classification steps are…

Machine Learning · Statistics 2018-06-28 Jens Behrmann , Christian Etmann , Tobias Boskamp , Rita Casadonte , Jörg Kriegsmann , Peter Maass

In this paper we consider the problem of recovering a high dimensional data matrix from a set of incomplete and noisy linear measurements. We introduce a new model that can efficiently restrict the degrees of freedom of the problem and is…

Information Theory · Computer Science 2012-11-22 Mohammad Golbabaee , Pierre Vandergheynst

Feature selection poses a challenge in small-sample high-dimensional datasets, where the number of features exceeds the number of observations, as seen in microarray, gene expression, and medical datasets. There isn't a universally optimal…

Machine Learning · Computer Science 2024-07-23 Hossein Nematzadeh , Joseph Mani , Zahra Nematzadeh , Ebrahim Akbari , Radziah Mohamad

In microarray experiments, it is often of interest to identify genes which have a pre-specified gene expression profile with respect to time. Methods available in the literature are, however, typically not stringent enough in identifying…

Applications · Statistics 2009-01-18 J. Tuke , G. F. V. Glonek , P. J. Solomon