English
Related papers

Related papers: Direct estimation and inference of higher-level co…

200 papers

In this work we consider the problem of estimating a high-dimensional $p \times p$ covariance matrix $\Sigma$, given $n$ observations of confounded data with covariance $\Sigma + \Gamma \Gamma^T$, where $\Gamma$ is an unknown $p \times q$…

Methodology · Statistics 2019-12-03 Rajen D. Shah , Benjamin Frot , Gian-Andrea Thanei , Nicolai Meinshausen

We present a procedure for effective estimation of entropy and mutual information from small-sample data, and apply it to the problem of inferring high-dimensional gene association networks. Specifically, we develop a James-Stein-type…

Machine Learning · Statistics 2009-08-08 Jean Hausser , Korbinian Strimmer

Extracting biomedical relations from large corpora of scientific documents is a challenging natural language processing task. Existing approaches usually focus on identifying a relation either in a single sentence (mention-level) or across…

Computation and Language · Computer Science 2020-11-23 Harshil Shah , Julien Fauqueur

This article deals with the analysis of high dimensional data that come from multiple sources (experiments) and thus have different possibly correlated responses, but share the same set of predictors. The measurements of the predictors may…

Methodology · Statistics 2020-07-01 Guorong Dai , Ursula U. Müller , Raymond J. Carroll

We propose generalized additive partial linear models for complex data which allow one to capture nonlinear patterns of some covariates, in the presence of linear components. The proposed method improves estimation efficiency and increases…

Statistics Theory · Mathematics 2014-05-26 Li Wang , Lan Xue , Annie Qu , Hua Liang

We propose a principal components regression method based on maximizing a joint pseudo-likelihood for responses and predictors. Our method uses both responses and predictors to select linear combinations of the predictors relevant for the…

Methodology · Statistics 2021-08-10 Karl Oskar Ekvall

This study proposes a novel method for estimation and hypothesis testing in high-dimensional single-index models. We address a common scenario where the sample size and the dimension of regression coefficients are large and comparable.…

Statistics Theory · Mathematics 2024-04-30 Kazuma Sawaya , Yoshimasa Uematsu , Masaaki Imaizumi

Recent research has shown growing interest in modeling hypergraphs, which capture polyadic interactions among entities beyond traditional dyadic relations. However, most existing methodologies for hypergraphs face significant limitations,…

Methodology · Statistics 2025-11-04 Shihao Wu , Gongjun Xu , Ji Zhu

In genetic association studies, a single marker is often associated with multiple, correlated phenotypes (e.g., obesity and cardiovascular disease, or nicotine dependence and lung cancer). A pervasive question is then whether that marker…

The analysis of data arising from environmental health studies which collect a large number of measures of exposure can benefit from using latent variable models to summarize exposure information. However, difficulties with estimation of…

Applications · Statistics 2009-08-21 Brisa N. Sánchez , Esben Budtz-Jørgensen , Louise M. Ryan

Background: Selecting feature genes to predict phenotypes is one of the typical tasks in analyzing genomics data. Though many general-purpose algorithms were developed for prediction, dealing with highly correlated genes in the prediction…

Applications · Statistics 2022-04-11 Li Xing , Songwan Joun , Kurt Mackay , Mary Lesperance , Xuekui Zhang

High-dimensional feature vectors are likely to contain sets of measurements that are approximate replicates of one another. In complex applications, or automated data collection, these feature sets are not known a priori, and need to be…

Methodology · Statistics 2020-10-07 Xin Bing , Florentina Bunea , Marten Wegkamp

This article focuses on covariance estimation for multi-study data. Popular approaches employ factor-analytic terms with shared and study-specific loadings that decompose the variance into (i) a shared low-rank component, (ii)…

Methodology · Statistics 2026-01-26 Lorenzo Mauri , Niccolò Anceschi , David B. Dunson

Probabilistic inference over large data sets is a challenging data management problem since exact inference is generally #P-hard and is most often solved approximately with sampling-based methods today. This paper proposes an alternative…

Databases · Computer Science 2016-06-15 Wolfgang Gatterbauer , Dan Suciu

Factor analysis provides linear factors that describe relationships between individual variables of a data set. We extend this classical formulation into linear factors that describe relationships between groups of variables, where each…

Machine Learning · Statistics 2014-12-03 Arto Klami , Seppo Virtanen , Eemeli Leppäaho , Samuel Kaski

High-dimensional phenotypes hold promise for richer findings in association studies, but testing of several phenotype traits aggravates the grand challenge of association studies, that of multiple testing. Several methods have recently been…

Methodology · Statistics 2013-05-14 Pekka Marttinen , Jussi Gillberg , Aki Havulinna , Jukka Corander , Samuel Kaski

High-dimensional mixed data as a combination of both continuous and ordinal variables are widely seen in many research areas such as genomic studies and survey data analysis. Estimating the underlying correlation among mixed data is hence…

Methodology · Statistics 2018-09-18 Xiaoyun Quan , James G. Booth , Martin T. Wells

We present a method for estimating sparse high-dimensional inverse covariance and partial correlation matrices, which exploits the connection between the inverse covariance matrix and linear regression. The method is a two-stage estimation…

Machine Learning · Statistics 2025-05-13 Samuel Erickson , Tobias Rydén

Factor analysis is a classical data reduction technique that seeks a potentially lower number of unobserved variables that can account for the correlations among the observed variables. This paper presents an extension of the factor…

Methodology · Statistics 2013-12-04 Tsung-I Lin , Pal H. Wu , Geoffrey J. McLachlan , Sharon X. Lee

Ancestry-specific proteome-wide association studies (PWAS) based on genetically predicted protein expression can reveal complex disease etiology specific to certain ancestral groups. These studies require ancestry-specific models for…

Applications · Statistics 2024-04-26 Aaron J. Molstad , Yanwei Cai , Alexander P. Reiner , Charles Kooperberg , Wei Sun , Li Hsu