English
Related papers

Related papers: Incorporating Covariates into Integrated Factor An…

200 papers

This paper presents a new modeling strategy for joint unsupervised analysis of multiple high-throughput biological studies. As in Multi-study Factor Analysis, our goals are to identify both common factors shared across studies and…

Applications · Statistics 2018-06-27 Roberta De Vito , Ruggero Bellio , Lorenzo Trippa , Giovanni Parmigiani

Technological advances in genotyping have given rise to hypothesis-based association studies of increasing scope. As a result, the scientific hypotheses addressed by these studies have become more complex and more difficult to address using…

With the increasing availability of various sensor technologies, we now have access to large amounts of multi-block (also called multi-set, multi-relational, or multi-view) data that need to be jointly analyzed to explore their latent…

Computational Engineering, Finance, and Science · Computer Science 2015-09-01 Guoxu Zhou , Qibin Zhao , Yu Zhang , Tülay Adalı , Shengli Xie , Andrzej Cichocki

A full parametric and linear specification may be insufficient to capture complicated patterns in studies exploring complex features, such as those investigating age-related changes in brain functional abilities. Alternatively, a partially…

Methodology · Statistics 2024-02-07 Jia Liang , Shuo Chen , Peter Kochunov , L Elliot Hong , Chixiang Chen

We study semiparametric factor models in high-dimensional panels where the factor loadings consist of a nonparametric component explained by observed covariates and an idiosyncratic component capturing unobserved heterogeneity. A key…

Methodology · Statistics 2025-12-09 Sijie Zheng

Transformations of covariates are widely used in applied statistics to improve interpretability and to satisfy assumptions required for valid inference. More broadly, feature engineering encompasses a wider set of practices aimed at…

Methodology · Statistics 2026-03-30 Claudia Collarin , Matteo Fasiolo , Yannig Goude , Simon N. Wood

The scale of functional magnetic resonance image data is rapidly increasing as large multi-subject datasets are becoming widely available and high-resolution scanners are adopted. The inherent low-dimensionality of the information in this…

We introduce a novel Bayesian hybrid matrix factorisation model (HMF) for data integration, based on combining multiple matrix factorisation methods, that can be used for in- and out-of-matrix prediction of missing values. The model is very…

Machine Learning · Statistics 2017-04-18 Thomas Brouwer , Pietro Lió

In recent years, a comprehensive study of multi-view datasets (e.g., multi-omics and imaging scans) has been a focus and forefront in biomedical research. State-of-the-art biomedical technologies are enabling us to collect multi-view…

Machine Learning · Statistics 2020-04-30 Md Ashad Alam , Chuan Qiu , Hui Shen , Yu-Ping Wang , Hong-Wen Deng

Suppose one is interested in estimating causal effects in the presence of potentially unmeasured confounding with the aid of a valid instrumental variable. This paper investigates the problem of making inferences about the average treatment…

Methodology · Statistics 2020-12-15 BaoLuo Sun , Wang Miao

We consider shared response modeling, a multi-view learning problem where one wants to identify common components from multiple datasets or views. We introduce Shared Independent Component Analysis (ShICA) that models each view as a linear…

Machine Learning · Computer Science 2021-10-27 Hugo Richard , Pierre Ablin , Bertrand Thirion , Alexandre Gramfort , Aapo Hyvärinen

We consider integrative modeling of multiple gene networks and diverse genomic data, including protein-DNA binding, gene expression and DNA sequence data, to accurately identify the regulatory target genes of a transcription factor (TF).…

Applications · Statistics 2012-03-21 Peng Wei , Wei Pan

Modern mobile health (mHealth) assessment combines self-reported measures of participants' health experiences with passively collected health behavior data throughout the day. These data are collected across multiple measurement scales,…

Methodology · Statistics 2026-03-13 Debangan Dey , Rahul Ghosal , Kathleen Merikangas , Vadim Zipunnikov

Non-linear dimensionality reduction can be performed by \textit{manifold learning} approaches, such as Stochastic Neighbour Embedding (SNE), Locally Linear Embedding (LLE) and Isometric Feature Mapping (ISOMAP). These methods aim to produce…

Machine Learning · Statistics 2021-12-09 Theodoulos Rodosthenous , Vahid Shahrezaei , Marina Evangelou

The R package GFA provides a full pipeline for factor analysis of multiple data sources that are represented as matrices with co-occurring samples. It allows learning dependencies between subsets of the data sources, decomposed into latent…

Mathematical Software · Computer Science 2016-11-08 Eemeli Leppäaho , Muhammad Ammad-ud-din , Samuel Kaski

The rapid advancement of high-throughput sequencing and other assay technologies has resulted in the generation of large and complex multi-omics datasets, offering unprecedented opportunities for advancing precision medicine strategies.…

Quantitative Methods · Quantitative Biology 2025-01-30 Ana R. Baião , Zhaoxiang Cai , Rebecca C Poulos , Phillip J. Robinson , Roger R Reddel , Qing Zhong , Susana Vinga , Emanuel Gonçalves

Studies often estimate associations between an outcome and multiple variates. For example, studies of diagnostic test accuracy estimate sensitivity and specificity, and studies of predictive and prognostic factors typically estimate…

High-dimensional multimodal data arises in many scientific fields. The integration of multimodal data becomes challenging when there is no known correspondence between the samples and the features of different datasets. To tackle this…

Quantitative Methods · Quantitative Biology 2023-04-11 Kathryn Dover , Zixuan Cang , Anna Ma , Qing Nie , Roman Vershynin

Recent neuroimaging studies that focus on predicting brain disorders via modern machine learning approaches commonly include a single modality and rely on supervised over-parameterized models.However, a single modality provides only a…

Missingness is a common issue for neuroimaging data, and neglecting it in downstream statistical analysis can introduce bias and lead to misguided inferential conclusions. It is therefore crucial to conduct appropriate statistical methods…

Methodology · Statistics 2025-03-25 Tong Lu , Chixiang Chen , Hsin-Hsiung Huang , Peter Kochunov , Elliot Hong , Shuo Chen