English
Related papers

Related papers: Factor analysis in high dimensional biological dat…

200 papers

We introduce a new family of one factor distributions for high-dimensional binary data. The model provides an explicit probability for each event, thus avoiding the numeric approximations often made by existing methods. Model interpretation…

Methodology · Statistics 2015-11-05 Matthieu Marbac , Mohammed Sedki

Human mortality data sets can be expressed as multiway data arrays, the dimensions of which correspond to categories by which mortality rates are reported, such as age, sex, country and year. Regression models for such data typically assume…

Methodology · Statistics 2014-04-15 Bailey K. Fosdick , Peter D. Hoff

Anomaly detection aims to identify observations that deviate from the typical pattern of data. Anomalous observations may correspond to financial fraud, health risks, or incorrectly measured data in practice. We show detecting anomalies in…

Machine Learning · Statistics 2020-05-26 Matthew Davidow , David S. Matteson

Multi-dimensional functional data arises in numerous modern scientific experimental and observational studies. In this paper we focus on longitudinal functional data, a structured form of multidimensional functional data. Operating within a…

Methodology · Statistics 2019-09-20 John Shamshoian , Damla Senturk , Shafali Jeste , Donatello Telesca

Forensic scientists are often criticised for the lack of quantitative support for the conclusions of their examinations. While scholars advocate for the use of a Bayes factor to quantify the weight of forensic evidence, it is often…

Latent factor models are widely used to measure unobserved latent traits in social and behavioral sciences, including psychology, education, and marketing. When used in a confirmatory manner, design information is incorporated, yielding…

Methodology · Statistics 2019-06-14 Yunxiao Chen , Xiaoou Li , Siliang Zhang

Over the past decades, statisticians and machine-learning researchers have developed literally thousands of new tools for the reduction of high-dimensional data in order to identify the variables most responsible for a particular trait.…

Machine Learning · Statistics 2012-05-31 Chamont Wang , Jana Gevertz , Chaur-Chin Chen , Leonardo Auslender

We propose modeling raw functional data as a mixture of a smooth function and a high-dimensional factor component. The conventional approach to retrieving the smooth function from the raw data is through various smoothing techniques.…

Methodology · Statistics 2022-04-13 Yuan Gao , Han Lin Shang , Yanrong Yang

Factor models are widely applied to the analysis of multivariate data across disparate fields of research. However, modern scientific data are often incomplete, and estimating a factor model from partially observed data can be very…

Methodology · Statistics 2026-02-24 Giuseppe Vinci

The scale of functional magnetic resonance image data is rapidly increasing as large multi-subject datasets are becoming widely available and high-resolution scanners are adopted. The inherent low-dimensionality of the information in this…

Treatment effect estimation from observational data has attracted significant attention across various research fields. However, many widely used methods rely on the unconfoundedness assumption, which is often unrealistic due to the…

Machine Learning · Computer Science 2025-02-21 Di Fan , Renlei Jiang , Yunhao Wen , Chuanhou Gao

A high-dimensional $r$-factor model for an $n$-dimensional vector time series is characterised by the presence of a large eigengap (increasing with $n$) between the $r$-th and the $(r+1)$-th largest eigenvalues of the covariance matrix.…

Methodology · Statistics 2021-03-09 Matteo Barigozzi , Haeran Cho

Data-dependent metrics are powerful tools for learning the underlying structure of high-dimensional data. This article develops and analyzes a data-dependent metric known as diffusion state distance (DSD), which compares points using a…

Machine Learning · Statistics 2020-03-10 Lenore Cowen , Kapil Devkota , Xiaozhe Hu , James M. Murphy , Kaiyi Wu

In modern drug development, the broader availability of high-dimensional observational data provides opportunities for scientist to explore subgroup heterogeneity, especially when randomized clinical trials are unavailable due to cost and…

Methodology · Statistics 2021-02-24 Xinzhou Guo , Linqing Wei , Chong Wu , Jingshen Wang

A fundamental challenge in observational causal inference is that assumptions about unconfoundedness are not testable from data. Assessing sensitivity to such assumptions is therefore important in practice. Unfortunately, some existing…

Methodology · Statistics 2019-01-15 Alexander Franks , Alexander D'Amour , Avi Feller

Understanding covariate-varying interdependencies among features is of great interest in various applications. Motivated by microbiome studies where microbial abundances and interactions vary with environmental factors, we develop a…

Methodology · Statistics 2026-03-16 Shuangjie Zhang , Michael L. Patnode , Juhee Lee

Observational data is increasingly used as a means for making individual-level causal predictions and intervention recommendations. The foremost challenge of causal inference from observational data is hidden confounding, whose presence…

Machine Learning · Statistics 2018-10-30 Nathan Kallus , Aahlad Manas Puli , Uri Shalit

Imaging and hyperspectral data analysis is central to progress across biology, medicine, chemistry, and physics. The core challenge lies in converting high-resolution or high-dimensional datasets into interpretable representations that…

Image and Video Processing · Electrical Eng. & Systems 2025-12-29 Kamyar Barakati , Yu Liu , Utkarsh Pratiush , Boris N. Slautin , Sergei V. Kalinin

The foremost challenge to causal inference with real-world data is to handle the imbalance in the covariates with respect to different treatment options, caused by treatment selection bias. To address this issue, recent literature has…

Machine Learning · Statistics 2022-02-23 Zhixuan Chu , Stephen Rathbun , Sheng Li

This paper deals with the dimension reduction for high-dimensional time series based on common factors. In particular we allow the dimension of time series $p$ to be as large as, or even larger than, the sample size $n$. The estimation for…

Statistics Theory · Mathematics 2010-06-15 Clifford Lam , Qiwei Yao , Neil Bathia