English
Related papers

Related papers: Factor analysis in high dimensional biological dat…

200 papers

In this paper, we set up the theoretical foundations for a high-dimensional functional factor model approach in the analysis of large cross-sections (panels) of functional time series (FTS). We first establish a representation result…

Statistics Theory · Mathematics 2021-04-14 Shahin Tavakoli , Gilles Nisol , Marc Hallin

Customer Satisfaction is the most important factors in the industry irrespective of domain. Key Driver Analysis is a common practice in data science to help the business to evaluate the same. Understanding key features, which influence the…

Machine Learning · Statistics 2018-05-29 Kumarjit Pathak , Jitin Kapila , Aasheesh Barvey

Compositional data represent a specific family of multivariate data, where the information of interest is contained in the ratios between parts rather than in absolute values of single parts. The analysis of such specific data is…

Estimations and applications of factor models often rely on the crucial condition that the number of latent factors is consistently estimated, which in turn also requires that factors be relatively strong, data are stationary and weak…

Statistics Theory · Mathematics 2020-06-05 Jianqing Fan , Yuan Liao

Large volume of Genomics data is produced on daily basis due to the advancement in sequencing technology. This data is of no value if it is not properly analysed. Different kinds of analytics are required to extract useful information from…

Other Quantitative Biology · Quantitative Biology 2017-07-25 M. Usman Ali , Shahzad Ahmed , Javed Ferzund , Atif Mehmood , Abbas Rehman

Factor analysis is a statistical technique that explains correlations among observed random variables with the help of a smaller number of unobserved factors. In traditional full factor analysis, each observed variable is influenced by…

Statistics Theory · Mathematics 2024-12-09 Mathias Drton , Alexandros Grosdos , Irem Portakal , Nils Sturma

A common problem in health research is that we have a large database with many variables measured on a large number of individuals. We are interested in measuring additional variables on a subsample; these measurements may be newly…

Methodology · Statistics 2022-03-22 Thomas Lumley , Tong Chen

Speech signals are complex intermingling of various informative factors, and this information blending makes decoding any of the individual factors extremely difficult. A natural idea is to factorize each speech frame into independent…

Sound · Computer Science 2017-06-27 Dong Wang , Lantian Li , Ying Shi , Yixiang Chen , Zhiyuan Tang

Measuring the statistical dependence between observed signals is a primary tool for scientific discovery. However, biological systems often exhibit complex non-linear interactions that currently cannot be captured without a priori knowledge…

The ability to collect and analyze large amounts of data is a growing problem within the scientific community. The growing gap between data and users calls for innovative tools that address the challenges faced by big data volume, velocity…

Databases · Computer Science 2016-08-01 Vijay Gadepally , Jeremy Kepner

Inferring causal relationships from observed data is an important task, yet it becomes challenging when the data is subject to various external interferences. Most of these interferences are the additional effects of external factors on…

Machine Learning · Computer Science 2025-11-14 Ruichu Cai , Xiaokai Huang , Wei Chen , Zijian Li , Zhifeng Hao

Latent factor model estimation typically relies on either using domain knowledge to manually pick several observed covariates as factor proxies, or purely conducting multivariate analysis such as principal component analysis. However, the…

Methodology · Statistics 2023-01-04 Runzhe Wan , Yingying Li , Wenbin Lu , Rui Song

In this work, we propose an approach for assessing sensitivity to unobserved confounding in studies with multiple outcomes. We demonstrate how prior knowledge unique to the multi-outcome setting can be leveraged to strengthen causal…

Methodology · Statistics 2023-01-26 Jiajing Zheng , Jiaxi Wu , Alexander D'Amour , Alexander Franks

High throughput metabolomics data are fraught with both non-ignorable missing observations and unobserved factors that influence a metabolite's measured concentration, and it is well known that ignoring either of these complications can…

Methodology · Statistics 2019-09-09 Chris McKennan , Carole Ober , Dan Nicolae

Pattern extraction algorithms are enabling insights into the ever-growing amount of today's datasets by translating reoccurring data properties into compact representations. Yet, a practical problem arises: With increasing data volumes and…

Information Retrieval · Computer Science 2018-07-05 Michael Behrisch , Robert Krueger , Fritz Lekschas , Tobias Schreck , Nils Gehlenborg , Hanspeter Pfister

We reconcile the two worlds of dense and sparse modeling by exploiting the positive aspects of both. We employ a factor model and assume {the dynamic of the factors is non-pervasive while} the idiosyncratic term follows a sparse vector…

Methodology · Statistics 2022-05-25 Jonas Krampe , Luca Margaritella

Identifying genes associated with complex human diseases is one of the main challenges of human genetics and computational medicine. To answer this question, millions of genetic variants get screened to identify a few of importance. To…

Genomics · Quantitative Biology 2015-09-01 Aziz M. Mezlini , Fabio Fuligni , Adam Shlien , Anna Goldenberg

Data quality is fundamentally important to ensure the reliability of data for stakeholders to make decisions. In real world applications, such as scientific exploration of extreme environments, it is unrealistic to require raw data…

Artificial Intelligence · Computer Science 2015-10-08 Dongping Fang , Elizabeth Oberlin , Wei Ding , Samuel P. Kounaves

In this work, we consider causal inference in various high-dimensional treatment settings, including for single multi-valued treatments and vector treatments with binary or continuous components, when the number of treatments can be…

Statistics Theory · Mathematics 2026-02-26 Patrick Kramer , Edward H. Kennedy , Isaac M. Opper

In a high-energy physics data analysis, the term "fake" backgrounds refers to events that would formally not satisfy the (signal) process selection criteria, but are accepted nonetheless due to mis-reconstructed particles. This can occur,…

High Energy Physics - Phenomenology · Physics 2026-01-29 Jan Gavranovič , Lara Čalić , Jernej Debevc , Else Lytken , Borut Paul Kerševan