English
Related papers

Related papers: Statistical Integration of Heterogeneous Data with…

200 papers

Two-phase outcome dependent sampling (ODS) is widely used in many fields, especially when certain covariates are expensive and/or difficult to measure. For two-phase ODS, the conditional maximum likelihood (CML) method is very attractive…

Methodology · Statistics 2022-12-21 Menglu Che , Peisong Han , Jerald F. Lawless

The integrative analysis of multiple datasets is an important strategy in data analysis. It is increasingly popular in genomics, which enjoys a wealth of publicly available datasets that can be compared, contrasted, and combined in order to…

Methodology · Statistics 2019-11-20 Sihai Dave Zhao

Graphical model estimation from multi-omics data requires a balance between statistical estimation performance and computational scalability. We introduce a novel pseudolikelihood-based graphical model framework that reparameterizes the…

Machine Learning · Statistics 2025-09-23 Sungdong Lee , Joshua Bang , Youngrae Kim , Hyungwon Choi , Sang-Yun Oh , Joong-Ho Won

Estimating the prevalence of a disease is necessary for evaluating and mitigating risks of its transmission within or between populations. Estimates that consider how prevalence changes with time provide more information about these risks…

Applications · Statistics 2021-11-12 Braden Scherting , Alison Peel , Raina Plowright , Andrew Hoegh

A ubiquitous feature of data of our era is their extra-large sizes and dimensions. Analyzing such high-dimensional data poses significant challenges, since the feature dimension is often much larger than the sample size. This thesis…

Statistics Theory · Mathematics 2025-09-11 Kai Yang

Identifying molecular signatures from complex disease patients with underlying symptomatic similarities is a significant challenge in the analysis of high dimensional multi-omics data. Topological data analysis (TDA) provides a way of…

Genomics · Quantitative Biology 2024-04-23 Davide Gurnari , Aldo Guzmán-Sáenz , Filippo Utro , Aritra Bose , Saugata Basu , Laxmi Parida

Data integration is a notoriously difficult and heuristic-driven process, especially when ground-truth data are not readily available. This paper presents a measure of uncertainty by providing maximal and minimal ranges of a query outcome…

Databases · Computer Science 2023-09-12 Deniz Turkcapar , Sanjay Krishnan

Subjects in clinical studies that investigate paired body parts can carry a disease on either both sides (bilateral) or a single side (unilateral) of the organs. Data in such studies may consist of both bilateral and unilateral records.…

Applications · Statistics 2025-11-06 Shuyi Liang , Chang-Xing Ma

Multidimensional scaling visualizes dissimilarities among objects and reduces data dimensionality. While many methods address symmetric proximity data, asymmetric and especially three-way proximity data (capturing relationships across…

Methodology · Statistics 2025-11-21 Aleix Alcacer , Rafael Benitez , Vicente J. Bolos , Irene Epifanio

Improvements in computational and experimental capabilities are rapidly increasing the amount of scientific data that is routinely generated. In applications that are constrained by memory and computational intensity, excessively large…

Machine Learning · Computer Science 2023-02-28 Malik Hassanaly , Bruce A. Perry , Michael E. Mueller , Shashank Yellapantula

Recovering dynamical equations from observed noisy data is the central challenge of system identification. We develop a statistical mechanics approach to analyze sparse equation discovery algorithms, which typically balance data fit and…

Statistical Mechanics · Physics 2025-09-16 Andrei A. Klishin , Joseph Bakarji , J. Nathan Kutz , Krithika Manohar

Partial least squares (PLS) regression combines dimensionality reduction and prediction using a latent variable model. Since partial least squares regression (PLS-R) does not require matrix inversion or diagonalization, it can be applied to…

Methodology · Statistics 2014-08-05 Tzu-Yu Liu , Laura Trinchera , Arthur Tenenhaus , Dennis Wei , Alfred O. Hero

Logistic regression models are a popular and effective method to predict the probability of categorical response data. However inference for these models can become computationally prohibitive for large datasets. Here we adapt ideas from…

Methodology · Statistics 2020-08-25 Tom Whitaker , Boris Beranger , Scott A. Sisson

Statistical depth, which measures the center-outward rank of a given sample with respect to its underlying distribution, has become a popular and powerful tool in nonparametric inference. In this paper, we investigate the use of statistical…

Methodology · Statistics 2025-11-25 Chifeng Shen , Yuejiao Fu , Michael Chen , Xiaoping Shi

Persistent homology is a powerful tool for characterizing the topology of a data set at various geometric scales. When applied to the description of molecular structures, persistent homology can capture the multiscale geometric features and…

Quantitative Methods · Quantitative Biology 2018-07-31 Zixuan Cang , Guo-Wei Wei

In engineering, accurately modeling nonlinear dynamic systems from data contaminated by noise is both essential and complex. Established Sequential Monte Carlo (SMC) methods, used for the Bayesian identification of these systems, facilitate…

Machine Learning · Statistics 2024-04-25 Joe D. Longbottom , Max D. Champneys , Timothy J. Rogers

High-dimensional data common in genomics, proteomics, and chemometrics often contains complicated correlation structures. Recently, partial least squares (PLS) and Sparse PLS methods have gained attention in these areas as dimension…

Machine Learning · Statistics 2012-04-19 Genevera I. Allen , Christine Peterson , Marina Vannucci , Mirjana Maletic-Savatic

We develop a new tool, the time inhomogeneous Poisson equation in the whole space and with a terminal condition at infinity, to study the asymptotic behavior of the non-autonomous multi-scale stochastic system with irregular coefficients,…

Probability · Mathematics 2024-12-13 Ling Wang , Pengcheng Xia , Longjie Xie , Li Yang

Both linear mixed models (LMMs) and sparse regression models are widely used in genetics applications, including, recently, polygenic modeling in genome-wide association studies. These two approaches make very different assumptions, so are…

Quantitative Methods · Quantitative Biology 2012-11-16 Xiang Zhou , Peter Carbonetto , Matthew Stephens

The advancement of polymer informatics has been significantly propelled by the integration of machine learning (ML) techniques, enabling the rapid prediction of polymer properties and expediting the discovery of high-performance polymeric…

Materials Science · Physics 2025-04-01 Jiaxin Xu , Gang Liu , Ruilan Guo , Meng Jiang , Tengfei Luo
‹ Prev 1 3 4 5 6 7 10 Next ›