English
Related papers

Related papers: ULV: A robust statistical method for clustered dat…

200 papers

Statistical analysis of magnetic resonance imaging (MRI) can help radiologists to detect pathologies that are otherwise likely to be missed. Deep learning (DL) has shown promise in modeling complex spatial data for brain anomaly detection.…

Computer Vision and Pattern Recognition · Computer Science 2020-11-26 Victor Saase , Holger Wenz , Thomas Ganslandt , Christoph Groden , Máté E. Maros

Motivated by high-throughput single-cell cytometry data with applications to vaccine development and immunological research, we consider statistical clustering in large-scale data that contain multiple rare clusters. We propose a new…

Methodology · Statistics 2016-06-30 Lin Lin , Jia Li

We propose a statistical framework built on latent variable modeling for scaling laws of large language models (LLMs). Our work is motivated by the rapid emergence of numerous new LLM families with distinct architectures and training…

Network models provide a powerful framework for analysing single-cell count data, facilitating the characterisation of cellular identities, disease mechanisms, and developmental trajectories. However, uncertainty modeling in unsupervised…

Genomics · Quantitative Biology 2026-04-27 Shanshan Ren , Thomas E. Bartlett , Lina Gerontogianni , Swati Chandna

Latent variable models are powerful statistical tools that can uncover relevant variation between patients or cells, by inferring unobserved hidden states from observable high-dimensional data. A major shortcoming of current methods,…

Machine Learning · Statistics 2022-04-12 Arber Qoku , Florian Buettner

The recent statistical finite element method (statFEM) provides a coherent statistical framework to synthesise finite element models with observed data. Through embedding uncertainty inside of the governing equations, finite element…

Computation · Statistics 2021-12-30 Ömer Deniz Akyildiz , Connor Duffin , Sotirios Sabanis , Mark Girolami

We developed a single factor model with measure-specific sample weights for multivariate data with multiple observed indicators clustered within a higher level subject. The factor is therefore a latent variable shared by multiple indicators…

Methodology · Statistics 2019-10-22 Chengan Du , Shu-Xia Li , Zhenqiu Lin , Haiqun Lin

Studying unified model averaging estimation for situations with complicated data structures, we propose a novel model averaging method based on cross-validation (MACV). MACV unifies a large class of new and existing model averaging…

Methodology · Statistics 2024-12-16 Dalei Yu , Xinyu Zhang , Hua Liang

Quantifying variable importance is essential for answering high-stakes questions in fields like genetics, public policy, and medicine. Current methods generally calculate variable importance for a given model trained on a given dataset.…

Machine Learning · Computer Science 2024-04-03 Jon Donnelly , Srikar Katta , Cynthia Rudin , Edward P. Browne

We establish sample complexity guarantees for estimating the covariance matrix of a strongly log-concave smooth distribution using the unadjusted Langevin algorithm (ULA). We quantitatively compare our complexity estimates on single-chain…

Probability · Mathematics 2026-02-16 Shogo Nakakita

We consider a class of latent Gaussian models with a univariate link function (ULLGMs). These are based on standard likelihood specifications (such as Poisson, Binomial, Bernoulli, Erlang, etc.) but incorporate a latent normal linear…

Methodology · Statistics 2025-04-11 Mark F. J. Steel , Gregor Zens

Advances in single-cell omics allow for unprecedented insights into the transcription profiles of individual cells. When combined with large-scale perturbation screens, through which specific biological mechanisms can be targeted, these…

Machine Learning · Computer Science 2023-10-24 Alejandro Tejada-Lapuerta , Paul Bertin , Stefan Bauer , Hananeh Aliee , Yoshua Bengio , Fabian J. Theis

Computational molecular modeling and visualization has seen significant progress in recent years with sev- eral molecular modeling and visualization software systems in use today. Nevertheless the molecular biology community lacks…

Computational Engineering, Finance, and Science · Computer Science 2016-05-20 Muhibur Rasheed , Nathan Clement , Abhishek Bhowmick , Chandrajit Bajaj

Two challenging problems in the clinical study of cancer are the characterization of cancer subtypes and the classification of individual patients according to those subtypes. Statistical approaches addressing these problems are hampered by…

Methodology · Statistics 2012-02-28 John A. Dawson , Christina Kendziorski

Large language models (LLMs) are increasingly used as decision-support tools in data-constrained scientific workflows, where correctness and validity are critical. However, evaluation practices often emphasize stability or reproducibility…

Machine Learning · Computer Science 2026-03-18 Nazia Riasat

The standard A/B testing approaches are mostly based on t-test in large scale industry applications. These standard approaches however suffers from low statistical power in business settings, due to nature of small sample-size or…

Methodology · Statistics 2025-12-30 Changshuai Wei , Phuc Nguyen , Benjamin Zelditch , Joyce Chen

The sample covariance matrix is a cornerstone of multivariate statistics, but it is highly sensitive to outliers. These can be casewise outliers, such as cases belonging to a different population, or cellwise outliers, which are deviating…

Methodology · Statistics 2025-05-27 Fabio Centofanti , Mia Hubert , Peter J. Rousseeuw

This paper is mainly concerned with asymptotic studies of weighted bootstrap for u- and v-statistics. We derive the consistency of the weighted bootstrap u- and v-statistics, based on i.i.d. and non i.i.d. observations, from some more…

Statistics Theory · Mathematics 2012-10-23 Miklos Csorgo , Masoud M. Nasari

Advancement in sequencing technology enables the study of association between complex disorders and rare variants with low minor allele frequencies. One of the major challenges in rare variant testing is lack of statistical power of…

Quantitative Methods · Quantitative Biology 2016-07-27 Rui Sun , Haoyi Weng , Inchi Hu , Junfeng Guo , William K. K. Wu , Benny Chung-Ying Zee , Maggie Haitian Wang

Compositional data, where only relative abundances are available, are common in microbiome and other high-throughput sequencing studies. Log ratios between groups of variables serve as key biomarkers in these settings. However, selecting…

Methodology · Statistics 2025-04-02 Jing Ma , Paizhe Xie , Kristyn Pantoja , David E. Jones