中文
相关论文

相关论文: ULV: A robust statistical method for clustered dat…

200 篇论文

Statistical analysis of magnetic resonance imaging (MRI) can help radiologists to detect pathologies that are otherwise likely to be missed. Deep learning (DL) has shown promise in modeling complex spatial data for brain anomaly detection.…

计算机视觉与模式识别 · 计算机科学 2020-11-26 Victor Saase , Holger Wenz , Thomas Ganslandt , Christoph Groden , Máté E. Maros

Motivated by high-throughput single-cell cytometry data with applications to vaccine development and immunological research, we consider statistical clustering in large-scale data that contain multiple rare clusters. We propose a new…

统计方法学 · 统计学 2016-06-30 Lin Lin , Jia Li

We propose a statistical framework built on latent variable modeling for scaling laws of large language models (LLMs). Our work is motivated by the rapid emergence of numerous new LLM families with distinct architectures and training…

Network models provide a powerful framework for analysing single-cell count data, facilitating the characterisation of cellular identities, disease mechanisms, and developmental trajectories. However, uncertainty modeling in unsupervised…

基因组学 · 定量生物学 2026-04-27 Shanshan Ren , Thomas E. Bartlett , Lina Gerontogianni , Swati Chandna

Latent variable models are powerful statistical tools that can uncover relevant variation between patients or cells, by inferring unobserved hidden states from observable high-dimensional data. A major shortcoming of current methods,…

机器学习 · 统计学 2022-04-12 Arber Qoku , Florian Buettner

The recent statistical finite element method (statFEM) provides a coherent statistical framework to synthesise finite element models with observed data. Through embedding uncertainty inside of the governing equations, finite element…

统计计算 · 统计学 2021-12-30 Ömer Deniz Akyildiz , Connor Duffin , Sotirios Sabanis , Mark Girolami

We developed a single factor model with measure-specific sample weights for multivariate data with multiple observed indicators clustered within a higher level subject. The factor is therefore a latent variable shared by multiple indicators…

统计方法学 · 统计学 2019-10-22 Chengan Du , Shu-Xia Li , Zhenqiu Lin , Haiqun Lin

Studying unified model averaging estimation for situations with complicated data structures, we propose a novel model averaging method based on cross-validation (MACV). MACV unifies a large class of new and existing model averaging…

统计方法学 · 统计学 2024-12-16 Dalei Yu , Xinyu Zhang , Hua Liang

Quantifying variable importance is essential for answering high-stakes questions in fields like genetics, public policy, and medicine. Current methods generally calculate variable importance for a given model trained on a given dataset.…

机器学习 · 计算机科学 2024-04-03 Jon Donnelly , Srikar Katta , Cynthia Rudin , Edward P. Browne

We establish sample complexity guarantees for estimating the covariance matrix of a strongly log-concave smooth distribution using the unadjusted Langevin algorithm (ULA). We quantitatively compare our complexity estimates on single-chain…

概率论 · 数学 2026-02-16 Shogo Nakakita

We consider a class of latent Gaussian models with a univariate link function (ULLGMs). These are based on standard likelihood specifications (such as Poisson, Binomial, Bernoulli, Erlang, etc.) but incorporate a latent normal linear…

统计方法学 · 统计学 2025-04-11 Mark F. J. Steel , Gregor Zens

Advances in single-cell omics allow for unprecedented insights into the transcription profiles of individual cells. When combined with large-scale perturbation screens, through which specific biological mechanisms can be targeted, these…

机器学习 · 计算机科学 2023-10-24 Alejandro Tejada-Lapuerta , Paul Bertin , Stefan Bauer , Hananeh Aliee , Yoshua Bengio , Fabian J. Theis

Computational molecular modeling and visualization has seen significant progress in recent years with sev- eral molecular modeling and visualization software systems in use today. Nevertheless the molecular biology community lacks…

计算工程、金融与科学 · 计算机科学 2016-05-20 Muhibur Rasheed , Nathan Clement , Abhishek Bhowmick , Chandrajit Bajaj

Two challenging problems in the clinical study of cancer are the characterization of cancer subtypes and the classification of individual patients according to those subtypes. Statistical approaches addressing these problems are hampered by…

统计方法学 · 统计学 2012-02-28 John A. Dawson , Christina Kendziorski

Large language models (LLMs) are increasingly used as decision-support tools in data-constrained scientific workflows, where correctness and validity are critical. However, evaluation practices often emphasize stability or reproducibility…

机器学习 · 计算机科学 2026-03-18 Nazia Riasat

The standard A/B testing approaches are mostly based on t-test in large scale industry applications. These standard approaches however suffers from low statistical power in business settings, due to nature of small sample-size or…

统计方法学 · 统计学 2025-12-30 Changshuai Wei , Phuc Nguyen , Benjamin Zelditch , Joyce Chen

The sample covariance matrix is a cornerstone of multivariate statistics, but it is highly sensitive to outliers. These can be casewise outliers, such as cases belonging to a different population, or cellwise outliers, which are deviating…

统计方法学 · 统计学 2025-05-27 Fabio Centofanti , Mia Hubert , Peter J. Rousseeuw

This paper is mainly concerned with asymptotic studies of weighted bootstrap for u- and v-statistics. We derive the consistency of the weighted bootstrap u- and v-statistics, based on i.i.d. and non i.i.d. observations, from some more…

统计理论 · 数学 2012-10-23 Miklos Csorgo , Masoud M. Nasari

Advancement in sequencing technology enables the study of association between complex disorders and rare variants with low minor allele frequencies. One of the major challenges in rare variant testing is lack of statistical power of…

定量方法 · 定量生物学 2016-07-27 Rui Sun , Haoyi Weng , Inchi Hu , Junfeng Guo , William K. K. Wu , Benny Chung-Ying Zee , Maggie Haitian Wang

Compositional data, where only relative abundances are available, are common in microbiome and other high-throughput sequencing studies. Log ratios between groups of variables serve as key biomarkers in these settings. However, selecting…

统计方法学 · 统计学 2025-04-02 Jing Ma , Paizhe Xie , Kristyn Pantoja , David E. Jones