中文
相关论文

相关论文: ULV: A robust statistical method for clustered dat…

200 篇论文

Recent studies have demonstrated the feasibility of modeling single-cell data as natural languages and the potential of leveraging powerful large language models (LLMs) for understanding cell biology. However, a comprehensive evaluation of…

定量方法 · 定量生物学 2025-05-14 Fan Zhang , Tianyu Liu , Zhihong Zhu , Hao Wu , Haixin Wang , Donghao Zhou , Yefeng Zheng , Kun Wang , Xian Wu , Pheng-Ann Heng

Latent variable models such as the Variational Auto-Encoder (VAE) have become a go-to tool for analyzing biological data, especially in the field of single-cell genomics. One remaining challenge is the interpretability of latent variables…

基因组学 · 定量生物学 2023-02-20 Romain Lopez , Nataša Tagasovska , Stephen Ra , Kyunghyn Cho , Jonathan K. Pritchard , Aviv Regev

Data generated in studies of cellular regulatory systems are often qualitative. For example, measurements of signaling readouts in the presence and absence of mutations may reveal a rank ordering of responses across conditions but not the…

定量方法 · 定量生物学 2026-04-15 Ely F. Miller , Abhishek Mallela , Jacob Neumann , Yen Ting Lin , William S. Hlavacek , Richard G. Posner

VARCLUST algorithm is proposed for clustering variables under the assumption that variables in a given cluster are linear combinations of a small number of hidden latent variables, corrupted by the random noise. The entire clustering task…

Finding relevant and high-quality datasets to train machine learning models is a major bottleneck for practitioners. Furthermore, to address ambitious real-world use-cases there is usually the requirement that the data come labelled with…

机器学习 · 计算机科学 2023-10-05 Georgios Papadopoulos , Fran Silavong , Sean Moran

The sophisticated and automated means of data collection used by an increasing number of institutions and companies leads to extremely large data sets. Subset selection in regression is essential when a huge number of covariates can…

应用统计 · 统计学 2013-04-22 Debbie J. Dupuis , Maria-Pia Victoria-Feser

Clustering algorithms are pivotal in data analysis, enabling the organization of data into meaningful groups. However, individual clustering methods often exhibit inherent limitations and biases, preventing the development of a universal…

神经与进化计算 · 计算机科学 2024-12-13 H. Jahani , F. Zamio

With the rapid global spread of COVID-19, more and more data related to this virus is becoming available, including genomic sequence data. The total number of genomic sequences that are publicly available on platforms such as GISAID is…

基因组学 · 定量生物学 2021-11-16 Sarwan Ali , Murray Patterson

In the context of multilevel longitudinal data, where sample units are collected in clusters, an important aspect that should be accounted for is the unobserved heterogeneity between sample units and between clusters. For this aim we…

统计理论 · 数学 2012-08-10 F. Bartolucci , M. Lupparelli

One of the most common analysis tasks in genomic research is to identify genes that are differentially expressed (DE) between experimental conditions. Empirical Bayes (EB) statistical tests using moderated genewise variances have been very…

应用统计 · 统计学 2016-07-28 Belinda Phipson , Stanley Lee , Ian J. Majewski , Warren S. Alexander , Gordon K. Smyth

Background: Randomized controlled trials (RCTs) are costly, time-consuming, and often infeasible, while treatment-effect estimation from observational data is limited by unobserved confounding. Methods: We developed a three-step framework…

Large Language Models (LLMs) have become increasingly pervasive, finding applications across many industries and disciplines. Ensuring the trustworthiness of LLM outputs is paramount, where Uncertainty Estimation (UE) plays a key role. In…

计算与语言 · 计算机科学 2025-11-06 Kevin Wang , Subre Abdoul Moktar , Jia Li , Kangshuo Li , Feng Chen

Selective inference aims at providing valid inference after a data-driven selection of models or hypotheses. It is essential to avoid overconfident results and replicability issues. While significant advances have been made in this area for…

统计方法学 · 统计学 2025-03-14 Matteo D'Alessandro , Magne Thoresen

Verifying hardware designs in embedded systems is crucial but often labor-intensive and time-consuming. While existing solutions have improved automation, they frequently rely on unrealistic assumptions. To address these challenges, we…

硬件体系结构 · 计算机科学 2024-11-26 Yuchen Hu , Junhao Ye , Ke Xu , Jialin Sun , Shiyue Zhang , Xinyao Jiao , Dingrong Pan , Jie Zhou , Ning Wang , Weiwei Shan , Xinwei Fang , Xi Wang , Nan Guan , Zhe Jiang

Ultra high-throughput sequencing of transcriptomes (RNA-Seq) is a widely used method for quantifying gene expression levels due to its low cost, high accuracy and wide dynamic range for detection. However, the nature of RNA-Seq makes it…

统计方法学 · 统计学 2016-08-30 Hui Jiang , Tianyu Zhan

Integrating heterogeneous datasets across different measurement platforms is a fundamental challenge in many scientific applications. A common example arises in deconvolution problems, such as cell type deconvolution, where one aims to…

统计方法学 · 统计学 2025-09-30 Dongyue Xie , Lin Gui , Jingshu Wang

Reasoning language models can solve increasingly complex tasks, but struggle to produce the calibrated confidence estimates necessary for reliable deployment. Existing calibration methods usually depend on labels or repeated sampling at…

机器学习 · 计算机科学 2026-04-22 Thomas Zollo , Jimmy Wang , Richard Zemel

Rare disease diagnosis requires matching variant-bearing genes to complex patient phenotypes across large and heterogeneous evidence sources. This process remains time-intensive in current clinical interpretation pipelines. To overcome…

基因组学 · 定量生物学 2026-03-09 Jaeyeon Lee , Lin Yao , Hyun-Hwan Jeong , Zhandong Liu

Matrix valued data has become increasingly prevalent in many applications. Most of the existing clustering methods for this type of data are tailored to the mean model and do not account for the dependence structure of the features, which…

机器学习 · 统计学 2023-12-07 Inbeom Lee , Siyi Deng , Yang Ning

Structured Latent Attribute Models (SLAMs) are a family of discrete latent variable models widely used in education, psychology, and epidemiology to model multivariate categorical data. A SLAM assumes that multiple discrete latent…

统计方法学 · 统计学 2021-07-12 Yuqi Gu , Gongjun Xu