English
Related papers

Related papers: Per-sample immunoglobulin germline inference from …

200 papers

Clustering is part of unsupervised analysis methods that consist in grouping samples into homogeneous and separate subgroups of observations also called clusters. To interpret the clusters, statistical hypothesis testing is often used to…

Methodology · Statistics 2022-10-25 Benjamin Hivert , Denis Agniel , Rodolphe Thiébaut , Boris P Hejblum

High throughput genome sequencing technologies such as RNA-Seq and Microarray have the potential to transform clinical decision making and biomedical research by enabling high-throughput measurements of the genome at a granular level.…

Recent advances in high-resolution sequencing have paved the way for population-scale analysis in single-cell RNA-sequencing (scRNA-seq) data. scRNA-seq data, in particular, have proven to be extremely powerful in profiling a variety of…

Methodology · Statistics 2025-10-30 Hanxuan Ye , Zachary Qian , Hongzhe Li

Intuitively, unfamiliarity should lead to lack of confidence. In reality, current algorithms often make highly confident yet wrong predictions when faced with relevant but unfamiliar examples. A classifier we trained to recognize gender is…

Computer Vision and Pattern Recognition · Computer Science 2020-09-09 Zhizhong Li , Derek Hoiem

Motivation: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands…

Genomics · Quantitative Biology 2025-09-23 Siying Yang , Neng Huang , Heng Li

High-throughput sequencing technology provides unprecedented opportunities to quantitatively explore human gut microbiome and its relation to diseases. Microbiome data are compositional, sparse, noisy, and heterogeneous, which pose serious…

Methodology · Statistics 2020-10-12 Fangting Zhou , Kejun He , Qiwei Li , Robert S. Chapkin , Yang Ni

Partial label learning (PLL) is a typical weakly supervised learning problem, where each training example is associated with a set of candidate labels among which only one is true. Most existing PLL approaches assume that the incorrect…

Machine Learning · Computer Science 2021-10-27 Ning Xu , Congyu Qiao , Xin Geng , Min-Ling Zhang

We proposed a machine learning approach to identify and distinguish dusty stellar sources employing supervised and unsupervised methods and categorizing point sources, mainly evolved stars, using photometric and spectroscopic data collected…

Reconstructing the evolutionary history relating a collection of molecular sequences is the main subject of modern Bayesian phylogenetic inference. However, the commonly used Markov chain Monte Carlo methods can be inefficient due to the…

Machine Learning · Statistics 2024-08-12 Tianyu Xie , Frederick A. Matsen , Marc A. Suchard , Cheng Zhang

The main contribution of this paper is a simple semi-supervised pipeline that only uses the original training set without collecting extra data. It is challenging in 1) how to obtain more training data only from the training set and 2) how…

Computer Vision and Pattern Recognition · Computer Science 2017-08-23 Zhedong Zheng , Liang Zheng , Yi Yang

In cancer biomarker development, a key objective is to evaluate whether a new biomarker, when combined with an established one, improves early cancer detection compared to using the established biomarker alone. Incremental value is often…

Methodology · Statistics 2025-11-21 Indrila Ganguly , Ying Huang

In light of the recent advancements in machine learning, we propose a novel approach to neutron source distribution estimation through the utilisation of probabilistic generative models. The estimation is based on a Monte Carlo particle…

Instrumentation and Detectors · Physics 2026-05-13 Jose Ignacio Robledo , Norberto Schmidt , Klaus Lieutenant , Jingjing Li , Stefan Kesselheim , Paul Zakalek

Repertoire-level analysis of T cell receptors offers a biologically grounded signal for disease detection and immune monitoring, yet practical deployment is impeded by label sparsity, cohort heterogeneity, and the computational burden of…

Machine Learning · Computer Science 2026-04-23 Rong Fu , Muge Qi , Yang Li , Yabin Jin , Jiekai Wu , Jiaxuan Lu , Chunlei Meng , Youjin Wang , Zeli Su , Juntao Gao , Li Bao , Qi Zhao , Wei Luo , Simon Fong

Biological sequence comparison is a key step in inferring the relatedness of various organisms and the functional similarity of their components. Thanks to the Next Generation Sequencing efforts, an abundance of sequence data is now…

Machine Learning · Computer Science 2016-09-13 Dhananjay Kimothi , Akshay Soni , Pravesh Biyani , James M. Hogan

The person re-identification task requires to robustly estimate visual similarities between person images. However, existing person re-identification models mostly estimate the similarities of different image pairs of probe and gallery…

Computer Vision and Pattern Recognition · Computer Science 2018-07-27 Yantao Shen , Hongsheng Li , Shuai Yi , Dapeng Chen , Xiaogang Wang

Clustering analysis is fundamental in single-cell RNA sequencing (scRNA-seq) data analysis for elucidating cellular heterogeneity and diversity. Recent graph-based scRNA-seq clustering methods, particularly graph neural networks (GNNs),…

Machine Learning · Computer Science 2025-07-15 Ping Xu , Pengfei Wang , Zhiyuan Ning , Meng Xiao , Min Wu , Yuanchun Zhou

The classical Gaussian mixture model assumes homogeneity within clusters, an assumption that often fails in real-world data where observations naturally exhibit varying scales or intensities. To address this, we introduce the…

Machine Learning · Statistics 2026-04-08 Huan Qing

Comprehensive discovery of structural variation (SV) in human genomes from DNA sequencing requires the integration of multiple alignment signals including read-pair, split-read and read-depth. However, owing to inherent technical…

Genomics · Quantitative Biology 2014-01-23 Ryan M. Layer , Ira M. Hall , Aaron R. Quinlan

Genetic variants identified to date by genome-wide association studies only explain a small fraction of total heritability. Gene-by-gene interaction is one important potential source of unexplained heritability. In the first part of this…

Methodology · Statistics 2016-05-10 Chen Lu

We present the methods and results of a two-stage modeling process that generates candidate gene-regulatory networks of the bacterium B. subtilis from experimentally obtained, yet mathematically underdetermined microchip array data. By…

Molecular Networks · Quantitative Biology 2009-11-13 C. Christensen , A. Gupta , C. D. Maranas , R. Albert
‹ Prev 1 8 9 10 Next ›