中文
相关论文

相关论文: Joint discovery of haplotype blocks and complex tr…

200 篇论文

The success of deep learning in high-dimensional settings is often attributed to the presence of low-dimensional structure in real-world data. While standard theoretical models typically assume that this structure lies in the target…

机器学习 · 计算机科学 2026-05-15 Elisabetta Cornacchia , Laurent Massoulié

Discovering all the genetic causes of a phenotype is an important goal in functional genomics. In this paper we combine an experimental design for multiple independent detections of the genetic causes of a phenotype, with a high-throughput…

定量方法 · 定量生物学 2014-03-18 Marc Harper , Luisa Gronenberg , James Liao , Christopher Lee

Genome-wide eQTL mapping explores the relationship between gene expression values and DNA variants to understand genetic causes of human disease. Due to the large number of genes and DNA variants that need to be assessed simultaneously,…

应用统计 · 统计学 2018-04-10 Jacob Rhyne , Jung-Ying Tzeng , Teng Zhang , X. Jessie Jeng

We develop a latent variable model and an efficient spectral algorithm motivated by the recent emergence of very large data sets of chromatin marks from multiple human cell types. A natural model for chromatin data in one cell type is a…

机器学习 · 统计学 2015-06-09 Chicheng Zhang , Jimin Song , Kevin C Chen , Kamalika Chaudhuri

The SNPs (Single Nucleotide Polymorphisms) genotyping platforms are of great value for gene mapping of complex diseases. Nowadays, the high-density of these molecular markers enables studies of dependence patterns between loci over the…

统计方法学 · 统计学 2013-02-25 André J. Bianchi , Suely R. Giolo , Júlia P. Soler , Florencia Leonardi

Tumor samples are heterogeneous. They consist of different subclones that are characterized by differences in DNA nucleotide sequences and copy numbers on multiple loci. Heterogeneity can be measured through the identification of the…

统计方法学 · 统计学 2014-09-26 Juhee Lee , Peter Mueller , Subhajit Sengupta , Kamalakar Gulukota , Yuan Ji

We are interested in the widespread problem of clustering documents and finding topics in large collections of written documents in the presence of metadata and hyperlinks. To tackle the challenge of accounting for these different types of…

社会与信息网络 · 计算机科学 2021-07-01 Charles C. Hyland , Yuanming Tao , Lamiae Azizi , Martin Gerlach , Tiago P. Peixoto , Eduardo G. Altmann

Genotype-to-phenotype mappings translate genotypic variations such as mutations into phenotypic changes. Neutrality is the observation that some mutations do not lead to phenotypic changes. Studying the search trajectories in genotypic and…

种群与进化 · 定量生物学 2023-06-26 Ting Hu , Gabriela Ochoa , Wolfgang Banzhaf

Recent studies reveal even the smallest genomes such as viruses evolve through complex and stochastic processes, and the assumption of independent alleles is not valid in most applications. Advances in sequencing technologies produce…

种群与进化 · 定量生物学 2017-10-30 Hyunjin Shim

High throughput sequencing is a technology that allows for the generation of millions of reads of genomic data regarding a study of interest, and data from high throughput sequencing platforms are usually count compositions. Subsequent…

定量方法 · 定量生物学 2017-04-07 Jia R. Wu , Jean M. Macklaim , Briana L. Genge , Gregory B. Gloor

Complexity metrics and machine learning (ML) models have been utilized to analyze the lengths of segmental genomic entities like: exons, introns, intergenic and repeat/unique DNA sequences, in each of the 22 human chromosomes. The purpose…

While progress has been made in identifying common genetic variants associated with human diseases, for most of common complex diseases, the identified genetic variants only account for a small proportion of heritability. Challenges remain…

应用统计 · 统计学 2025-08-18 Olga A. Vsevolozhskaya , Dmitri V. Zaykin , Mark C. Greenwood , Changshuai Wei , Qing Lu

Predicting genetic perturbations enables the identification of potentially crucial genes prior to wet-lab experiments, significantly improving overall experimental efficiency. Since genes are the foundation of cellular life, building gene…

定量方法 · 定量生物学 2025-05-09 Changxi Chi , Jun Xia , Jingbo Zhou , Jiabei Cheng , Chang Yu , Stan Z. Li

RNA molecules are known to form complex secondary structures including pseudoknots. A systematic framework for the enumeration, classification and prediction of secondary structures is critical to determine the biological significance of…

生物大分子 · 定量生物学 2025-12-24 Rayan Ibrahim , Allison H. Moore

In community detection on graphs, the semi-supervised learning problem entails inferring the ground-truth membership of each node in a graph, given the connectivity structure and a limited number of revealed node labels. Different subsets…

无序系统与神经网络 · 物理学 2022-03-22 Hugo Cui , Luca Saglietti , Lenka Zdeborová

We propose a generalized stochastic block model to explore the mesoscopic structures in signed networks by grouping vertices that exhibit similar positive and negative connection profiles into the same cluster. In this model, the group…

社会与信息网络 · 计算机科学 2015-06-17 Jonathan Q. Jiang

Mutation rate variation across loci is well known to cause difficulties, notably identifiability issues, in the reconstruction of evolutionary trees from molecular sequences. Here we introduce a new approach for estimating general…

概率论 · 数学 2011-09-30 Elchanan Mossel , Sebastien Roch

Gene annotation addresses the problem of predicting unknown associations between gene and functions (e.g., biological processes) of a specific organism. Despite recent advances, the cost and time demanded by annotation procedures that rely…

机器学习 · 计算机科学 2022-05-02 Miguel Romero , Oscar Ramírez , Jorge Finke , Camilo Rocha

The stochastic block model (SBM) is a random graph model with different group of vertices connecting differently. It is widely employed as a canonical model to study clustering and community detection, and provides a fertile ground to study…

概率论 · 数学 2023-10-26 Emmanuel Abbe

High-throughput sequencing (HTS) is revolutionizing biological research by enabling scientists to quickly and cheaply query variation at a genomic scale. Despite the increasing ease of obtaining such data, using these data effectively still…

基因组学 · 定量生物学 2012-11-09 Sonal Singhal