中文
相关论文

相关论文: CompostBin: A DNA composition-based algorithm for …

200 篇论文

Understanding the relationship between atomic structure (order) and chemical composition (chemistry) is critical for advancing materials science, yet traditional spectroscopic techniques can be slow and damaging to sensitive samples.…

材料科学 · 物理学 2025-08-29 Mridul Kumar , Yevgeny Rakita

Background: In the metagenome assembly of a microbiome community, we may think abundant species would be easier to assemble due to their deeper coverage. However, this conjucture is rarely tested. We often do not know how many abundant…

基因组学 · 定量生物学 2022-11-23 Xiaowen Feng , Heng Li

Genome sequence analysis, which examines the DNA sequences of organisms, drives advances in many critical medical and biotechnological fields. Given its importance and the exponentially growing volumes of genomic sequence data, there are…

In this paper a novel biclustering algorithm based on artificial intelligence (AI) is introduced. The method called EBIC aims to detect biologically meaningful, order-preserving patterns in complex data. The proposed algorithm is probably…

机器学习 · 计算机科学 2018-07-27 Patryk Orzechowski , Moshe Sipper , Xiuzhen Huang , Jason H. Moore

Evolutionary crystal structure prediction proved to be a powerful approach for studying a wide range of materials. Here, we present a specifically designed algorithm for the prediction of the structure of complex crystals consisting of…

材料科学 · 物理学 2012-05-21 Qiang Zhu , Artem R. Oganov , Colin W. Glass , Harold T. Stokes

Ensembles are popular methods for solving practical supervised learning problems. They reduce the risk of having underperforming models in production-grade software. Although critical, methods for learning heterogeneous regression ensembles…

机器学习 · 计算机科学 2018-04-18 Jihed Khiari , Luis Moreira-Matias , Ammar Shaker , Bernard Zenko , Saso Dzeroski

The combination of data science and materials informatics has significantly propelled the advancement of multi-component compound synthesis research. This study employs atomic-level data to predict miscibility in binary compounds using…

材料科学 · 物理学 2024-09-05 Chiwen Feng , Yanwei Liang , Jiaying Sun , Renhai Wang , Huaijun Sun , Huafeng Dong

Adequate read filtering is critical when processing high-throughput data in marker-gene-based studies. Sequencing errors can cause the mis-clustering of otherwise similar reads, artificially increasing the number of retrieved Operational…

定量方法 · 定量生物学 2015-06-02 Fernando Puente-Sánchez , Jacobo Aguirre , Víctor Parro

Composite DNA is a recent novel method to increase the information capacity of DNA-based data storage above the theoretical limit of 2 bits/symbol. In this method, every composite symbol does not store a single DNA nucleotide but a mixture…

信息论 · 计算机科学 2025-01-22 Tuan Thanh Nguyen , Chen Wang , Kui Cai , Yiwei Zhang , Zohar Yakhini

In the context of variable selection, ensemble learning has gained increasing interest due to its great potential to improve selection accuracy and to reduce false discovery rate. A novel ordering-based selective ensemble learning strategy…

机器学习 · 统计学 2017-04-28 Chunxia Zhang , Yilei Wu , Mu Zhu

Mixed-membership (MM) models such as Latent Dirichlet Allocation (LDA) have been applied to microbiome compositional data to identify latent subcommunities of microbial species. These subcommunities are informative for understanding the…

应用统计 · 统计学 2022-05-18 Patrick LeBlanc , Li Ma

DNA has immense potential as an emerging data storage medium. The principle of DNA storage is the conversion and flow of digital information between binary code stream, quaternary base, and actual DNA fragments. This process will inevitably…

信息检索 · 计算机科学 2022-10-21 Yun Qin , Fei Zhu , Bo Xi

When analyzing communities of microorganisms from their sequenced DNA, an important task is taxonomic profiling: enumerating the presence and relative abundance of all organisms, or merely of all taxa, contained in the sample. This task can…

基因组学 · 定量生物学 2020-01-24 Simon Foucart , David Koslicki

Metagenomics is the study of environments through genetic sampling of their microbiota. Metagenomic studies produce large datasets that are estimated to grow at a faster rate than the available computational capacity. A key step in the…

分布式、并行与集群计算 · 计算机科学 2013-10-04 Freddie Sunarso , Srikumar Venugopal , Federico Lauro

Genome-resolved metagenomics has contributed largely to discovering prokaryotic genomes. When applied to microscopic eukaryotes, challenges such as the high number of introns and repeat regions found in nuclear genomes have hampered the…

基因组学 · 定量生物学 2025-10-14 Yuhao Tong , Vanessa Rossetto Marcelino , Robert Turnbull , Heroen Verbruggen

DNA sequencing, especially of microbial genomes and metagenomes, has been at the core of recent research advances in large-scale comparative genomics. The data deluge has resulted in exponential growth in genomic datasets over the past…

Advances in next-generation sequencing technology have enabled the high-throughput profiling of metagenomes and accelerated the microbiome study. Recently, there has been a rise in quantitative studies that aim to decipher the microbiome…

统计方法学 · 统计学 2023-08-29 Kevin C. Lutz , Michael L. Neugent , Tejasv Bedi , Nicole J. De Nisco , Qiwei Li

In shotgun sequencing, the input string (typically, a long DNA sequence composed of nucleotide bases) is sequenced as multiple overlapping fragments of much shorter lengths (called \textit{reads}). Modelling the shotgun sequencing pipeline…

信息论 · 计算机科学 2024-05-14 Hrishi Narayanan , Prasad Krishnan , Nita Parekh

Metagenome, a mixture of different genomes (as a rule, bacterial), represents a pattern, and the analysis of its composition is, currently, one of the challenging problems of bioinformatics. In the present study, the possibility of…

定量方法 · 定量生物学 2016-11-04 Valery Kirzhner , Zeev Volkovich , Renata Avros , Katerina Korenblat

Motivation: Estimation of bacterial community composition from a high-throughput sequenced sample is an important task in metagenomics applications. Since the sample sequence data typically harbors reads of variable lengths and different…