English
Related papers

Related papers: Protein-to-genome alignment with miniprot

200 papers

Multiple sequence alignment (MSA) has been one of the most important problems in bioinformatics for more decades and it is still heavily examined by many mathematicians and biologists. However, mostly because of the practical motivation of…

Quantitative Methods · Quantitative Biology 2015-11-17 Kristóf Takács

Advances in deep learning have opened an era of abundant and accurate predicted protein structures; however, similar progress in protein ensembles has remained elusive. This review highlights several recent research directions towards…

Biomolecules · Quantitative Biology 2025-09-23 Bowen Jing , Bonnie Berger , Tommi Jaakkola

The knowledge regarding the function of proteins is necessary as it gives a clear picture of biological processes. Nevertheless, there are many protein sequences found and added to the databases but lacks functional annotation. The…

Quantitative Methods · Quantitative Biology 2018-09-13 Anu Vazhayil , Vinayakumar R , Soman KP

Genomics is changing our understanding of humans, evolution, diseases, and medicines to name but a few. As sequencing technology is developed collecting DNA sequences takes less time thereby generating more genetic data every day. Today the…

Quantitative Methods · Quantitative Biology 2020-07-29 Sahand Salamat , Tajana Rosing

Addressing the growing demands of artificial intelligence (AI) and data analytics requires new computing approaches. In this paper, we propose a reconfigurable hardware accelerator designed specifically for AI and data-intensive…

Hardware Architecture · Computer Science 2026-02-05 Md Rownak Hossain Chowdhury , Mostafizur Rahman

Current AI-assisted protein design mainly utilizes protein sequential and structural information. Meanwhile, there exists tremendous knowledge curated by humans in the text format describing proteins' high-level functionalities. Yet,…

Sequence alignments are fundamental to bioinformatics which has resulted in a variety of optimized implementations. Unfortunately, the vast majority of them are hand-tuned and specific to certain architectures and execution models. This not…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-05-17 André Müller , Bertil Schmidt , Andreas Hildebrandt , Richard Membarth , Roland Leißa , Matthis Kruse , Sebastian Hack

Classification of proteins based on their structure provides a valuable resource for studying protein structure, function and evolutionary relationships. With the rapidly increasing number of known protein structures, manual and…

Computational Engineering, Finance, and Science · Computer Science 2009-07-14 Oktie Hassanzadeh

The two most common data-structures for genome indexing, FM-indices and hash-tables, exhibit a fundamental trade-off between memory footprint and performance. We present Ranger, a new indexing technique for nucleotide sequences that is both…

Data Structures and Algorithms · Computer Science 2023-08-09 Alon Rashelbach , Ori Rottensterich , Mark Silberstien

Motivation Protein fold recognition is an important problem in structural bioinformatics. Almost all traditional fold recognition methods use sequence (homology) comparison to indirectly predict the fold of a tar get protein based on the…

Machine Learning · Computer Science 2017-06-06 Jie Hou , Badri Adhikari , Jianlin Cheng

Generating protein sequences conditioned on protein structures is an impactful technique for protein engineering. When synthesizing engineered proteins, they are commonly translated into DNA and expressed in an organism such as yeast. One…

Machine Learning · Computer Science 2024-09-27 Hannes Stark , Umesh Padia , Julia Balla , Cameron Diao , George Church

Deciphering protein function remains a fundamental challenge in protein representation learning. The task presents significant difficulties for protein language models (PLMs) due to the sheer volume of functional annotation categories and…

Biomolecules · Quantitative Biology 2025-06-10 Zixuan Jiang , Renjing Xu

The Gene or DNA sequence in every cell does not control genetic properties on its own; Rather, this is done through translation of DNA into protein and subsequent formation of a certain 3D structure. The biological function of a protein is…

Computational Engineering, Finance, and Science · Computer Science 2019-05-30 Leila Khalatbari , Mohammad Reza Kangavari , Saeid Hosseini , Hongzhi Yin , Ngai-Man Cheung

We introduce an improved version of RECKONER, an error corrector for Illumina whole genome sequencing data. By modifying its workflow we reduce the computation time even 10 times. We also propose a new method of determination of $k$-mer…

Genomics · Quantitative Biology 2017-03-03 Maciej Dlugosz , Sebastian Deorowicz , Marek Kokot

Various methods have been developed to analyze the association between organisms and their genomic sequences. Among them, sequence alignment is the most frequently used for comparative analysis of biological genomes. However, the…

Quantitative Methods · Quantitative Biology 2020-10-27 Yong Joon Song , Dong Jin Ji , Hye In Seo , Gyu Bum Han , Dong Ho Cho

Aligning protein-protein interaction (PPI) networks of different species has drawn a considerable interest recently. This problem is important to investigate evolutionary conserved pathways or protein complexes across species, and to help…

Optimization and Control · Mathematics 2009-05-08 Mikhail Zaslavskiy , Francis Bach , Jean-Philippe Vert

Calculating the similarities between a pair of genomic sequences is one of the most fundamental computational steps in genomic analysis. This step -- called sequence alignment -- is the computational bottleneck because: (1) it is…

Genomics · Quantitative Biology 2019-10-10 Mohammed H K Alser

Conformational ensembles of protein structures are immensely important both for understanding protein function and drug discovery in novel modalities such as cryptic pockets. Current techniques for sampling ensembles such as molecular…

Biological Physics · Physics 2025-10-24 Ameya Daigavane , Bodhi P. Vani , Darcy Davidson , Saeed Saremi , Joshua Rackers , Joseph Kleinhenz

Pairwise alignment of DNA sequencing data is a ubiquitous task in bioinformatics and typically represents a heavy computational burden. State-of-the-art approaches to speed up this task use hashing to identify short segments (k-mers) that…

Machine Learning · Computer Science 2021-02-16 Govinda M. Kamath , Tavor Z. Baharav , Ilan Shomorony

Designing DNA and protein sequences with improved function has the potential to greatly accelerate synthetic biology. Machine learning models that accurately predict biological fitness from sequence are becoming a powerful tool for…

Machine Learning · Computer Science 2022-03-17 Johannes Linder , Georg Seelig