English
Related papers

Related papers: Protein-to-genome alignment with miniprot

200 papers

The third-generation long reads sequencing technologies, such as PacBio and Nanopore, have great advantages over second-generation Illumina sequencing in de novo assembly studies. However, due to the inherent low base accuracy,…

Genomics · Quantitative Biology 2020-03-27 Hengchao Wang , Bo Liu , Yan Zhang , Fan Jiang , Yuwei Ren , Lijuan Yin , Hangwei Liu , Sen Wang , Wei Fan

Background: To understand protein function, it is important to study protein- protein interaction networks. These networks can be represented in network diagrams called protein interaction maps that can lead to better understanding by…

Computational Engineering, Finance, and Science · Computer Science 2014-07-23 Mine Edes , Can Özturan , Türkan Haliloğlu , Augustin Luna , Ruth Nussinov

Genome sequencing has become a central focus in computational biology. A genome study typically begins with sequencing, which produces millions to billions of short DNA fragments known as reads. Read mapping aligns these reads to a…

The protein design problem involves finding polypeptide sequences folding into a given threedimensional structure. Its rigorous algorithmic solution is computationally demanding, involving a nested search in sequence and structure spaces.…

Quantum Physics · Physics 2024-07-11 Veronica Panizza , Philipp Hauke , Cristian Micheletti , Pietro Faccioli

Aligning Large Language Models (LLMs) with human values and preferences is essential for making them helpful and safe. However, building efficient tools to perform alignment can be challenging, especially for the largest and most competent…

Existing preference optimization objectives for language model alignment require additional hyperparameters that must be extensively tuned to achieve optimal performance, increasing both the complexity and time required for fine-tuning…

Machine Learning · Computer Science 2025-02-21 Teng Xiao , Yige Yuan , Zhengyu Chen , Mingxiao Li , Shangsong Liang , Zhaochun Ren , Vasant G Honavar

While DeepMind has tentatively solved protein folding, its inverse problem -- protein design which predicts protein sequences from their 3D structures -- still faces significant challenges. Particularly, the lack of large-scale standardized…

Quantitative Methods · Quantitative Biology 2022-02-15 Zhangyang Gao , Cheng Tan , Stan Z. Li

The goal of protein representation learning is to extract knowledge from protein databases that can be applied to various protein-related downstream tasks. Although protein sequence, structure, and function are the three key modalities for…

Biomolecules · Quantitative Biology 2024-05-14 Eunji Ko , Seul Lee , Minseon Kim , Dongki Kim

Human preference data is essential for aligning large language models (LLMs) with human values, but collecting such data is often costly and inefficient-motivating the need for efficient data selection methods that reduce annotation costs…

Computation and Language · Computer Science 2026-04-21 Seohyeong Lee , Eunwon Kim , Hwaran Lee , Buru Chang

We describe a new global multiple alignment program capable of aligning a large number of genomic regions. Our progressive alignment approach incorporates the following ideas: maximum-likelihood inference of ancestral sequences, automatic…

Genomics · Quantitative Biology 2007-05-23 Nicolas Bray , Lior Pachter

We propose a metric for the space of multiple sequence alignments that can be used to compare two alignments to each other. In the case where one of the alignments is a reference alignment, the resulting accuracy measure improves upon…

Quantitative Methods · Quantitative Biology 2011-11-09 Ariel S. Schwartz , Eugene W. Myers , Lior Pachter

Multiple genome alignment remains a challenging problem. Effects of recombination including rearrangement, segmental duplication, gain, and loss can create a mosaic pattern of homology even among closely related organisms. We describe a…

Genomics · Quantitative Biology 2009-11-02 Aaron E. Darling , Bob Mau , Nicole T. Perna

In the present article, we propose a paradigm shift on evolving Artificial Neural Networks (ANNs) towards a new bio-inspired design that is grounded on the structural properties, interactions, and dynamics of protein networks (PNs): the…

Neural and Evolutionary Computing · Computer Science 2024-06-10 Oscar Lao , Konstantinos Zacharopoulos , Apostolos Fournaris , Rossano Schifanella , Ioannis Arapakis

For protein sequence datasets, unlabeled data has greatly outpaced labeled data due to the high cost of wet-lab characterization. Recent deep-learning approaches to protein prediction have shown that pre-training on unlabeled data can yield…

Machine Learning · Computer Science 2020-12-02 Pascal Sturmfels , Jesse Vig , Ali Madani , Nazneen Fatema Rajani

Motivation: The gene content regulates the biology of an organism. It varies between species and between individuals of the same species. Although tools have been developed to identify gene content changes in bacterial genomes, none is…

Genomics · Quantitative Biology 2024-05-30 Heng Li , Maximillian Marin , Maha Reda Farhat

Excitement at the prospect of using data-driven generative models to sample configurational ensembles of biomolecular systems stems from the extraordinary success of these models on a diverse set of high-dimensional sampling tasks. Unlike…

Statistical Mechanics · Physics 2024-02-06 Shriram Chennakesavalu , Grant M. Rotskoff

Accurately annotating and controlling protein function from sequence data remains a major challenge, particularly within homologous families where annotated sequences are scarce and structural variation is minimal. We present a two-stage…

Quantitative Methods · Quantitative Biology 2025-07-22 Lorenzo Rosset , Martin Weigt , Francesco Zamponi

Gene finding is the task of identifying the locations of coding sequences within the vast amount of genetic code contained in the genome. With an ever increasing quantity of raw genome sequences, gene finding is an important avenue towards…

Genomics · Quantitative Biology 2025-05-07 Frederikke I. Marin , Dennis Pultz , Wouter Boomsma

Motivation: The ability to generate massive amounts of sequencing data continues to overwhelm the processing capability of existing algorithms and compute infrastructures. In this work, we explore the use of hardware/software co-design and…

Computational Engineering, Finance, and Science · Computer Science 2020-10-29 Mohammed Alser , Hasan Hassan , Akash Kumar , Onur Mutlu , Can Alkan

A molecular understanding of how protein function is related to protein structure will require an ability to understand large conformational changes between multiple states. Unfortunately these states are often separated by high free energy…

Biological Physics · Physics 2011-08-08 Juan R. Perilla , Thomas B. Woolf