English
Related papers

Related papers: CoPRA: Bridging Cross-domain Pretrained Sequence M…

200 papers

Less than 1% of protein sequences are structurally and functionally annotated. Natural Language Processing (NLP) community has recently embraced self-supervised learning as a powerful approach to learn representations from unlabeled text,…

Biomolecules · Quantitative Biology 2020-12-08 Modestas Filipavicius , Matteo Manica , Joris Cadow , Maria Rodriguez Martinez

Geometric deep learning has recently achieved great success in non-Euclidean domains, and learning on 3D structures of large biomolecules is emerging as a distinct research area. However, its efficacy is largely constrained due to the…

Machine Learning · Computer Science 2023-10-31 Fang Wu , Lirong Wu , Dragomir Radev , Jinbo Xu , Stan Z. Li

Protein language models (PLMs) have demonstrated remarkable capabilities in learning relationships between protein sequences and functions. However, finetuning these large models requires substantial computational resources, often with…

Machine Learning · Computer Science 2025-12-09 Shuo Zhang , Jian K. Liu

Recent advances in Artificial Intelligence have enabled multi-modal systems to model and translate diverse information spaces. Extending beyond text and vision, we introduce OneProt, a multi-modal AI for proteins that integrates structural,…

Pre-trained models have been successful in many protein engineering tasks. Most notably, sequence-based models have achieved state-of-the-art performance on protein fitness prediction while structure-based models have been used…

Machine Learning · Computer Science 2023-07-25 Antonia Boca , Simon Mathis

Structural prediction has long been considered critical in RNA research, especially following the success of AlphaFold2 in protein studies, which has drawn significant attention to the field. While recent advances in machine learning and…

Biomolecules · Quantitative Biology 2024-09-26 Jiaxing Yang

In recent years, there has been a surge in the development of 3D structure-based pre-trained protein models, representing a significant advancement over pre-trained protein language models in various downstream tasks. However, most existing…

Machine Learning · Computer Science 2024-06-04 Jiale Zhao , Wanru Zhuang , Jia Song , Yaqi Li , Shuqi Lu

Accurate identification of protein nucleic-acid-binding residues poses a significant challenge with important implications for various biological processes and drug design. Many typical computational methods for protein analysis rely on a…

Biomolecules · Quantitative Biology 2023-12-21 Linglin Jing , Sheng Xu , Yifan Wang , Yuzhe Zhou , Tao Shen , Zhigang Ji , Hui Fang , Zhen Li , Siqi Sun

Protein-ligand structure prediction is an essential task in drug discovery, predicting the binding interactions between small molecules (ligands) and target proteins (receptors). Recent advances have incorporated deep learning techniques to…

In this paper, we use the biological domain knowledge incorporated into stochastic models for ab initio RNA secondary-structure prediction to improve the state of the art in joint compression of RNA sequence and structure data (Liu et al.,…

Biomolecules · Quantitative Biology 2023-02-24 Evarista Onokpasa , Sebastian Wild , Prudence W. H. Wong

Empirical scoring functions based on either molecular force fields or cheminformatics descriptors are widely used, in conjunction with molecular docking, during the early stages of drug discovery to predict potency and binding affinity of a…

Machine Learning · Computer Science 2017-03-31 Joseph Gomes , Bharath Ramsundar , Evan N. Feinberg , Vijay S. Pande

RNA binding proteins play a crucial role in post-transcriptional gene regulation by controlling the transport, processing, and translation of their target RNAs. Post-transcriptional gene regulation leads to the differential expression of…

Biomolecules · Quantitative Biology 2026-03-24 Danielle Wampler , Ralf Bundschuh

Protein language models (pLMs) pre-trained on vast protein sequence databases excel at various downstream tasks but often lack the structural knowledge essential for some biological applications. To address this, we introduce a method to…

Recent work on extending coreference resolution across domains and languages relies on annotated data in both the target domain and language. At the same time, pre-trained large language models (LMs) have been reported to exhibit strong…

Computation and Language · Computer Science 2023-11-16 Nghia T. Le , Alan Ritter

Fine-tuning large-scale pretrained models has led to tremendous progress in well-studied modalities such as vision and NLP. However, similar gains have not been observed in many other modalities due to a lack of relevant pretrained models.…

Machine Learning · Computer Science 2023-03-21 Junhong Shen , Liam Li , Lucio M. Dery , Corey Staten , Mikhail Khodak , Graham Neubig , Ameet Talwalkar

Protein-protein interactions (PPIs) are fundamental to numerous cellular processes, and their characterization is vital for understanding disease mechanisms and guiding drug discovery. While protein language models (PLMs) have demonstrated…

Understanding how protein mutations affect protein-nucleic acid binding is critical for unraveling disease mechanisms and advancing therapies. Current experimental approaches are laborious, and computational methods remain limited in…

Quantitative Methods · Quantitative Biology 2025-05-30 Xiang Liu , Junjie Wee , Guo-Wei Wei

Protein language models have shown remarkable success in learning biological information from protein sequences. However, most existing models are limited by either autoencoding or autoregressive pre-training objectives, which makes them…

Quantitative Methods · Quantitative Biology 2024-12-10 Bo Chen , Xingyi Cheng , Pan Li , Yangli-ao Geng , Jing Gong , Shen Li , Zhilei Bei , Xu Tan , Boyan Wang , Xin Zeng , Chiming Liu , Aohan Zeng , Yuxiao Dong , Jie Tang , Le Song

Genomic language models (gLMs) face a fundamental efficiency challenge: either maintain separate specialized models for each biological modality (DNA and RNA) or develop large multi-modal architectures. Both approaches impose significant…

Genomics · Quantitative Biology 2025-08-08 Shiyi Du , Litian Liang , Jiayi Li , Carl Kingsford

Protein-ligand binding affinity (PLBA) prediction is the fundamental task in drug discovery. Recently, various deep learning-based models predict binding affinity by incorporating the three-dimensional structure of protein-ligand complexes…

Biomolecules · Quantitative Biology 2023-12-21 Jiaxian Yan , Zhaofeng Ye , Ziyi Yang , Chengqiang Lu , Shengyu Zhang , Qi Liu , Jiezhong Qiu