English
Related papers

Related papers: RITA: a Study on Scaling Up Generative Protein Seq…

200 papers

Biological machine learning is often bottlenecked by a lack of scaled data. One promising route to relieving data bottlenecks is through high throughput screens, which can experimentally test the activity of $10^6-10^{12}$ protein sequences…

Machine Learning · Statistics 2025-10-21 Eli N. Weinstein , Andrei Slabodkin , Mattia G. Gollub , Elizabeth B. Wood

Generalization beyond training data remains a central challenge in machine learning for biology. A common way to enhance generalization is self-supervised pre-training on large datasets. However, aiming to perform well on all possible…

We propose Gradient Inversion Transcript (GIT), a novel generative approach for reconstructing training data from leaked gradients. GIT employs a generative attack model, whose architecture is tailored to align with the structure of the…

Machine Learning · Computer Science 2025-05-27 Xinping Chen , Chen Liu

Proteins are macromolecules that mediate a significant fraction of the cellular processes that underlie life. An important task in bioengineering is designing proteins with specific 3D structures and chemical properties which enable…

Quantitative Methods · Quantitative Biology 2022-05-31 Namrata Anand , Tudor Achim

Despite the impressive capabilities of large language models across various tasks, their continued scaling is severely hampered not only by data scarcity but also by the performance degradation associated with excessive data repetition…

Computation and Language · Computer Science 2025-05-20 Xintong Hao , Ruijie Zhu , Ge Zhang , Ke Shen , Chenggang Li

We analyse a simple discrete-time stochastic process for the theoretical modeling of the evolution of protein lengths. At every step of the process a new protein is produced as a modification of one of the proteins already existing and its…

Populations and Evolution · Quantitative Biology 2009-11-13 C. Destri , C. Miccio

We present the MSA-to-protein transformer, a generative model of protein sequences conditioned on protein families represented by multiple sequence alignments (MSAs). Unlike existing approaches to learning generative models of protein…

Biomolecules · Quantitative Biology 2022-04-05 Soumya Ram , Tristan Bepler

Designing protein sequences with desired biological function is crucial in biology and chemistry. Recent machine learning methods use a surrogate sequence-function model to replace the expensive wet-lab validation. How can we efficiently…

Biomolecules · Quantitative Biology 2024-12-04 Zhenqiao Song , Lei Li

Protein generative models have shown remarkable promise in protein design, yet their success rates remain constrained by reliance on curated sequence-structure datasets and by misalignment between supervised objectives and real design…

Machine Learning · Computer Science 2026-03-03 Ziwen Wang , Jiajun Fan , Ruihan Guo , Thao Nguyen , Heng Ji , Ge Liu

Protein design has become a critical method in advancing significant potential for various applications such as drug development and enzyme engineering. However, protein design methods utilizing large language models with solely pretraining…

Artificial Intelligence · Computer Science 2024-12-06 Xiao-Yu Guo , Yi-Fan Li , Yuan Liu , Xiaoyong Pan , Hong-Bin Shen

Generative machine learning models offer a powerful framework for therapeutic design by efficiently exploring large spaces of biological sequences enriched for desirable properties. Unlike supervised learning methods, which require both…

Proteins are fundamental biological entities that play a key role in life activities. The amino acid sequences of proteins can be folded into stable 3D structures in the real physicochemical world, forming a special kind of…

Machine Learning · Computer Science 2023-01-04 Lirong Wu , Yufei Huang , Haitao Lin , Stan Z. Li

SE(3)-based generative models have shown great promise in protein geometry modeling and effective structure design. However, the field currently lacks a modularized benchmark to enable comprehensive investigation and fair comparison of…

Machine Learning · Computer Science 2025-07-29 Lang Yu , Zhangyang Gao , Cheng Tan , Qin Chen , Jie Zhou , Liang He

Accurately measuring protein-RNA binding affinity is crucial in many biological processes and drug design. Previous computational methods for protein-RNA binding affinity prediction rely on either sequence or structure features, unable to…

Biomolecules · Quantitative Biology 2025-01-06 Rong Han , Xiaohong Liu , Tong Pan , Jing Xu , Xiaoyu Wang , Wuyang Lan , Zhenyu Li , Zixuan Wang , Jiangning Song , Guangyu Wang , Ting Chen

In this work, we introduce Reinforcement Pre-Training (RPT) as a new scaling paradigm for large language models and reinforcement learning (RL). Specifically, we reframe next-token prediction as a reasoning task trained using RL, where it…

Computation and Language · Computer Science 2025-06-10 Qingxiu Dong , Li Dong , Yao Tang , Tianzhu Ye , Yutao Sun , Zhifang Sui , Furu Wei

Protein is linked to almost every life process. Therefore, analyzing the biological structure and property of protein sequences is critical to the exploration of life, as well as disease detection and drug discovery. Traditional protein…

Machine Learning · Computer Science 2021-12-08 Yijia Xiao , Jiezhong Qiu , Ziang Li , Chang-Yu Hsieh , Jie Tang

With the growth of the academic engines, the mining and analysis acquisition of massive researcher data, such as collaborator recommendation and researcher retrieval, has become indispensable. It can improve the quality of services and…

Information Retrieval · Computer Science 2022-03-02 Ziyue Qiao , Yanjie Fu , Pengyang Wang , Meng Xiao , Zhiyuan Ning , Denghui Zhang , Yi Du , Yuanchun Zhou

Large Language Models (LLMs) have achieved remarkable success through imitation learning on vast text corpora, but this paradigm creates a training-generation gap and limits robust reasoning. Reinforcement learning (RL) offers a more…

Computation and Language · Computer Science 2026-04-13 Zhepeng Cen , Haolin Chen , Shiyu Wang , Zuxin Liu , Zhiwei Liu , Jielin Qiu , Ding Zhao , Silvio Savarese , Caiming Xiong , Huan Wang , Weiran Yao

Test-time augmentation (TTA) has become a promising approach for mitigating data sparsity in sequential recommendation by improving inference accuracy without requiring costly model retraining. However, existing TTA methods typically rely…

Information Retrieval · Computer Science 2026-04-20 Xibo Li , Liang Zhang

The rapid advancement of DNA sequencing has produced vast genomic datasets, yet interpreting and engineering genomic function remain fundamental challenges. Recent large language models have opened new avenues for genomic analysis, but…