English
Related papers

Related papers: Prot2Token: A Unified Framework for Protein Modeli…

200 papers

Self-supervised learning has been actively studied in time series domain recently, especially for masked reconstruction. Most of these methods follow the "Pre-training + Fine-tuning" paradigm in which a new decoder replaces the pre-trained…

Machine Learning · Computer Science 2023-11-08 Hao Liu , Jinrui Gan , Xiaoxuan Fan , Yi Zhang , Chuanxian Luo , Jing Zhang , Guangxin Jiang , Yucheng Qian , Changwei Zhao , Huan Ma , Zhenyu Guo

Federated Learning (FL) enables collaborative training of Large Language Models (LLMs) across distributed data sources while preserving privacy. However, when federated LLMs are deployed in critical applications, it remains unclear which…

Machine Learning · Computer Science 2026-01-29 Waris Gill , Ahmad Humayun , Ali Anwar , Muhammad Ali Gulzar

Proteins are essential macromolecules defined by their amino acid sequences, which determine their three-dimensional structures and, consequently, their functions in all living organisms. Therefore, generative protein modeling necessitates…

Machine Learning · Computer Science 2024-10-18 Xinyou Wang , Zaixiang Zheng , Fei Ye , Dongyu Xue , Shujian Huang , Quanquan Gu

Multi-token prediction (MTP) is a recently proposed pre-training objective for language models. Rather than predicting only the next token (NTP), MTP predicts the next $k$ tokens at each prediction step, using multiple prediction heads. MTP…

Computation and Language · Computer Science 2025-05-30 Ansar Aynetdinov , Alan Akbik

Protein function prediction is currently achieved by encoding its sequence or structure, where the sequence-to-function transcendence and high-quality structural data scarcity lead to obvious performance bottlenecks. Protein domains are…

Biomolecules · Quantitative Biology 2024-12-03 Mingqing Wang , Zhiwei Nie , Yonghong He , Athanasios V. Vasilakos , Zhixiang Ren

We introduce BigBang-Proton, a unified sequence-based architecture for auto-regressive language modeling pretrained on cross-scale, cross-structure, cross-discipline real-world scientific tasks to construct a scientific multi-task learner.…

Understanding and leveraging the 3D structures of proteins is central to a variety of biological and drug discovery tasks. While deep learning has been applied successfully for structure-based protein function prediction tasks, current…

Machine Learning · Computer Science 2024-04-03 Rong Han , Wenbing Huang , Lingxiao Luo , Xinyan Han , Jiaming Shen , Zhiqiang Zhang , Jun Zhou , Ting Chen

Large language models such as GPT and Llama are trained with a next-token prediction loss. In this work, we suggest that training language models to predict multiple future tokens at once results in higher sample efficiency. More…

Computation and Language · Computer Science 2026-03-03 Athul Radhakrishnan , Siddhant Mohan , Mahima Sachdeva

Document parsing, as a fundamental yet crucial vision task, is being revolutionized by vision-language models (VLMs). However, the autoregressive (AR) decoding inherent to VLMs creates a significant bottleneck, severely limiting parsing…

Computation and Language · Computer Science 2026-03-17 Lei Li , Ze Zhao , Meng Li , Zhongwang Lun , Yi Yuan , Xingjing Lu , Zheng Wei , Jiang Bian , Zang Li

Recent years have witnessed a surge in the development of protein foundation models, significantly improving performance in protein prediction and generative tasks ranging from 3D structure prediction and protein design to conformational…

Quantitative Methods · Quantitative Biology 2024-10-08 Fei Ye , Zaixiang Zheng , Dongyu Xue , Yuning Shen , Lihao Wang , Yiming Ma , Yan Wang , Xinyou Wang , Xiangxin Zhou , Quanquan Gu

Multi-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were…

Machine Learning · Computer Science 2024-12-31 Hanjing Zhou , Mingze Yin , Wei Wu , Mingyang Li , Kun Fu , Jintai Chen , Jian Wu , Zheng Wang

Code completion is one of the most useful features in the Integrated Development Environments (IDEs), which can accelerate software development by suggesting the next probable token based on the contextual code in real-time. Recent studies…

Software Engineering · Computer Science 2021-01-01 Fang Liu , Ge Li , Yunfei Zhao , Zhi Jin

Predicting protein properties, functions and localizations are important tasks in bioinformatics. Recent progress in machine learning offers an opportunities for improving existing methods. We developed a new approach called ProtBoost,…

Quantitative Methods · Quantitative Biology 2024-12-09 Alexander Chervov , Anton Vakhrushev , Sergei Fironov , Loredana Martignetti

Recently, extensive deep learning architectures and pretraining strategies have been explored to support downstream protein applications. Additionally, domain-specific models incorporating biological knowledge have been developed to enhance…

Biomolecules · Quantitative Biology 2026-03-03 Shuo Yan , Yuliang Yan , Bin Ma , Chenao Li , Haochun Tang , Jiahua Lu , Minhua Lin , Yuyuan Feng , Enyan Dai

Generative modeling for protein engineering is key to solving fundamental problems in synthetic biology, medicine, and material science. We pose protein engineering as an unsupervised sequence generation problem in order to leverage the…

Protein structure tokenization converts 3D structures into discrete or vectorized representations, enabling the integration of structural and sequence data. Despite many recent works on structure tokenization, the properties of the…

Machine Learning · Computer Science 2025-11-14 Zijing Liu , Bin Feng , He Cao , Yu Li

Understanding protein solubility is essential for their functional applications. Computational methods for predicting protein solubility are crucial for reducing experimental costs and enhancing the efficiency and success rates of protein…

Quantitative Methods · Quantitative Biology 2024-07-01 Yang Tan , Jia Zheng , Liang Hong , Bingxin Zhou

We present an approach to pose object recognition as next token prediction. The idea is to apply a language decoder that auto-regressively predicts the text tokens from image embeddings to form labels. To ground this prediction process in…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Kaiyu Yue , Bor-Chun Chen , Jonas Geiping , Hengduo Li , Tom Goldstein , Ser-Nam Lim

Consistency and reliability are crucial for conducting AI research. Many famous research fields, such as object detection, have been compared and validated with solid benchmark frameworks. After AlphaFold2, the protein folding task has…

Biomolecules · Quantitative Biology 2023-08-01 Jaemyung Lee , Kyeongtak Han , Jaehoon Kim , Hasun Yu , Youhan Lee

In recent years, there has been a surge in the development of 3D structure-based pre-trained protein models, representing a significant advancement over pre-trained protein language models in various downstream tasks. However, most existing…

Machine Learning · Computer Science 2024-06-04 Jiale Zhao , Wanru Zhuang , Jia Song , Yaqi Li , Shuqi Lu