中文
相关论文

相关论文: Scaling Up ESM2 Architectures for Long Protein Seq…

200 篇论文

Predicting the structure of interacting chains is crucial for understanding biological systems and developing new drugs. Large-scale pre-trained Protein Language Models (PLMs), such as ESM2, have shown impressive abilities in extracting…

生物大分子 · 定量生物学 2023-12-05 Shuxian Zou , Hui Li , Shentong Mo , Xingyi Cheng , Eric Xing , Le Song

Existing Protein Language Models (PLMs) often suffer from limited adaptability to multiple tasks and exhibit poor generalization across diverse biological contexts. In contrast, general-purpose Large Language Models (LLMs) lack the…

机器学习 · 计算机科学 2026-02-23 Yujia Wang , Jihong Guan , Wengen Li , Shuigeng Zhou , Xuhong Wang

Considering the significance of proteins, computational protein science has always been a critical scientific field, dedicated to revealing knowledge and developing applications within the protein sequence-structure-function paradigm. In…

计算工程、金融与科学 · 计算机科学 2025-01-28 Wenqi Fan , Yi Zhou , Shijie Wang , Yuyao Yan , Hui Liu , Qian Zhao , Le Song , Qing Li

Machine learning is widely used to analyze biological sequence data. Non-sequential models such as SVMs or feed-forward neural networks are often used although they have no natural way of handling sequences of varying length. Recurrent…

定量方法 · 定量生物学 2016-03-14 Søren Kaae Sønderby , Casper Kaae Sønderby , Henrik Nielsen , Ole Winther

Research on long non-coding RNAs (lncRNAs) has garnered significant attention due to their critical roles in gene regulation and disease mechanisms. However, the complexity and diversity of lncRNA sequences, along with the limited knowledge…

基因组学 · 定量生物学 2024-11-07 Wei Wang , Zhichao Hou , Xiaorui Liu , Xinxia Peng

Text generating capabilities have undergone a substantial transformation with the introduction of large language models (LLMs). Electroencephalography (EEG)-based text production is still difficult, though, because it requires a lot of data…

人机交互 · 计算机科学 2025-11-18 Khushiyant

Learning effective protein representations is critical in a variety of tasks in biology such as predicting protein functions. Recent sequence representation learning methods based on Protein Language Models (PLMs) excel in sequence-based…

定量方法 · 定量生物学 2023-10-19 Zuobai Zhang , Chuanrui Wang , Minghao Xu , Vijil Chenthamarakshan , Aurélie Lozano , Payel Das , Jian Tang

A major limitation for the broader scope of problems solvable by transformers is the quadratic scaling of computational complexity with input size. In this study, we investigate the recurrent memory augmentation of pre-trained transformer…

计算与语言 · 计算机科学 2024-02-07 Aydar Bulatov , Yuri Kuratov , Yermek Kapushev , Mikhail S. Burtsev

Large Language Models (LLMs) are revolutionizing bioinformatics, enabling advanced analysis of DNA, RNA, proteins, and single-cell data. This survey provides a systematic review of recent advancements, focusing on genomic sequence modeling,…

计算与语言 · 计算机科学 2026-03-03 Zhenyu Wang , Zikang Wang , Jiyue Jiang , Pengan Chen , Xiangyu Shi , Yu Li

Large language models (LLMs) have demonstrated remarkable capabilities across diverse domains, but their heavy resource demands make quantization-reducing precision to lower-bit formats-critical for efficient serving. While many…

性能 · 计算机科学 2025-08-26 Tianyao Shi , Yi Ding

Proteins perform much of the work in living organisms, and consequently the development of efficient computational methods for protein representation is essential for advancing large-scale biological research. Most current approaches…

定量方法 · 定量生物学 2023-06-09 Francesco Ceccarelli , Lorenzo Giusti , Sean B. Holden , Pietro Liò

Protein language models (PLMs) have emerged as powerful tools to detect complex patterns of protein sequences. However, the capability of PLMs to fully capture information on protein sequences might be limited by focusing on single…

机器学习 · 计算机科学 2025-05-27 Hazem Alsamkary , Mohamed Elshaffei , Mohamed Elkerdawy , Ahmed Elnaggar

This paper investigates the application of the transformer architecture in protein folding, as exemplified by DeepMind's AlphaFold project, and its implications for the understanding of so-called large language models. The prevailing…

计算机与社会 · 计算机科学 2024-12-10 Fabian Offert , Paul Kim , Qiaoyu Cai

Large language models have made remarkable progress in the field of molecular science, particularly in understanding and generating functional small molecules. This success is largely attributed to the effectiveness of molecular…

生物大分子 · 定量生物学 2025-03-14 Zicheng Ma , Chuanliu Fan , Zhicong Wang , Zhenyu Chen , Xiaohan Lin , Yanheng Li , Shihao Feng , Jun Zhang , Ziqiang Cao , Yi Qin Gao

Robust and effective scaling of models from small to large width typically requires the precise adjustment of many algorithmic and architectural details, such as parameterization and optimizer choices. In this work, we propose a new…

Deep learning has contributed to major advances in the prediction of protein structure from sequence, a fundamental problem in structural bioinformatics. With predictions now approaching the accuracy of crystallographic resolution in some…

定量方法 · 定量生物学 2022-01-26 Mu Gao , Mark Coletti , Russell B. Davidson , Ryan Prout , Subil Abraham , Benjamin Hernandez , Ada Sedova

Resolving and rewriting references is fundamental in programming languages. Motivated by a real-world decompilation task, we abstract reference rewriting into the problems of direct and indirect indexing by permutation. We create synthetic…

机器学习 · 计算机科学 2026-04-16 Gergő Szalay , Gergely Zsolt Kovács , Sándor Teleki , Balázs Pintér , Tibor Gregorics

Transformer-based sequence-to-sequence architectures, while achieving state-of-the-art results on a large number of NLP tasks, can still suffer from overfitting during training. In practice, this is usually countered either by applying…

计算与语言 · 计算机科学 2022-01-04 Dušan Variš , Ondřej Bojar

Protein language models have excelled in a variety of tasks, ranging from structure prediction to protein engineering. However, proteins are highly diverse in functions and structures, and current state-of-the-art models including the…

生物大分子 · 定量生物学 2023-02-27 Chang Ma , Haiteng Zhao , Lin Zheng , Jiayi Xin , Qintong Li , Lijun Wu , Zhihong Deng , Yang Lu , Qi Liu , Lingpeng Kong

Recent advancements in Large Language Models (LLMs)-based text embedding models primarily focus on data scaling or synthesis, yet limited exploration of training techniques and data quality, thereby constraining performance. In this work,…