中文
相关论文

相关论文: Isotropy and Geometry of Pretrained Protein LMs

200 篇论文

Protein language models (pLMs) produce per-residue representations that capture evolutionary and structural information, yet their mean-pooled sequence embeddings are not explicitly trained to reflect functional, evolutionary or structural…

机器学习 · 计算机科学 2026-05-11 Dan Ofer , Oriel Perets , Michal Linial , Nadav Rappoport

Fine-tuning pre-trained language models (PTLMs), such as BERT and its better variant RoBERTa, has been a common practice for advancing performance in natural language understanding (NLU) tasks. Recent advance in representation learning…

计算与语言 · 计算机科学 2021-02-05 Wenxuan Zhou , Bill Yuchen Lin , Xiang Ren

Computational biology and bioinformatics provide vast data gold-mines from protein sequences, ideal for Language Models taken from NLP. These LMs reach for new prediction frontiers at low inference costs. Here, we trained two…

Latent representation alignment has become a foundational technique for constructing multimodal large language models (MLLM) by mapping embeddings from different modalities into a shared space, often aligned with the embedding space of…

机器学习 · 计算机科学 2025-03-06 Dong Shu , Bingbing Duan , Kai Guo , Kaixiong Zhou , Jiliang Tang , Mengnan Du

Pre-trained LLMs have demonstrated substantial capabilities across a range of conventional natural language processing (NLP) tasks, such as summarization and entity recognition. In this paper, we explore the application of LLMs in the…

定量方法 · 定量生物学 2024-08-14 Kamyar Zeinalipour , Neda Jamshidi , Monica Bianchini , Marco Maggini , Marco Gori

Pre-trained language models such as BERT have become a more common choice of natural language processing (NLP) tasks. Research in word representation shows that isotropic embeddings can significantly improve performance on downstream tasks.…

计算与语言 · 计算机科学 2021-08-30 Yuxin Liang , Rui Cao , Jie Zheng , Jie Ren , Ling Gao

We propose ProtLLM, a versatile cross-modal large language model (LLM) for both protein-centric and protein-language tasks. ProtLLM features a unique dynamic protein mounting mechanism, enabling it to handle complex inputs where the natural…

生物大分子 · 定量生物学 2024-03-14 Le Zhuo , Zewen Chi , Minghao Xu , Heyan Huang , Heqi Zheng , Conghui He , Xian-Ling Mao , Wentao Zhang

The parallels between protein sequences and natural language in their sequential structures have inspired the application of large language models (LLMs) to protein understanding. Despite the success of LLMs in NLP, their effectiveness in…

定量方法 · 定量生物学 2024-07-09 Yiqing Shen , Zan Chen , Michail Mamalakis , Luhan He , Haiyang Xia , Tianbin Li , Yanzhou Su , Junjun He , Yu Guang Wang

Several studies have explored various advantages of multilingual pre-trained models (such as multilingual BERT) in capturing shared linguistic knowledge. However, less attention has been paid to their limitations. In this paper, we…

计算与语言 · 计算机科学 2022-03-18 Sara Rajaee , Mohammad Taher Pilehvar

Proteins adopt multiple structural conformations to perform their diverse biological functions, and understanding these conformations is crucial for advancing drug discovery. Traditional physics-based simulation methods often struggle with…

生物大分子 · 定量生物学 2025-03-14 Jiarui Lu , Xiaoyin Chen , Stephen Zhewen Lu , Chence Shi , Hongyu Guo , Yoshua Bengio , Jian Tang

Protein research is crucial in various fundamental disciplines, but understanding their intricate structure-function relationships remains challenging. Recent Large Language Models (LLMs) have made significant strides in comprehending…

计算工程、金融与科学 · 计算机科学 2025-01-24 Chao Wang , Hehe Fan , Ruijie Quan , Yi Yang

Recent advances in protein language models (PLMs) have demonstrated remarkable capabilities in understanding protein sequences. However, the extent to which different model architectures capture antibody-specific biological properties…

机器学习 · 计算机科学 2025-12-11 Mengren , Liu , Yixiang Zhang , Yiming , Zhang

Understanding biological processes, drug development, and biotechnological advancements requires a detailed analysis of protein structures and functions, a task that is inherently complex and time-consuming in traditional protein research.…

人工智能 · 计算机科学 2025-04-21 Yijia Xiao , Edward Sun , Yiqiao Jin , Qifan Wang , Wei Wang

Recent advances in protein large language models, such as ProtTeX, represent both side-chain amino acids and backbone structure as discrete token sequences of residue length. While this design enables unified modeling of multimodal protein…

机器学习 · 计算机科学 2025-08-19 Chuanliu Fan , Zicheng Ma , Jun Gao , Nan Yu , Jun Zhang , Ziqiang Cao , Yi Qin Gao , Guohong Fu

The representation space of pretrained Language Models (LMs) encodes rich information about words and their relationships (e.g., similarity, hypernymy, polysemy) as well as abstract semantic notions (e.g., intensity). In this paper, we…

计算与语言 · 计算机科学 2023-06-02 Qing Lyu , Marianna Apidianaki , Chris Callison-Burch

Large language models (LLMs) demonstrate impressive results in natural language processing tasks but require a significant amount of computational and memory resources. Structured matrix representations are a promising way for reducing the…

计算与语言 · 计算机科学 2025-06-04 Ekaterina Grishina , Mikhail Gorbunov , Maxim Rakhuba

Masked Language Modeling (MLM) is widely used to pretrain language models. The standard random masking strategy in MLM causes the pre-trained language models (PLMs) to be biased toward high-frequency tokens. Representation learning of rare…

计算与语言 · 计算机科学 2023-05-25 Linhan Zhang , Qian Chen , Wen Wang , Chong Deng , Xin Cao , Kongzhang Hao , Yuxin Jiang , Wei Wang

Geometric deep learning has recently achieved great success in non-Euclidean domains, and learning on 3D structures of large biomolecules is emerging as a distinct research area. However, its efficacy is largely constrained due to the…

机器学习 · 计算机科学 2023-10-31 Fang Wu , Lirong Wu , Dragomir Radev , Jinbo Xu , Stan Z. Li

It is widely accepted that fine-tuning pre-trained language models usually brings about performance improvements in downstream tasks. However, there are limited studies on the reasons behind this effectiveness, particularly from the…

计算与语言 · 计算机科学 2021-09-13 Sara Rajaee , Mohammad Taher Pilehvar

Protein language models (pLMs) pre-trained on vast protein sequence databases excel at various downstream tasks but often lack the structural knowledge essential for some biological applications. To address this, we introduce a method to…

‹ 上一页 1 2 3 10 下一页 ›