中文
相关论文

相关论文: PROTOCOL: Late Interaction Retrieval for Protein H…

200 篇论文

Protein language models (pLMs) produce per-residue representations that capture evolutionary and structural information, yet their mean-pooled sequence embeddings are not explicitly trained to reflect functional, evolutionary or structural…

机器学习 · 计算机科学 2026-05-11 Dan Ofer , Oriel Perets , Michal Linial , Nadav Rappoport

Multimodal molecular representation learning, which jointly models molecular graphs and their textual descriptions, enhances predictive accuracy and interpretability by enabling more robust and reliable predictions of drug toxicity,…

机器学习 · 计算机科学 2025-10-21 Yingxu Wang , Kunyu Zhang , Jiaxin Huang , Nan Yin , Siwei Liu , Eran Segal

Deep learning has become a crucial tool in studying proteins. While the significance of modeling protein structure has been discussed extensively in the literature, amino acid types are typically included in the input as a default operation…

定量方法 · 定量生物学 2024-07-01 Yang Tan , Lirong Zheng , Bozitao Zhong , Liang Hong , Bingxin Zhou

Protein language models are a powerful tool for learning protein representations through pre-training on vast protein sequence datasets. However, traditional protein language models lack explicit structural supervision, despite its…

生物大分子 · 定量生物学 2024-02-09 Zuobai Zhang , Jiarui Lu , Vijil Chenthamarakshan , Aurélie Lozano , Payel Das , Jian Tang

Large pretrained language models have transformed natural language processing, and their adaptation to protein sequences -- viewed as strings of amino acid characters -- has advanced protein analysis. However, the distinct properties of…

其他定量生物学 · 定量生物学 2025-10-14 Sheikh Azizul Hakim , Kowshic Roy , M Saifur Rahman

Protein language models often take into consideration the alignment between a protein sequence and its textual description. However, they do not take structural information into consideration. Traditional methods treat sequence and…

机器学习 · 计算机科学 2026-03-10 Aditya Ranganath , Hasin Us Sami , Kowshik Thopalli , Bhavya Kailkhura , Wesam Sakla

Understanding protein sequences is vital and urgent for biology, healthcare, and medicine. Labeling approaches are expensive yet time-consuming, while the amount of unlabeled data is increasing quite faster than that of the labeled data due…

计算与语言 · 计算机科学 2021-11-01 Liang He , Shizhuo Zhang , Lijun Wu , Huanhuan Xia , Fusong Ju , He Zhang , Siyuan Liu , Yingce Xia , Jianwei Zhu , Pan Deng , Bin Shao , Tao Qin , Tie-Yan Liu

Protein language models (pLMs) have emerged as powerful predictors of protein structure and function. However, the computational circuits underlying their predictions remain poorly understood. Recent mechanistic interpretability methods…

机器学习 · 计算机科学 2026-05-14 Darin Tsui , Kunal Talreja , Daniel Saeedi , Amirali Aghazadeh

Protein retrieval, which targets the deconstruction of the relationship between sequences, structures and functions, empowers the advancing of biology. Basic Local Alignment Search Tool (BLAST), a sequence-similarity-based algorithm, has…

信息检索 · 计算机科学 2025-01-06 Yuxuan Wu , Xiao Yi , Yang Tan , Huiqun Yu , Guisheng Fan , Gaowei Zheng

The design of protein sequences with desired functionalities is a fundamental task in protein engineering. Deep generative methods, such as autoregressive models and diffusion models, have greatly accelerated the discovery of novel protein…

机器学习 · 计算机科学 2025-04-16 Zitai Kong , Yiheng Zhu , Yinlong Xu , Hanjing Zhou , Mingzhe Yin , Jialu Wu , Hongxia Xu , Chang-Yu Hsieh , Tingjun Hou , Jian Wu

The prediction of protein-protein interactions (PPIs) is crucial for understanding biological functions and diseases. Previous machine learning approaches to PPI prediction mainly focus on direct physical interactions, ignoring the broader…

生物大分子 · 定量生物学 2024-07-15 Mingyu Jin , Haochen Xue , Zhenting Wang , Boming Kang , Ruosong Ye , Kaixiong Zhou , Mengnan Du , Yongfeng Zhang

With the development of pre-trained language models, the dense retrieval models have become promising alternatives to the traditional retrieval models that rely on exact match and sparse bag-of-words representations. Different from most…

信息检索 · 计算机科学 2024-03-21 Qi Liu , Gang Guo , Jiaxin Mao , Zhicheng Dou , Ji-Rong Wen , Hao Jiang , Xinyu Zhang , Zhao Cao

Motivation: Proteins are of great significance in living organisms. However, understanding their functions encounters numerous challenges, such as insufficient integration of multimodal information, a large number of training parameters,…

机器学习 · 计算机科学 2025-05-23 Zhicong Wang , Zicheng Ma , Ziqiang Cao , Changlong Zhou , Jun Zhang , Yiqin Gao

Proteins inherently possess a consistent sequence-structure duality. The abundance of protein sequence data, which can be readily represented as discrete tokens, has driven fruitful developments in protein language models (pLMs). A key…

计算工程、金融与科学 · 计算机科学 2026-05-29 Yi Zhou , Haohao Qu , Yunqing Liu , Shanru Lin , Le Song , Wenqi Fan

We propose ProtLLM, a versatile cross-modal large language model (LLM) for both protein-centric and protein-language tasks. ProtLLM features a unique dynamic protein mounting mechanism, enabling it to handle complex inputs where the natural…

生物大分子 · 定量生物学 2024-03-14 Le Zhuo , Zewen Chi , Minghao Xu , Heyan Huang , Heqi Zheng , Conghui He , Xian-Ling Mao , Wentao Zhang

Predicting the binding affinity of protein protein complexes directly from sequence remains a challenging problem, particularly in the absence of reliable structural information. Here I present ProtT Affinity, a sequence only model that…

定量方法 · 定量生物学 2025-11-21 Hongfu Lou

Protein research is crucial in various fundamental disciplines, but understanding their intricate structure-function relationships remains challenging. Recent Large Language Models (LLMs) have made significant strides in comprehending…

计算工程、金融与科学 · 计算机科学 2025-01-24 Chao Wang , Hehe Fan , Ruijie Quan , Yi Yang

Sequence-based protein homology detection has been extensively studied and so far the most sensitive method is based upon comparison of protein sequence profiles, which are derived from multiple sequence alignment (MSA) of sequence homologs…

定量方法 · 定量生物学 2015-06-18 Jianzhu Ma , Sheng Wang , Zhiyong Wang , Jinbo Xu

Recent advances in protein large language models, such as ProtTeX, represent both side-chain amino acids and backbone structure as discrete token sequences of residue length. While this design enables unified modeling of multimodal protein…

机器学习 · 计算机科学 2025-08-19 Chuanliu Fan , Zicheng Ma , Jun Gao , Nan Yu , Jun Zhang , Ziqiang Cao , Yi Qin Gao , Guohong Fu

Vector embeddings from pre-trained language models form a core component in Neural Information Retrieval systems across a multitude of knowledge extraction tasks. The paradigm of late interaction, introduced in ColBERT, demonstrates high…

信息检索 · 计算机科学 2026-03-27 Raj Nath Patel , Sourav Dutta
‹ 上一页 1 2 3 10 下一页 ›