中文
相关论文

相关论文: PepTriX: A Framework for Explainable Peptide Analy…

200 篇论文

Understanding the binding specificity between T-cell receptors (TCRs) and peptide-major histocompatibility complexes (pMHCs) is central to immunotherapy and vaccine development. However, current predictive models struggle with…

定量方法 · 定量生物学 2025-12-29 Cong Qi , Hanzhang Fang , Siqi jiang , Tianxing Hu , Zhi Wei

This paper demonstrates that language models are strong structure-based protein designers. We present LM-Design, a generic approach to reprogramming sequence-based protein language models (pLMs), that have learned massive sequential…

机器学习 · 计算机科学 2023-02-10 Zaixiang Zheng , Yifan Deng , Dongyu Xue , Yi Zhou , Fei YE , Quanquan Gu

Molecular property prediction is a crucial foundation for drug discovery. In recent years, pre-trained deep learning models have been widely applied to this task. Some approaches that incorporate prior biological domain knowledge into the…

机器学习 · 计算机科学 2024-08-20 Tianyu Zhang , Yuxiang Ren , Chengbin Hou , Hairong Lv , Xuegong Zhang

Understanding the internal representations of large language models (LLMs) can help explain models' behavior and verify their alignment with human values. Given the capabilities of LLMs in generating human-understandable text, we propose…

计算与语言 · 计算机科学 2024-06-10 Asma Ghandeharioun , Avi Caciularu , Adam Pearce , Lucas Dixon , Mor Geva

Latent representation alignment has become a foundational technique for constructing multimodal large language models (MLLM) by mapping embeddings from different modalities into a shared space, often aligned with the embedding space of…

机器学习 · 计算机科学 2025-03-06 Dong Shu , Bingbing Duan , Kai Guo , Kaixiong Zhou , Jiliang Tang , Mengnan Du

Representation learning for protein biochemical space faces a difficult trade-off: protein language models excel at capturing long-range biological semantics but often miss fine-grained chemical details. Conversely, chemical language models…

In research areas with scarce data, representation learning plays a significant role. This work aims to enhance representation learning for clinical time series by deriving universal embeddings for clinical features, such as heart rate and…

机器学习 · 计算机科学 2024-02-07 Yurong Hu , Manuel Burger , Gunnar Rätsch , Rita Kuznetsova

Existing works show that augmenting the training data of pre-trained language models (PLMs) for classification tasks fine-tuned via parameter-efficient fine-tuning methods (PEFT) using both clean and adversarial examples can enhance their…

计算与语言 · 计算机科学 2024-06-18 Tuc Nguyen , Thai Le

Proteins are biomolecules of life. They fold into a great variety of three-dimensional (3D) shapes. Underlying these folding patterns are many recurrent structural fragments or building blocks (analogous to `LEGO bricks'). This paper…

定量方法 · 定量生物学 2013-10-08 Arun S. Konagurthu , Arthur M. Lesk , David Abramson , Peter J. Stuckey , Lloyd Allison

Among these, D-peptides are resistant to proteolysis, exhibit greater in vivo stability, and are easier to synthesize. Despite advances in deep learning for peptide discovery, the scarcity of natural D-protein data limits the transfer of…

计算工程、金融与科学 · 计算机科学 2026-05-04 Fang Wu , Shuting Jin , Xiangru Tang , Junlin Xu , Mark Gerstein , James Zou

Proteins are essential biological macromolecules that execute life functions. Local structural motifs, such as active sites, are the most critical components for linking structure to function and are key to understanding protein evolution…

定量方法 · 定量生物学 2026-04-10 Zhiyu Wang , Bingxin Zhou , Jing Wang , Yang Tan , Weishu Zhao , Pietro Liò , Liang Hong

Understanding how small molecules perturb gene expression is essential for uncovering drug mechanisms, predicting off-target effects, and identifying repurposing opportunities. While prior deep learning frameworks have integrated multimodal…

机器学习 · 计算机科学 2026-01-01 Pascal Passigan , Kevin Zhu , Angelina Ning

Pre-trained Large Language Models (LLMs) often struggle on out-of-domain datasets like healthcare focused text. We explore specialized pre-training to adapt smaller LLMs to different healthcare datasets. Three methods are assessed:…

计算与语言 · 计算机科学 2024-04-01 Niall Taylor , Dan Schofield , Andrey Kormilitzin , Dan W Joyce , Alejo Nevado-Holgado

Existing spectral benchmarks are limited in scale, modality alignment, and evaluation scope, and typically focus on either specialized models or multimodal language models (MLLMs). We introduce SpecX, a large-scale benchmark for multi-modal…

图像与视频处理 · 电气工程与系统科学 2026-05-20 Chengrui Xiang , Tengfei Ma , Yujie Chen , Tong Wang , Haowen Chen , Xiangxiang Zeng

With the exponential increase of the protein sequence databases over time, multiple-sequence alignment (MSA) methods, like PSI-BLAST, perform exhaustive and time-consuming database search to retrieve evolutionary information. The resulting…

定量方法 · 定量生物学 2023-08-21 Issar Arab

Drug repurposing is often framed as a candidate identification task, but existing approaches provide limited guidance for distinguishing biologically plausible candidates from historically well-connected ones. Here we introduce DrugKLM, a…

Protein language models (PLMs) learn contextual representations from protein sequences and are profoundly impacting various scientific disciplines spanning protein design, drug discovery, and structural predictions. One particular research…

定量方法 · 定量生物学 2024-02-07 Andreas Dounas , Tudor-Stefan Cotet , Alexander Yermanos

Fine-tuning of Large Language Models (LLMs) has become the default practice for improving model performance on a given task. However, performance improvement comes at the cost of training on vast amounts of annotated data which could be…

计算与语言 · 计算机科学 2025-04-25 Jose G. Moreno , Jesus Lovon , M'Rick Robin-Charlet , Christine Damase-Michel , Lynda Tamine

Large Language Models (LLMs) possess strong representation and reasoning capabilities, but their application to structure-based drug design (SBDD) is limited by insufficient understanding of protein structures and unpredictable molecular…

机器学习 · 计算机科学 2026-01-27 Xuanning Hu , Anchen Li , Qianli Xing , Jinglong Ji , Hao Tuo , Bo Yang

Motivation: Peptides have attracted the attention in this century due to their remarkable therapeutic properties. Computational tools are being developed to take advantage of existing information, encapsulating knowledge and making it…