中文
相关论文

相关论文: PepMLM: Target Sequence-Conditioned Generation of …

200 篇论文

Generative modeling for protein engineering is key to solving fundamental problems in synthetic biology, medicine, and material science. We pose protein engineering as an unsupervised sequence generation problem in order to leverage the…

Recently, 3D generative models have shown promising performances in structure-based drug design by learning to generate ligands given target binding sites. However, only modeling the target-ligand distribution can hardly fulfill one of the…

生物大分子 · 定量生物学 2024-03-22 Xiangxin Zhou , Xiwei Cheng , Yuwei Yang , Yu Bao , Liang Wang , Quanquan Gu

The imperfect modeling of ternary complexes has limited the application of computer-aided drug discovery tools in PROTAC research and development. In this study, an AI-assisted approach for PROTAC molecule design pipeline named LM-PROTAC…

定量方法 · 定量生物学 2024-12-16 Jinsong Shao , Qineng Gong , Zeyu Yin , Yu Chen , Yajie Hao , Lei Zhang , Linlin Jiang , Min Yao , Jinlong Li , Fubo Wang , Li Wang

Language Models (LMs) excel in understanding textual descriptions of proteins, as evident in biomedical question-answering tasks. However, their capability falters with raw protein data, such as amino acid sequences, due to a deficit in…

定量方法 · 定量生物学 2024-05-22 Zhiyuan Liu , An Zhang , Hao Fei , Enzhi Zhang , Xiang Wang , Kenji Kawaguchi , Tat-Seng Chua

Masked generative models (MGMs) can generate tokens in parallel and in any order, unlike autoregressive models (ARMs), which decode one token at a time, left-to-right. However, MGMs process the full-length sequence at every sampling step,…

机器学习 · 计算机科学 2026-02-18 Justin Deschenaux , Lan Tran , Caglar Gulcehre

Understanding the relationships between protein sequence, structure and function is a long-standing biological challenge with manifold implications from drug design to our understanding of evolution. Recently, protein language models have…

定量方法 · 定量生物学 2024-01-29 Dexiong Chen , Philip Hartout , Paolo Pellizzoni , Carlos Oliver , Karsten Borgwardt

Decoding protein-protein interactions (PPIs) at the residue level is crucial for understanding cellular mechanisms and developing targeted therapeutics. We present Seq2Bind Webserver, a computational framework that leverages fine-tuned…

定量方法 · 定量生物学 2025-06-18 Xiang Ma , Supantha Dey , Vaishnavey SR , Casey Zelinski , Qi Li , Ratul Chowdhury

We propose a specialized string kernel for small bio-molecules, peptides and pseudo-sequences of binding interfaces. The kernel incorporates physico-chemical properties of amino acids and elegantly generalize eight kernels, such as the…

定量方法 · 定量生物学 2014-01-29 Sébastien Giguère , Mario Marchand , François Laviolette , Alexandre Drouin , Jacques Corbeil

Geometric deep learning has recently achieved great success in non-Euclidean domains, and learning on 3D structures of large biomolecules is emerging as a distinct research area. However, its efficacy is largely constrained due to the…

机器学习 · 计算机科学 2023-10-31 Fang Wu , Lirong Wu , Dragomir Radev , Jinbo Xu , Stan Z. Li

Recent advances in Language Models have enabled the protein modeling community with a powerful tool since protein sequences can be represented as text. Specifically, by taking advantage of Transformers, sequence-to-property prediction will…

生物大分子 · 定量生物学 2023-09-07 Chakradhar Guntuboina , Adrita Das , Parisa Mollaei , Seongwon Kim , Amir Barati Farimani

Predicting the structure of interacting chains is crucial for understanding biological systems and developing new drugs. Large-scale pre-trained Protein Language Models (PLMs), such as ESM2, have shown impressive abilities in extracting…

生物大分子 · 定量生物学 2023-12-05 Shuxian Zou , Hui Li , Shentong Mo , Xingyi Cheng , Eric Xing , Le Song

Molecular docking is a key task in computational biology that has attracted increasing interest from the machine learning community. While existing methods have achieved success, they generally treat each protein-ligand pair in isolation.…

生物大分子 · 定量生物学 2025-01-28 Jiaqi Guan , Jiahan Li , Xiangxin Zhou , Xingang Peng , Sheng Wang , Yunan Luo , Jian Peng , Jianzhu Ma

Finding drug-like compounds with high bioactivity is essential for drug discovery, but the task is complicated by the high cost of chemical synthesis and validation. With their outstanding performance in de novo drug design, deep generative…

定量方法 · 定量生物学 2023-01-03 Yibo Li , Jianfeng Pei , Luhua Lai

Pretraining DNA language models (DNALMs) on the full human genome is resource-intensive, yet often considered necessary for strong downstream performance. Inspired by recent findings in NLP and long-context modeling, we explore an…

基因组学 · 定量生物学 2025-06-24 Sohan Mupparapu , Parameswari Krishnamurthy , Ratish Puduppully

Aptamers, short synthetic RNA/DNA molecules binding specific targets with high affinity and specificity, are utilized in an increasing spectrum of bio-medical applications. Aptamers are identified in vitro via the Systematic Evolution of…

We propose to pre-train a unified language model for both autoencoding and partially autoregressive language modeling tasks using a novel training procedure, referred to as a pseudo-masked language model (PMLM). Given an input text with…

计算与语言 · 计算机科学 2020-03-02 Hangbo Bao , Li Dong , Furu Wei , Wenhui Wang , Nan Yang , Xiaodong Liu , Yu Wang , Songhao Piao , Jianfeng Gao , Ming Zhou , Hsiao-Wuen Hon

Designing protein sequences that fold into a target 3-D structure, termed as the inverse folding problem, is central to protein engineering. However, it remains challenging due to the vast sequence space and the importance of local…

定量方法 · 定量生物学 2026-03-17 Sazan Mahbub , Souvik Kundu , Eric P. Xing

Protein language models (PLMs) encode rich biological information, yet their internal neuron representations are poorly understood. We introduce the first automated framework for labeling every neuron in a PLM with biologically grounded…

机器学习 · 计算机科学 2025-07-10 Arjun Banerjee , David Martinez , Camille Dang , Ethan Tam

Traditional non-biological storage media, such as hard drives, face limitations in both storage density and lifespan due to the rapid growth of data in the big data era. Mirror-image peptides composed of D-amino acids have emerged as a…

定量方法 · 定量生物学 2026-02-09 Yilong Lu , Si Chen , Songyan Gao , Han Liu , Xin Dong , Wenfeng Shen , Guangtai Ding

Large language models (LLMs) are widely applied in various natural language processing tasks such as question answering and machine translation. However, due to the lack of labeled data and the difficulty of manual annotation for…

人工智能 · 计算机科学 2026-05-27 Xuan Lin , Long Chen , Yile Wang , Yangyang Chen , Xiangxiang Zeng
‹ 上一页 1 8 9 10 下一页 ›