中文
相关论文

相关论文: PepMLM: Target Sequence-Conditioned Generation of …

200 篇论文

Background: The inception of next generations sequencing technologies have exponentially increased the volume of biological sequence data. Protein sequences, being quoted as the `language of life', has been analyzed for a multitude of…

定量方法 · 定量生物学 2020-12-08 Nabil Ibtehaz , S. M. Shakhawat Hossain Sourav , Md. Shamsuzzoha Bayzid , M. Sohel Rahman

In recent years, protein-text models have gained significant attention for their potential in protein generation and understanding. Current approaches focus on integrating protein-related knowledge into large language models through…

计算与语言 · 计算机科学 2025-11-11 Juntong Wu , Zijing Liu , He Cao , Hao Li , Bin Feng , Zishan Shu , Ke Yu , Li Yuan , Yu Li

Target-specific peptides, such as conotoxins, exhibit exceptional binding affinity and selectivity toward ion channels and receptors. However, their therapeutic potential remains underutilized due to the limited diversity of natural…

生物大分子 · 定量生物学 2025-05-07 Cheng Ge , Han-Shen Tae , Zhenqiang Zhang , Lu Lu , Zhijie Huang , Yilin Wang , Tao Jiang , Wenqing Cai , Shan Chang , David J. Adams , Rilei Yu

Targeted protein degradation (TPD) induced by small molecules has emerged as a rapidly evolving modality in drug discovery, targeting proteins traditionally considered "undruggable". Proteolysis-targeting chimeras (PROTACs) and molecular…

生物大分子 · 定量生物学 2025-02-27 Fanglei Xue , Meihan Zhang , Shuqi Li , Xinyu Gao , James A. Wohlschlegel , Wenbing Huang , Yi Yang , Weixian Deng

Masked language modeling (MLM) is the standard objective for training protein language models, typically implemented by randomly masking individual residues at a fixed rate (e.g., 15%). This practice implicitly assumes that all sequence…

机器学习 · 计算机科学 2026-05-19 Thomas Walton , Ayan Goel , Amirali Aghazadeh

Autoregressive models have transformed protein engineering by enabling the generation of novel protein sequences beyond those found in nature. However, their sequential inference introduces significant latency, limiting their utility in…

机器学习 · 计算机科学 2025-09-29 Thomas Walton , Darin Tsui , Aryan Musharaf , Amirali Aghazadeh

Pre-trained language models (PrLM) have to carefully manage input units when training on a very large text with a vocabulary consisting of millions of words. Previous works have shown that incorporating span-level information over…

计算与语言 · 计算机科学 2021-09-16 Rongzhou Bao , Zhuosheng Zhang , Hai Zhao

Peptide compounds demonstrate considerable potential as therapeutic agents due to their high target affinity and low toxicity, yet their drug development is constrained by their low membrane permeability. Molecular weight and peptide length…

机器学习 · 计算机科学 2025-05-26 Shuang Wu , Meijie Wang , Lun Yu

Adeno-associated viral (AAV) vectors are widely used delivery platforms in gene therapy, and the design of improved capsids is key to expanding their therapeutic potential. A central challenge in AAV bioengineering, as in protein design…

Designing protein binders targeting specific sites, which requires to generate realistic and functional interaction patterns, is a fundamental challenge in drug discovery. Current structure-based generative models are limited in generating…

机器学习 · 计算机科学 2025-10-17 Zishen Zhang , Xiangzhe Kong , Wenbing Huang , Yang Liu

Generating molecules with high binding affinities to target proteins (a.k.a. structure-based drug design) is a fundamental and challenging task in drug discovery. Recently, deep generative models have achieved remarkable success in…

生物大分子 · 定量生物学 2023-05-24 Zaixi Zhang , Qi Liu

Large Language Models (LLMs) excel at generating fluent text but struggle to enforce external constraints because they generate tokens sequentially without explicit control mechanisms. GenCP addresses this limitation by combining LLM…

计算与语言 · 计算机科学 2025-06-02 Alexandre Bonlarron , Florian Régin , Elisabetta De Maria , Jean-Charles Régin

The design of novel protein sequences with targeted functionalities underpins a central theme in protein engineering, impacting diverse fields such as drug discovery and enzymatic engineering. However, navigating this vast combinatorial…

生物大分子 · 定量生物学 2024-02-19 Yiheng Zhu , Zitai Kong , Jialu Wu , Weize Liu , Yuqiang Han , Mingze Yin , Hongxia Xu , Chang-Yu Hsieh , Tingjun Hou

Current token-sequence-based Large Language Models (LLMs) are not well-suited for directly processing 3D Boundary Representation (Brep) models that contain complex geometric and topological information. We propose BrepLLM, the first…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Liyuan Deng , Hao Guo , Yunpeng Bai , Yongkang Dai , Huaxi Huang , Yilei Shi

In recent years, the scientific community has become increasingly interested on peptides with non-canonical amino acids due to their superior stability and resistance to proteolytic degradation. These peptides present promising…

生物大分子 · 定量生物学 2023-11-09 Ruochi Zhang , Haoran Wu , Yuting Xiu , Kewei Li , Ningning Chen , Yu Wang , Yan Wang , Xin Gao , Fengfeng Zhou

Peptide therapeutics are widely regarded as the "third generation" of drugs, yet progress in peptide Machine Learning (ML) are hindered by the absence of standardized benchmarks. Here we present PepBenchmark, which unifies datasets,…

机器学习 · 计算机科学 2026-04-14 Jiahui Zhang , Rouyi Wang , Kuangqi Zhou , Tianshu Xiao , Lingyan Zhu , Yaosen Min , Yang Wang

Designing proteins with specific attributes offers an important solution to address biomedical challenges. Pre-trained protein large language models (LLMs) have shown promising results on protein sequence generation. However, to control…

人工智能 · 计算机科学 2025-01-28 Xiangyu Liu , Yi Liu , Silei Chen , Wei Hu

Efficient design and discovery of target-driven molecules is a critical step in facilitating lead optimization in drug discovery. Current approaches to develop molecules for a target protein are intuition-driven, hampered by slow iterative…

机器学习 · 计算机科学 2022-05-24 Andrew D. McNaughton , Mridula S. Bontha , Carter R. Knutson , Jenna A. Pope , Neeraj Kumar

For protein sequence datasets, unlabeled data has greatly outpaced labeled data due to the high cost of wet-lab characterization. Recent deep-learning approaches to protein prediction have shown that pre-training on unlabeled data can yield…

机器学习 · 计算机科学 2020-12-02 Pascal Sturmfels , Jesse Vig , Ali Madani , Nazneen Fatema Rajani

Prediction of ligand binding sites of proteins is a fundamental and important task for understanding the function of proteins and screening potential drugs. Most existing methods require experimentally determined protein holo-structures as…

定量方法 · 定量生物学 2023-12-07 Shuo Zhang , Lei Xie