中文
相关论文

相关论文: Copy-Augmented Representation for Structure Invari…

200 篇论文

Retrosynthesis planning remains a central challenge in molecular discovery due to the vast and complex chemical reaction space. While traditional template-based methods offer tractability, they suffer from poor scalability and limited…

机器学习 · 计算机科学 2025-07-31 Nguyen Xuan-Vu , Daniel P Armstrong , Zlatko Jončev , Philippe Schwaller

Recent years have seen rapid development of descriptor generation based on representation learning of extremely diverse molecules, especially those that apply natural language processing (NLP) models to SMILES, a literal representation of…

机器学习 · 计算机科学 2024-02-20 Yasuhiro Yoshikai , Tadahaya Mizuno , Shumpei Nemoto , Hiroyuki Kusuhara

Retrosynthesis -- the process of identifying a set of reactants to synthesize a target molecule -- is of vital importance to material design and drug discovery. Existing machine learning approaches based on language models and graph neural…

化学物理 · 物理学 2021-12-10 Ruoxi Sun , Hanjun Dai , Li Li , Steven Kearnes , Bo Dai

With the recent advances in machine learning for quantum chemistry, it is now possible to predict the chemical properties of compounds and to generate novel molecules. Existing generative models mostly use a string- or graph-based…

生物大分子 · 定量生物学 2020-10-14 Vitali Nesterov , Mario Wieser , Volker Roth

There is more and more evidence that machine learning can be successfully applied in materials science and related fields. However, datasets in these fields are often quite small ($\ll1000$ samples). It makes the most advanced machine…

计算物理 · 物理学 2022-02-25 Guillaume Lambard , Ekaterina Gracheva

In drug discovery, predicting the absorption, distribution, metabolism, excretion, and toxicity (ADMET) properties of small-molecule drugs is critical for ensuring safety and efficacy. However, the process of accurately predicting these…

机器学习 · 计算机科学 2026-03-27 Bohao Xu , Yingzhou Lu , Chenhao Li , Ling Yue , Xiao Wang , Tianfan Fu , Minjie Shen , Lulu Chen

Chemical databases store information in text representations, and the SMILES format is a universal standard used in many cheminformatics software. Encoded in each SMILES string is structural information that can be used to predict complex…

机器学习 · 统计学 2018-08-16 Garrett B. Goh , Nathan O. Hodas , Charles Siegel , Abhinav Vishnu

We introduce Group SELFIES, a molecular string representation that leverages group tokens to represent functional groups or entire substructures while maintaining chemical robustness guarantees. Molecular string representations, such as…

机器学习 · 计算机科学 2023-10-19 Austin Cheng , Andy Cai , Santiago Miret , Gustavo Malkomes , Mariano Phielipp , Alán Aspuru-Guzik

Molecule generation is key to drug discovery and materials science, enabling the design of novel compounds with specific properties. Large language models (LLMs) can learn to perform a wide range of tasks from just a few examples. However,…

计算与语言 · 计算机科学 2025-09-30 Wen Tao , Jing Tang , Alvin Chan , Bryan Hooi , Baolong Bi , Nanyun Peng , Yuansheng Liu , Yiwei Wang

AI-based computer-aided synthesis planning (CASP) systems are in demand as components of AI-driven drug discovery workflows. However, the high latency of such CASP systems limits their utility for high-throughput synthesizability screening…

机器学习 · 计算机科学 2025-08-05 Mikhail Andronov , Natalia Andronova , Michael Wand , Jürgen Schmidhuber , Djork-Arné Clevert

Retrosynthesis is a major task for drug discovery. It is formulated as a graph-generating problem by many existing approaches. Specifically, these methods firstly identify the reaction center, and break target molecule accordingly to…

机器学习 · 计算机科学 2022-09-28 Jiahan Liu , Chaochao Yan , Yang Yu , Chan Lu , Junzhou Huang , Le Ou-Yang , Peilin Zhao

Accurate molecular property prediction (MPP) is a critical step in modern drug development. However, the scarcity of experimental validation data poses a significant challenge to AI-driven research paradigms. Under few-shot learning…

机器学习 · 计算机科学 2025-05-20 Yifan Dai , Xuanbai Ren , Tengfei Ma , Qipeng Yan , Yiping Liu , Yuansheng Liu , Xiangxiang Zeng

Language models demonstrate fundamental abilities in syntax, semantics, and reasoning, though their performance often depends significantly on the inputs they process. This study introduces TSIS (Simplified TSID) and its variants:TSISD…

人工智能 · 计算机科学 2024-11-19 Juan-Ni Wu , Tong Wang , Li-Juan Tang , Hai-Long Wu , Ru-Qin Yu

The detailed analysis of molecular structures and properties holds great potential for drug development discovery through machine learning. Developing an emergent property in the model to understand molecules would broaden the horizons for…

Functional groups and moieties are chemical descriptors of biomolecules that can be used to interpret their properties and functions, leading to the understanding of chemical or biological mechanisms. These chemical building blocks, or…

生物大分子 · 定量生物学 2021-11-08 Yasemin Yesiltepe , Ryan S. Renslow , Thomas O. Metz

Molecule representation learning (MRL) methods aim to embed molecules into a real vector space. However, existing SMILES-based (Simplified Molecular-Input Line-Entry System) or GNN-based (Graph Neural Networks) MRL methods either take…

机器学习 · 计算机科学 2021-09-23 Hongwei Wang , Weijiang Li , Xiaomeng Jin , Kyunghyun Cho , Heng Ji , Jiawei Han , Martin D. Burke

Most molecular diagram parsers recover chemical structure from raster images (e.g., PNGs). However, many PDFs include commands giving explicit locations and shapes for characters, lines, and polygons. We present a new parser that uses these…

计算机视觉与模式识别 · 计算机科学 2025-02-27 Ayush Kumar Shah , Bryan Manrique Amador , Abhisek Dey , Ming Creekmore , Blake Ocampo , Scott Denmark , Richard Zanibbi

Template matching is a fundamental task in computer vision and has been studied for decades. It plays an essential role in manufacturing industry for estimating the poses of different parts, facilitating downstream tasks such as robotic…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Zhirui Gao , Renjiao Yi , Zheng Qin , Yunfan Ye , Chenyang Zhu , Kai Xu

Large-scale pre-training methodologies for chemical language models represent a breakthrough in cheminformatics. These methods excel in tasks such as property prediction and molecule generation by learning contextualized representations of…

Molecule generation is a task made very difficult by the complex ways in which we represent molecules computationally. A common technique used in molecular generative modeling is to use SMILES strings with recurrent neural networks built…

生物大分子 · 定量生物学 2024-02-28 Divahar Sivanesan