中文
相关论文

相关论文: MarkushGrapher: Joint Visual and Textual Recogniti…

200 篇论文

Automatically extracting chemical structures from documents is essential for the large-scale analysis of the literature in chemistry. Automatic pipelines have been developed to recognize molecules represented either in figures or in text…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Tim Strohmeyer , Lucas Morin , Gerhard Ingmar Meijer , Valéry Weber , Ahmed Nassar , Peter Staar

The automatic analysis of chemical literature has immense potential to accelerate the discovery of new materials and drugs. Much of the critical information in patent documents and scientific articles is contained in figures, depicting the…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Lucas Morin , Martin Danelljan , Maria Isabel Agea , Ahmed Nassar , Valery Weber , Ingmar Meijer , Peter Staar , Fisher Yu

In recent decades, chemistry publications and patents have increased rapidly. A significant portion of key information is embedded in molecular structure figures, complicating large-scale literature searches and limiting the application of…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Xi Fang , Jiankun Wang , Xiaochen Cai , Shangqian Chen , Shuwen Yang , Haoyi Tao , Nan Wang , Lin Yao , Linfeng Zhang , Guolin Ke

In drug discovery, knowledge of the graph structure of chemical compounds is essential. Many thousands of scientific articles in chemistry and pharmaceutical sciences have investigated chemical compounds, but in cases the details of the…

机器学习 · 统计学 2020-09-16 Martijn Oldenhof , Adam Arany , Yves Moreau , Jaak Simm

Searching for novel molecules with desired chemical properties is crucial in drug discovery. Existing work focuses on developing neural models to generate either molecular sequences or chemical graphs. However, it remains a big challenge to…

生物大分子 · 定量生物学 2021-03-22 Yutong Xie , Chence Shi , Hao Zhou , Yuwei Yang , Weinan Zhang , Yong Yu , Lei Li

For several decades, chemical knowledge has been published in written text, and there have been many attempts to make it accessible, for example, by transforming such natural language text to a structured format. Although the discovered…

计算机视觉与模式识别 · 计算机科学 2022-02-22 Sanghyun Yoo , Ohyun Kwon , Hoshik Lee

The problem of accelerating drug discovery relies heavily on automatic tools to optimize precursor molecules to afford them with better biochemical properties. Our work in this paper substantially extends prior state-of-the-art on…

化学物理 · 物理学 2019-10-22 Wengong Jin , Regina Barzilay , Tommi Jaakkola

Materials science literature contains millions of materials synthesis procedures described in unstructured natural language text. Large-scale analysis of these synthesis procedures would facilitate deeper scientific understanding of…

Abstractive summarization of scientific papers has always been a research focus, yet existing methods face two main challenges. First, most summarization models rely on Encoder-Decoder architectures that treat papers as sequences of words,…

计算与语言 · 计算机科学 2025-05-21 Tong Bao , Heng Zhang , Chengzhi Zhang

Text images contain both visual and linguistic information. However, existing pre-training techniques for text recognition mainly focus on either visual representation learning or linguistic knowledge learning. In this paper, we propose a…

计算机视觉与模式识别 · 计算机科学 2023-10-11 Pengyuan Lyu , Chengquan Zhang , Shanshan Liu , Meina Qiao , Yangliu Xu , Liang Wu , Kun Yao , Junyu Han , Errui Ding , Jingdong Wang

Knowledge in materials science is widely dispersed across extensive scientific literature, posing significant challenges to the efficient discovery and integration of new materials. Traditional methods, often reliant on costly and…

计算与语言 · 计算机科学 2025-05-16 Yanpeng Ye , Jie Ren , Shaozhou Wang , Yuwei Wan , Imran Razzak , Bram Hoex , Haofen Wang , Tong Xie , Wenjie Zhang

Parsing chemical reaction diagrams from scientific literature is challenging due to heterogeneous layouts, intertwined visual elements, and the difficulty of integrating recognition and reasoning. Existing vision-language models advance…

人工智能 · 计算机科学 2026-05-28 Chuang Tang , Chenhao Lin , Yin Xu , Hao Wang , Jinrui Zhou , Xin Li , Mingjun Xiao , Enhong Chen

Most of the knowledge in materials science literature is in the form of unstructured data such as text and images. Here, we present a framework employing natural language processing, which automates text and image comprehension and…

数字图书馆 · 计算机科学 2021-01-06 Vineeth Venugopal , Sourav Sahoo , Mohd Zaki , Manish Agarwal , Nitya Nand Gosvami , N. M. Anoop Krishnan

Large language models record impressive performance on many natural language processing tasks. However, their knowledge capacity is limited to the pretraining corpus. Retrieval augmentation offers an effective solution by retrieving context…

计算与语言 · 计算机科学 2023-11-22 Sai Munikoti , Anurag Acharya , Sridevi Wagle , Sameera Horawalavithana

There is increasing adoption of artificial intelligence in drug discovery. However, existing studies use machine learning to mainly utilize the chemical structures of molecules but ignore the vast textual knowledge available in chemistry.…

机器学习 · 计算机科学 2024-01-31 Shengchao Liu , Weili Nie , Chengpeng Wang , Jiarui Lu , Zhuoran Qiao , Ling Liu , Jian Tang , Chaowei Xiao , Anima Anandkumar

Retrieval-augmented generation (RAG) systems have predominantly focused on text-based retrieval, limiting their effectiveness in handling visually-rich documents that encompass text, images, tables, and charts. To bridge this gap, we…

信息检索 · 计算机科学 2025-05-07 Mingjun Xu , Zehui Wang , Hengxing Cai , Renxin Zhong

Accurately identifying the synthesis conditions of metal-organic frameworks (MOFs) is essential for guiding experimental design, yet remains challenging because relevant information in the literature is often scattered, inconsistent, and…

Graphs are ubiquitous data structures for representing interactions between entities. With an emphasis on the use of graphs to represent chemical molecules, we explore the task of learning to generate graphs that conform to a distribution…

机器学习 · 计算机科学 2019-03-08 Qi Liu , Miltiadis Allamanis , Marc Brockschmidt , Alexander L. Gaunt

While many NLP pipelines assume raw, clean texts, many texts we encounter in the wild, including a vast majority of legal documents, are not so clean, with many of them being visually structured documents (VSDs) such as PDFs. Conventional…

计算与语言 · 计算机科学 2021-11-09 Yuta Koreeda , Christopher D. Manning

We present MMOCR-an open-source toolbox which provides a comprehensive pipeline for text detection and recognition, as well as their downstream tasks such as named entity recognition and key information extraction. MMOCR implements 14…

计算机视觉与模式识别 · 计算机科学 2021-08-17 Zhanghui Kuang , Hongbin Sun , Zhizhong Li , Xiaoyu Yue , Tsui Hin Lin , Jianyong Chen , Huaqiang Wei , Yiqin Zhu , Tong Gao , Wenwei Zhang , Kai Chen , Wayne Zhang , Dahua Lin
‹ 上一页 1 2 3 10 下一页 ›