中文
相关论文

相关论文: ScanBank: A Benchmark Dataset for Figure Extractio…

200 篇论文

Non-textual components such as charts, diagrams and tables provide key information in many scientific documents, but the lack of large labeled datasets has impeded the development of data-driven methods for scientific figure extraction. In…

数字图书馆 · 计算机科学 2018-06-01 Noah Siegel , Nicholas Lourie , Russell Power , Waleed Ammar

Electronic Theses and Dissertations (ETDs) contain domain knowledge that can be used for many digital library tasks, such as analyzing citation networks and predicting research trends. Automatic metadata extraction is important to build…

数字图书馆 · 计算机科学 2021-07-02 Muntabir Hasan Choudhury , Himarsha R. Jayanetti , Jian Wu , William A. Ingram , Edward A. Fox

Electronic theses and dissertations (ETDs) have been proposed, advocated, and generated for more than 25 years. Although ETDs are hosted by commercial or institutional digital library repositories, they are still an understudied type of…

计算机视觉与模式识别 · 计算机科学 2023-11-09 Muntabir Hasan Choudhury , Lamia Salsabil , William A. Ingram , Edward A. Fox , Jian Wu

Document layout analysis usually relies on computer vision models to understand documents while ignoring textual information that is vital to capture. Meanwhile, high quality labeled datasets with both visual and textual information are…

计算与语言 · 计算机科学 2020-11-12 Minghao Li , Yiheng Xu , Lei Cui , Shaohan Huang , Furu Wei , Zhoujun Li , Ming Zhou

Scientific papers use schematic diagrams to communicate methods, workflows, and system structure, yet existing scientific-figure corpora often mix them with plots, screenshots, and photographs and rarely preserve document context. We…

信息检索 · 计算机科学 2026-05-28 Ling Yue , Tingwen Zhang , Jiaying Wang , Zhen Xu , Shaowu Pan

Extracting information from academic PDF documents is crucial for numerous indexing, retrieval, and analysis use cases. Choosing the best tool to extract specific content elements is difficult because many, technically diverse tools are…

信息检索 · 计算机科学 2023-03-20 Norman Meuschke , Apurva Jagdale , Timo Spinde , Jelena Mitrović , Bela Gipp

Important information that relates to a specific topic in a document is often organized in tabular format to assist readers with information retrieval and comparison, which may be difficult to provide in natural language. However, tabular…

计算机视觉与模式识别 · 计算机科学 2020-03-05 Xu Zhong , Elaheh ShafieiBavani , Antonio Jimeno Yepes

We present TableBank, a new image-based table detection and recognition dataset built with novel weak supervision from Word and Latex documents on the internet. Existing research for image-based table detection and recognition usually…

计算机视觉与模式识别 · 计算机科学 2020-07-07 Minghao Li , Lei Cui , Shaohan Huang , Furu Wei , Ming Zhou , Zhoujun Li

Pool of knowledge available to the mankind depends on the source of learning resources, which can vary from ancient printed documents to present electronic material. The rapid conversion of material available in traditional libraries to…

计算机视觉与模式识别 · 计算机科学 2014-12-25 Akmal Jahan Mac , Roshan G Ragel

Manual digitization of bibliographic metadata is time consuming and labor intensive, especially for historical and real-world archives with highly variable formatting across documents. Despite advances in machine learning, the absence of…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Jan Kohút , Martin Dočekal , Michal Hradiš , Marek Vaško

Document Structured Extraction (DSE) aims to extract structured content from raw documents. Despite the emergence of numerous DSE systems, their unified evaluation remains inadequate, significantly hindering the field's advancement. This…

计算与语言 · 计算机科学 2025-07-15 Zichao Li , Aizier Abulaiti , Yaojie Lu , Xuanang Chen , Jia Zheng , Hongyu Lin , Xianpei Han , Le Sun

Traditional archival practices for describing electronic theses and dissertations (ETDs) rely on broad, high-level metadata schemes that fail to capture the depth, complexity, and interdisciplinary nature of these long scholarly works. The…

数字图书馆 · 计算机科学 2025-02-05 Bipasha Banerjee , William A. Ingram , Edward A. Fox

This poster addresses accessibility issues of electronic theses and dissertations (ETDs) in digital libraries (DLs). ETDs are available primarily as PDF files, which present barriers to equitable access, especially for users with visual…

数字图书馆 · 计算机科学 2023-10-31 William A. Ingram , Jian Wu , Edward A. Fox

Table Extraction (TE) consists in extracting tables from PDF documents, in a structured format which can be automatically processed. While numerous TE tools exist, the variety of methods and techniques makes it difficult for users to choose…

数据库 · 计算机科学 2025-11-21 Marijan Soric , Cécile Gracianne , Ioana Manolescu , Pierre Senellart

Table extraction from document images is a challenging AI problem, and labelled data for many content domains is difficult to come by. Existing table extraction datasets often focus on scientific tables due to the vast amount of academic…

机器学习 · 计算机科学 2024-12-06 Ethan Bradley , Muhammad Roman , Karen Rafferty , Barry Devereux

The development of artificial intelligence systems for colonoscopy analysis often necessitates expert-annotated image datasets. However, limitations in dataset size and diversity impede model performance and generalisation. Image-text…

计算机视觉与模式识别 · 计算机科学 2023-10-18 Shuo Wang , Yan Zhu , Xiaoyuan Luo , Zhiwei Yang , Yizhe Zhang , Peiyao Fu , Manning Wang , Zhijian Song , Quanlin Li , Pinghong Zhou , Yike Guo

Eye movements in reading play a crucial role in psycholinguistic research studying the cognitive mechanisms underlying human language processing. More recently, the tight coupling between eye movements and cognition has also been leveraged…

计算与语言 · 计算机科学 2023-10-25 Lena S. Bolliger , David R. Reich , Patrick Haller , Deborah N. Jakobi , Paul Prasse , Lena A. Jäger

The scientific literature is growing faster than ever. Finding an expert in a particular scientific domain has never been as hard as today because of the increasing amount of publications and because of the ever growing diversity of…

信息检索 · 计算机科学 2020-04-09 Robin Brochier , Antoine Gourru , Adrien Guille , Julien Velcin

Highly specific datasets of scientific literature are important for both research and education. However, it is difficult to build such datasets at scale. A common approach is to build these datasets reductively by applying topic modeling…

Within the past few decades we have witnessed digital revolution, which moved scholarly communication to electronic media and also resulted in a substantial increase in its volume. Nowadays keeping track with the latest scientific…

数字图书馆 · 计算机科学 2017-10-30 Dominika Tkaczyk
‹ 上一页 1 2 3 10 下一页 ›