中文
相关论文

相关论文: Logical segmentation for article extraction in dig…

200 篇论文

Extracting text objects from the PDF images is a challenging problem. The text data present in the PDF images contain certain useful information for automatic annotation, indexing etc. However variations of the text due to differences in…

计算机视觉与模式识别 · 计算机科学 2012-10-02 D. Sasirekha , E. Chandra

Archived collections of documents (like newspaper archives) serve as important information sources for historians, journalists, sociologists and other interested parties. Semantic Layers over such digital archives allow describing and…

信息检索 · 计算机科学 2022-10-19 Pavlos Fafalios , Vaibhav Kasturia , Wolfgang Nejdl

Data acquisition forms the primary step in all empirical research. The availability of data directly impacts the quality and extent of conclusions and insights. In particular, larger and more detailed datasets provide convincing answers…

计算机视觉与模式识别 · 计算机科学 2021-02-08 Christian M. Dahl , Torben S. D. Johansen , Emil N. Sørensen , Christian E. Westermann , Simon F. Wittrock

Chemical structure extraction from documents remains a hard problem due to both false positive identification of structures during segmentation and errors in the predicted structures. Current approaches rely on handcrafted rules and…

机器学习 · 计算机科学 2018-02-15 Joshua Staker , Kyle Marshall , Robert Abel , Carolyn McQuaw

With the rapid advancement of tool-use capabilities in Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) is shifting from static, one-shot retrieval toward autonomous, multi-turn evidence acquisition. However, existing…

人工智能 · 计算机科学 2026-02-13 Zhanli Li , Huiwen Tian , Lvzhou Luo , Yixuan Cao , Ping Luo

The most recent advances in medical imaging that have transformed diagnosis, especially in the case of interpreting X-ray images, are actively involved in the healthcare sector. The advent of digital image processing technology and the…

图像与视频处理 · 电气工程与系统科学 2024-06-21 Abhishek Swami , Snehal Farande , Atharv Patil , Atharva Parle , Vivekanand Mane , Prathamesh Thorat

The biggest challenge in the field of image processing is to recognize documents both in printed and handwritten format. Optical Character Recognition OCR is a type of document image analysis where scanned digital image that contains either…

计算机视觉与模式识别 · 计算机科学 2016-12-05 Singh Vijendra , Nisha Vasudeva , Hem Jyotsana Parashar

Image segmentation is the process of partitioning an image into a set of meaningful regions according to some criteria. Hierarchical segmentation has emerged as a major trend in this regard as it favors the emergence of important regions at…

计算机视觉与模式识别 · 计算机科学 2017-03-10 Amin Fehri , Santiago Velasco-Forero , Fernand Meyer

Understanding visually situated language requires interpreting complex layouts of textual and visual elements. Pre-processing tools, such as optical character recognition (OCR), can map document image inputs to textual tokens, then large…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Wang Zhu , Alekh Agarwal , Mandar Joshi , Robin Jia , Jesse Thomason , Kristina Toutanova

Content analysis of news stories (whether manual or automatic) is a cornerstone of the communication studies field. However, much research is conducted at the level of individual news articles, despite the fact that news events (especially…

社会与信息网络 · 计算机科学 2024-10-31 Tom Nicholls , Jonathan Bright

In the digital era, the exponential growth of scientific publications has made it increasingly difficult for researchers to efficiently identify and access relevant work. This paper presents an automated framework for research article…

信息检索 · 计算机科学 2025-10-08 Shadikur Rahman , Hasibul Karim Shanto , Umme Ayman Koana , Syed Muhammad Danish

Image segmentation refers to the process to divide an image into nonoverlapping meaningful regions according to human perception, which has become a classic topic since the early ages of computer vision. A lot of research has been conducted…

计算机视觉与模式识别 · 计算机科学 2015-02-04 Hongyuan Zhu , Fanman Meng , Jianfei Cai , Shijian Lu

Online news media provides aggregated news and stories from different sources all over the world and up-to-date news coverage. The main goal of this study is to have a solution that considered as a homogeneous source for the news and to…

信息检索 · 计算机科学 2019-06-25 Aboubakr Aqle , Dena Al-Thani , Ali Jaoua

Academic literature retrieval is concerned with the selection of papers that are most likely to match a user's information needs. Most of the retrieval systems are limited to list-output models, in which the retrieval results are isolated…

信息检索 · 计算机科学 2017-11-27 Danping Liao , Yuntao Qian

On the one hand, nowadays, fake news articles are easily propagated through various online media platforms and have become a grand threat to the trustworthiness of information. On the other hand, our understanding of the language of fake…

计算与语言 · 计算机科学 2019-04-11 Hamid Karimi , Jiliang Tang

Text line detection is crucial for any application associated with Automatic Text Recognition or Keyword Spotting. Modern algorithms perform good on well-established datasets since they either comprise clean data or simple/homogeneous page…

计算机视觉与模式识别 · 计算机科学 2017-12-12 Tobias Grüning , Roger Labahn , Markus Diem , Florian Kleber , Stefan Fiel

We present docExtractor, a generic approach for extracting visual elements such as text lines or illustrations from historical documents without requiring any real data annotation. We demonstrate it provides high-quality performances as an…

计算机视觉与模式识别 · 计算机科学 2020-12-16 Tom Monnier , Mathieu Aubry

With the recent developments in digitisation, there are increasing number of documents available online. There are several information extraction tools that are available to extract information from digitised documents. However, identifying…

News will be biased so long as people have opinions. As social media becomes the primary entry point for news and partisan differences increase, it is increasingly important for informed citizens to be able to recognize bias. If people are…

计算与语言 · 计算机科学 2025-05-22 Jessica Zhu , Iain Cruickshank , Michel Cukier

Information extraction systems often produce hundreds to thousands of strings on a specific topic. We present a method that facilitates better consumption of these strings, in an exploratory setting in which a user wants to both get a broad…

计算与语言 · 计算机科学 2023-09-20 Itay Yair , Hillel Taub-Tabib , Yoav Goldberg