中文
相关论文

相关论文: Automatic Removal of Marginal Annotations in Print…

200 篇论文

The pressing need for digitization of historical documents has led to a strong interest in designing computerised image processing methods for automatic handwritten text recognition. However, not much attention has been paid on studying the…

计算机视觉与模式识别 · 计算机科学 2024-01-31 Liang Cheng , Jonas Frankemölle , Adam Axelsson , Ekta Vats

Handwritten text recognition has been widely studied in the last decades for its numerous applications. Nowadays, the state-of-the-art approach consists in a three-step process. The document is segmented into text lines, which are then…

计算机视觉与模式识别 · 计算机科学 2022-10-21 Denis Coquenet

The performance of information retrieval algorithms depends upon the availability of ground truth labels annotated by experts. This is an important prerequisite, and difficulties arise when the annotated ground truth labels are incorrect or…

信息检索 · 计算机科学 2018-02-22 Ekta Vats , Anders Hast

Large amounts of annotated data have become more important than ever, especially since the rise of deep learning techniques. However, manual annotations are costly. We propose a tool that enables researchers to create large, high-quality,…

数字图书馆 · 计算机科学 2021-12-23 Franziska Weeber , Felix Hamborg , Karsten Donnay , Bela Gipp

With the surging inclination towards carrying out tasks on computational devices and digital mediums, any method that converts a task that was previously carried out manually, to a digitized version, is always welcome. Irrespective of the…

计算机视觉与模式识别 · 计算机科学 2023-04-19 Pranav Guruprasad , Sujith Kumar S , Vigneswaran C , V. Srinivasa Chakravarthy

Camera-captured document images usually suffer from perspective and geometric deformations. It is of great value to rectify them when considering poor visual aesthetics and the deteriorated performance of OCR systems. Recent learning-based…

计算机视觉与模式识别 · 计算机科学 2023-07-13 Jiaxin Zhang , Canjie Luo , Lianwen Jin , Fengjun Guo , Kai Ding

This paper presents a framework for semi-automatic transcription of large-scale historical handwritten documents and proposes a simple user-friendly text extractor tool, TexT for transcription. The proposed approach provides a quick and…

数字图书馆 · 计算机科学 2018-02-22 Anders Hast , Per Cullhed , Ekta Vats

Image restoration is very crucial computer vision task. This paper describes two novel methods for the restoration of old degraded handwritten documents using deep neural network. In addition to that, a small-scale dataset of 26 heritage…

计算机视觉与模式识别 · 计算机科学 2020-01-27 Mayank Wadhwani , Debapriya Kundu , Deepayan Chakraborty , Bhabatosh Chanda

It is a common phenomenon in day to day life; where in some of the document gets damaged. Out of several reasons, the main reason for documents getting damaged is shredding by hands. Recovery of such documents is essential. Manual recovery…

计算机视觉与模式识别 · 计算机科学 2015-03-10 Waheeda Dhokley , Khan Munifa , Shaikh Nazia , Shaikh Saiqua

OCR has been an active research area since last few decades. OCR performs the recognition of the text in the scanned document image and converts it into editable form. The OCR process can have several stages like pre-processing,…

计算机视觉与模式识别 · 计算机科学 2013-06-07 Kanika Bansal , Rajiv Kumar

The aim of the paper is to separate handwritten and printed text from a real document embedded with noise, graphics including annotations. Relying on run-length smoothing algorithm (RLSA), the extracted pseudo-lines and pseudo-words are…

计算机视觉与模式识别 · 计算机科学 2013-03-20 Abdel Belaïd , K. C. Santosh , Vincent Poulain D'Andecy

One of the elements of legal research is looking for cases where judges have extended the meaning of a legal concept by providing interpretations of what a concept means or does not mean. This allow legal professionals to use such…

计算与语言 · 计算机科学 2025-06-18 Aleksander Smywiński-Pohl , Tomer Libal , Adam Kaczmarczyk , Magdalena Król

This paper presents an analysis of annotation using an automatic pre-annotation for a mid-level annotation complexity task -- dependency syntax annotation. It compares the annotation efforts made by annotators using a pre-annotated version…

计算与语言 · 计算机科学 2023-06-16 Marie Mikulová , Milan Straka , Jan Štěpánek , Barbora Štěpánková , Jan Hajič

Annotated data is an essential ingredient in natural language processing for training and evaluating machine learning models. It is therefore very desirable for the annotations to be of high quality. Recent work, however, has shown that…

计算与语言 · 计算机科学 2022-09-27 Jan-Christoph Klie , Bonnie Webber , Iryna Gurevych

Text documents, including programs, typically have human-readable semantic structure. Historically, programmatic access to these semantics has required explicit in-document tagging. Especially in systems where the text has an execution…

计算与语言 · 计算机科学 2024-03-07 Edward Misback , Zachary Tatlock , Steven L. Tanimoto

Document dewarping from a distorted camera-captured image is of great value for OCR and document understanding. The document boundary plays an important role which is more evident than the inner region in document dewarping. Current…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Beiya Dai , Xing li , Qunyi Xie , Yulin Li , Xiameng Qin , Chengquan Zhang , Kun Yao , Junyu Han

Pool of knowledge available to the mankind depends on the source of learning resources, which can vary from ancient printed documents to present electronic material. The rapid conversion of material available in traditional libraries to…

计算机视觉与模式识别 · 计算机科学 2014-12-25 Akmal Jahan Mac , Roshan G Ragel

One of the problems on the way to successful implementation of neural networks is the quality of annotation. For instance, different annotators can annotate images in a different way and very often their decisions do not match exactly and…

计算机视觉与模式识别 · 计算机科学 2018-07-25 Roman Khudorozhkov , Alexander Koryagin , Alexey Kozhevin

There is a huge amount of historical documents in libraries and in various National Archives that have not been exploited electronically. Although automatic reading of complete pages remains, in most cases, a long-term objective, tasks such…

计算机视觉与模式识别 · 计算机科学 2007-05-23 Laurence Likforman-Sulem , Abderrazak Zahour , Bruno Taconet

Automated paper reproduction has emerged as a promising approach to accelerate scientific research, employing multi-step workflow frameworks to systematically convert academic papers into executable code. However, existing frameworks often…

人工智能 · 计算机科学 2025-12-03 Zijie Lin , Qilin Cai , Liang Shen , Mingjun Xiao
‹ 上一页 1 2 3 10 下一页 ›