English
Related papers

Related papers: DriveThru: a Document Extraction Platform and Benc…

200 papers

Billions of public domain documents remain trapped in hard copy or lack an accurate digitization. Modern natural language processing methods cannot be used to index, retrieve, and summarize their texts; conduct computational textual…

Computer Vision and Pattern Recognition · Computer Science 2023-10-17 Tom Bryan , Jacob Carlson , Abhishek Arora , Melissa Dell

As the Internet help us cross language and cultural border by providing different types of translation tools, cross language plagiarism, also known as translation plagiarism are bound to arise. Especially among the academic works, such…

Other Computer Science · Computer Science 2009-12-22 Chow Kok Kent , Naomie Salim

We present the largest publicly available synthetic OCR benchmark dataset for Indic languages. The collection contains a total of 90k images and their ground truth for 23 Indic languages. OCR model validation in Indic languages require a…

Computer Vision and Pattern Recognition · Computer Science 2022-05-06 Naresh Saini , Promodh Pinto , Aravinth Bheemaraj , Deepak Kumar , Dhiraj Daga , Saurabh Yadav , Srihari Nagaraj

Weak supervision has emerged as a promising approach for rapid and large-scale dataset creation in response to the increasing demand for accelerated NLP development. By leveraging labeling functions, weak supervision allows practitioners to…

Computation and Language · Computer Science 2023-10-25 Mega Fransiska , Diah Pitaloka , Saripudin , Satrio Putra , Lintang Sutawika

Dysgraphia is a learning disorder that affects handwriting abilities, making it challenging for children to write legibly and consistently. Early detection and monitoring are crucial for providing timely support and interventions. This…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Vydeki D , Divyansh Bhandari , Pranav Pratap Patil , Aarush Anand Kulkarni

E-commerce provides an efficient and effective way to exchange goods between sellers and customers. E-commerce has been a popular method for doing business, because of its simplicity of having commerce activity transparently available,…

General Economics · Economics 2021-02-19 Andry Alamsyah , Nurlisa Laksmiani , Lies Anisa Rahimi

Manchu, a critically endangered language essential for understanding early modern Eastern Eurasian history, lacks effective OCR systems that can handle real-world historical documents. This study develops high-performing OCR systems by…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Yan Hon Michael Chung , Donghyeok Choi

Information extraction from copy-heavy documents, characterized by massive volumes of structurally similar content, represents a critical yet understudied challenge in enterprise document processing. We present a systematic framework that…

Computation and Language · Computer Science 2025-10-14 Zilong Wang , Xiaoyu Shen

In this paper, we present an Optical Character Recognition (OCR) system specifically designed for the accurate recognition and digitization of Greek polytonic texts. By leveraging the combined strengths of convolutional layers for feature…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Perifanos Konstantinos , Goutsos Dionisis

This paper accompanies the software documentation data set for machine translation, a parallel evaluation data set of data originating from the SAP Help Portal, that we released to the machine translation community for research purposes. It…

Computation and Language · Computer Science 2020-11-13 Bianka Buschbeck , Miriam Exel

Solving the problem of Optical Character Recognition (OCR) on printed text for Latin and its derivative scripts can now be considered settled due to the volumes of research done on English and other High-Resourced Languages (HRL). However,…

Computation and Language · Computer Science 2025-08-26 Nevidu Jayatilleke , Nisansa de Silva

Neural retrieval and GPT-style generative models rely on large, high-quality supervised data, which is still scarce for low-resource languages such as Amharic. We release an Amharic data resource consisting of two datasets that supports…

Computation and Language · Computer Science 2026-02-11 Tilahun Yeshambel , Moncef Garouani , Josiane Mothe

The Yunshan Cup 2020 track focused on creating a framework for evaluating different methods of part-of-speech (POS). There were two tasks for this track: (1) POS tagging for the Indonesian language, and (2) POS tagging for the Lao tagging.…

Computation and Language · Computer Science 2022-04-07 Yingwen Fu , Jinyi Chen , Nankai Lin , Xixuan Huang , Xinying Qiu , Shengyi Jiang

This paper provides an overall introduction of our Automatic Speech Recognition (ASR) systems for Southeast Asian languages. As not much existing work has been carried out on such regional languages, a few difficulties should be addressed…

Computation and Language · Computer Science 2022-10-10 Lei Wang , Rong Tong , Cheung Chi Leung , Sunil Sivadas , Chongjia Ni , Bin Ma

Information extraction (IE) from unstructured documents remains a critical challenge in data processing pipelines. Traditional optical character recognition (OCR) methods and conventional parsing engines demonstrate limited effectiveness…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Aditya Parikh

While the NLP community is generally aware of resource disparities among languages, we lack research that quantifies the extent and types of such disparity. Prior surveys estimating the availability of resources based on the number of…

Computation and Language · Computer Science 2022-11-29 Xinyan Velocity Yu , Akari Asai , Trina Chatterjee , Junjie Hu , Eunsol Choi

Visually Rich Document Understanding (VRDU) has become a pivotal area of research, driven by the need to automatically interpret documents that contain intricate visual, textual, and structural elements. Recently, Multimodal Large Language…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Yihao Ding , Siwen Luo , Yue Dai , Yanbei Jiang , Zechuan Li , Qiang Sun , Geoffrey Martin , Wei Liu , Yifan Peng

Big data research in Indonesia is constrained by a fundamental fragmentation: relevant data is scattered across social media, news portals, e-commerce platforms, review sites, and academic databases, each with different formats, access…

This short paper is intended as an additional progress report to share our experiences in Indonesia on collecting, integrating and disseminating both global and local scientific data across the country through the web technology. Our recent…

Computers and Society · Computer Science 2009-03-05 L. T. Handoko

Multilingual task-oriented dialogue (ToD) facilitates access to services and information for many (communities of) speakers. Nevertheless, the potential of this technology is not fully realised, as current datasets for multilingual ToD -…

Computation and Language · Computer Science 2023-05-24 Olga Majewska , Evgeniia Razumovskaia , Edoardo Maria Ponti , Ivan Vulić , Anna Korhonen
‹ Prev 1 4 5 6 7 8 10 Next ›