中文
相关论文

相关论文: DriveThru: a Document Extraction Platform and Benc…

200 篇论文

Billions of public domain documents remain trapped in hard copy or lack an accurate digitization. Modern natural language processing methods cannot be used to index, retrieve, and summarize their texts; conduct computational textual…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Tom Bryan , Jacob Carlson , Abhishek Arora , Melissa Dell

As the Internet help us cross language and cultural border by providing different types of translation tools, cross language plagiarism, also known as translation plagiarism are bound to arise. Especially among the academic works, such…

其他计算机科学 · 计算机科学 2009-12-22 Chow Kok Kent , Naomie Salim

We present the largest publicly available synthetic OCR benchmark dataset for Indic languages. The collection contains a total of 90k images and their ground truth for 23 Indic languages. OCR model validation in Indic languages require a…

计算机视觉与模式识别 · 计算机科学 2022-05-06 Naresh Saini , Promodh Pinto , Aravinth Bheemaraj , Deepak Kumar , Dhiraj Daga , Saurabh Yadav , Srihari Nagaraj

Weak supervision has emerged as a promising approach for rapid and large-scale dataset creation in response to the increasing demand for accelerated NLP development. By leveraging labeling functions, weak supervision allows practitioners to…

计算与语言 · 计算机科学 2023-10-25 Mega Fransiska , Diah Pitaloka , Saripudin , Satrio Putra , Lintang Sutawika

Dysgraphia is a learning disorder that affects handwriting abilities, making it challenging for children to write legibly and consistently. Early detection and monitoring are crucial for providing timely support and interventions. This…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Vydeki D , Divyansh Bhandari , Pranav Pratap Patil , Aarush Anand Kulkarni

E-commerce provides an efficient and effective way to exchange goods between sellers and customers. E-commerce has been a popular method for doing business, because of its simplicity of having commerce activity transparently available,…

综合经济学 · 经济学 2021-02-19 Andry Alamsyah , Nurlisa Laksmiani , Lies Anisa Rahimi

Manchu, a critically endangered language essential for understanding early modern Eastern Eurasian history, lacks effective OCR systems that can handle real-world historical documents. This study develops high-performing OCR systems by…

计算机视觉与模式识别 · 计算机科学 2025-07-10 Yan Hon Michael Chung , Donghyeok Choi

Information extraction from copy-heavy documents, characterized by massive volumes of structurally similar content, represents a critical yet understudied challenge in enterprise document processing. We present a systematic framework that…

计算与语言 · 计算机科学 2025-10-14 Zilong Wang , Xiaoyu Shen

In this paper, we present an Optical Character Recognition (OCR) system specifically designed for the accurate recognition and digitization of Greek polytonic texts. By leveraging the combined strengths of convolutional layers for feature…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Perifanos Konstantinos , Goutsos Dionisis

This paper accompanies the software documentation data set for machine translation, a parallel evaluation data set of data originating from the SAP Help Portal, that we released to the machine translation community for research purposes. It…

计算与语言 · 计算机科学 2020-11-13 Bianka Buschbeck , Miriam Exel

Solving the problem of Optical Character Recognition (OCR) on printed text for Latin and its derivative scripts can now be considered settled due to the volumes of research done on English and other High-Resourced Languages (HRL). However,…

计算与语言 · 计算机科学 2025-08-26 Nevidu Jayatilleke , Nisansa de Silva

Neural retrieval and GPT-style generative models rely on large, high-quality supervised data, which is still scarce for low-resource languages such as Amharic. We release an Amharic data resource consisting of two datasets that supports…

计算与语言 · 计算机科学 2026-02-11 Tilahun Yeshambel , Moncef Garouani , Josiane Mothe

The Yunshan Cup 2020 track focused on creating a framework for evaluating different methods of part-of-speech (POS). There were two tasks for this track: (1) POS tagging for the Indonesian language, and (2) POS tagging for the Lao tagging.…

计算与语言 · 计算机科学 2022-04-07 Yingwen Fu , Jinyi Chen , Nankai Lin , Xixuan Huang , Xinying Qiu , Shengyi Jiang

This paper provides an overall introduction of our Automatic Speech Recognition (ASR) systems for Southeast Asian languages. As not much existing work has been carried out on such regional languages, a few difficulties should be addressed…

计算与语言 · 计算机科学 2022-10-10 Lei Wang , Rong Tong , Cheung Chi Leung , Sunil Sivadas , Chongjia Ni , Bin Ma

Information extraction (IE) from unstructured documents remains a critical challenge in data processing pipelines. Traditional optical character recognition (OCR) methods and conventional parsing engines demonstrate limited effectiveness…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Aditya Parikh

While the NLP community is generally aware of resource disparities among languages, we lack research that quantifies the extent and types of such disparity. Prior surveys estimating the availability of resources based on the number of…

计算与语言 · 计算机科学 2022-11-29 Xinyan Velocity Yu , Akari Asai , Trina Chatterjee , Junjie Hu , Eunsol Choi

Visually Rich Document Understanding (VRDU) has become a pivotal area of research, driven by the need to automatically interpret documents that contain intricate visual, textual, and structural elements. Recently, Multimodal Large Language…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Yihao Ding , Siwen Luo , Yue Dai , Yanbei Jiang , Zechuan Li , Qiang Sun , Geoffrey Martin , Wei Liu , Yifan Peng

Big data research in Indonesia is constrained by a fundamental fragmentation: relevant data is scattered across social media, news portals, e-commerce platforms, review sites, and academic databases, each with different formats, access…

This short paper is intended as an additional progress report to share our experiences in Indonesia on collecting, integrating and disseminating both global and local scientific data across the country through the web technology. Our recent…

计算机与社会 · 计算机科学 2009-03-05 L. T. Handoko

Multilingual task-oriented dialogue (ToD) facilitates access to services and information for many (communities of) speakers. Nevertheless, the potential of this technology is not fully realised, as current datasets for multilingual ToD -…

计算与语言 · 计算机科学 2023-05-24 Olga Majewska , Evgeniia Razumovskaia , Edoardo Maria Ponti , Ivan Vulić , Anna Korhonen