English
Related papers

Related papers: SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.…

200 papers

Knowledge is central to human and scientific developments. Natural Language Processing (NLP) allows automated analysis and creation of knowledge. Data is a crucial NLP and machine learning ingredient. The scarcity of open datasets is a…

Computation and Language · Computer Science 2022-10-19 Istiak Ahmad , Fahad AlQurashi , Rashid Mehmood

This paper introduces the Balanced Arabic Readability Evaluation Corpus (BAREC), a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 69,441 sentences spanning 1+ million words, carefully curated to cover…

Computation and Language · Computer Science 2025-06-17 Khalid N. Elmadani , Nizar Habash , Hanada Taha-Thomure

Representing words and phrases into dense vectors of real numbers which encode semantic and syntactic properties is a vital constituent in natural language processing (NLP). The success of neural network (NN) models in NLP largely rely on…

Computation and Language · Computer Science 2021-01-01 Wazir Ali , Jay Kumar , Junyu Lu , Zenglin Xu

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse…

This paper presents the first publicly available treebank of Odia, a morphologically rich low resource Indian language. The treebank contains approx. 1082 tokens (100 sentences) in Odia selected from "Samantar", the largest available…

Computation and Language · Computer Science 2022-05-25 Shantipriya Parida , Kalyanamalini Sahoo , Atul Kr. Ojha , Saraswati Sahoo , Satya Ranjan Dash , Bijayalaxmi Dash

Deep learning based natural language processing model is proven powerful, but need large-scale dataset. Due to the significant gap between the real-world tasks and existing Chinese corpus, in this paper, we introduce a large-scale corpus of…

Computation and Language · Computer Science 2018-11-27 Jianyu Zhao , Zhuoran Ji

This paper presents the challenges in creating and managing large parallel corpora of 12 major Indian languages (which is soon to be extended to 23 languages) as part of a major consortium project funded by the Department of Information…

Computation and Language · Computer Science 2021-12-06 Ritesh Kumar , Shiv Bhusan Kaushik , Pinkey Nainwani , Girish Nath Jha

This paper presents UD-NewsCrawl, the largest Tagalog treebank to date, containing 15.6k trees manually annotated according to the Universal Dependencies framework. We detail our treebank development process, including data collection,…

Computation and Language · Computer Science 2025-05-28 Angelina A. Aquino , Lester James V. Miranda , Elsie Marie T. Or

Many languages have vast amounts of handwritten texts, such as ancient scripts about folktale stories and historical narratives or contemporary documents and letters. Digitization of those texts has various applications, such as daily…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Ameer Majeed , Hossein Hassani

Over the past years, interest in discourse analysis and discourse parsing has steadily grown, and many discourse-annotated corpora and, as a result, discourse parsers have been built. In this paper, we present a discourse-annotated corpus…

Computation and Language · Computer Science 2021-06-29 Sara Shahmohammadi , Hadi Veisi , Ali Darzi

Dyslexia in adults remains an under-researched and under-served area, particularly in non-English-speaking contexts, despite its significant impact on personal and professional lives. This work addresses that gap by focusing on Sinhala, a…

Computation and Language · Computer Science 2025-10-07 Peshala Perera , Deshan Sumanathilaka

This paper presents SwissCrawl, the largest Swiss German text corpus to date. Composed of more than half a million sentences, it was generated using a customized web scraping tool that could be applied to other low-resource languages as…

Computation and Language · Computer Science 2020-06-17 Lucy Linder , Michael Jungo , Jean Hennebert , Claudiu Musat , Andreas Fischer

We introduce a corpus of 7,032 sentences rated by human annotators for formality, informativeness, and implicature on a 1-7 scale. The corpus was annotated using Amazon Mechanical Turk. Reliability in the obtained judgments was examined by…

Computation and Language · Computer Science 2016-09-29 Shibamouli Lahiri

We present ASCAT (Arabic Scientific Corpus for Advanced Translation), a high-quality English-Arabic parallel benchmark corpus designed for scientific translation evaluation constructed through a systematic multi-engine translation and human…

Computation and Language · Computer Science 2026-04-02 Serry Sibaee , Khloud Al Jallad , Zineb Yousfi , Israa Elsayed Elhosiny , Yousra El-Ghawi , Batool Balah , Omer Nacar

In this paper, we present a new Russian and Kazakh database (with about 95% of Russian and 5% of Kazakh words/sentences respectively) for offline handwriting recognition. A few pre-processing and segmentation procedures have been developed…

Computer Vision and Pattern Recognition · Computer Science 2021-08-31 Daniyar Nurseitov , Kairat Bostanbekov , Daniyar Kurmankhojayev , Anel Alimova , Abdelrahman Abdallah

This paper proposes OCR++, an open-source framework designed for a variety of information extraction tasks from scholarly articles including metadata (title, author names, affiliation and e-mail), structure (section headings and body text,…

In this paper, we introduce Libriheavy, a large-scale ASR corpus consisting of 50,000 hours of read English speech derived from LibriVox. To the best of our knowledge, Libriheavy is the largest freely-available corpus of speech with…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-17 Wei Kang , Xiaoyu Yang , Zengwei Yao , Fangjun Kuang , Yifan Yang , Liyong Guo , Long Lin , Daniel Povey

While strides have been made in deep learning based Bengali Optical Character Recognition (OCR) in the past decade, the absence of large Document Layout Analysis (DLA) datasets has hindered the application of OCR in document transcription,…

Modern VLMs have achieved near-saturation accuracy in English document visual question-answering (VQA). However, this task remains challenging in lower resource languages due to a dearth of suitable training and evaluation data. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Jonathan Li , Zoltan Csaki , Nidhi Hiremath , Etash Guha , Fenglu Hong , Edward Ma , Urmish Thakker

We present the largest publicly available synthetic OCR benchmark dataset for Indic languages. The collection contains a total of 90k images and their ground truth for 23 Indic languages. OCR model validation in Indic languages require a…

Computer Vision and Pattern Recognition · Computer Science 2022-05-06 Naresh Saini , Promodh Pinto , Aravinth Bheemaraj , Deepak Kumar , Dhiraj Daga , Saurabh Yadav , Srihari Nagaraj