English
Related papers

Related papers: SiDiaC: Sinhala Diachronic Corpus

200 papers

In this work, we introduce the MOldavian and ROmanian Dialectal COrpus (MOROCO), which is freely available for download at https://github.com/butnaruandrei/MOROCO. The corpus contains 33564 samples of text (with over 10 million tokens)…

Computation and Language · Computer Science 2019-06-04 Andrei M. Butnaru , Radu Tudor Ionescu

This study demonstrates how hybrid neural-symbolic methods can yield significant new insights into the evolution of a morphologically rich, low-resource language. We challenge the naive assumption that linguistic change is simplification by…

Computation and Language · Computer Science 2025-12-08 Ananth Hariharan , David Mortensen

Kurdish is a less-resourced language consisting of different dialects written in various scripts. Approximately 30 million people in different countries speak the language. The lack of corpora is one of the main obstacles in Kurdish…

Computation and Language · Computer Science 2019-09-26 Roshna Omer Abdulrahman , Hossein Hassani , Sina Ahmadi

Dyslexia in adults remains an under-researched and under-served area, particularly in non-English-speaking contexts, despite its significant impact on personal and professional lives. This work addresses that gap by focusing on Sinhala, a…

Computation and Language · Computer Science 2025-10-07 Peshala Perera , Deshan Sumanathilaka

We present spINAch, a large diachronic corpus of French speech from radio and television archives, balanced by speakers' gender, age (20-95 years old), and spanning 60 years from 1955 to 2015. The dataset includes over 320 hours of…

The Ramayana is among the most influential literary traditions of South and Southeast Asia, transmitted across numerous linguistic and cultural contexts over two millennia. Despite extensive scholarship on regional Ramayana traditions,…

Computation and Language · Computer Science 2026-04-16 Sumesh VP

The Facebook network allows its users to record their reactions to text via a typology of emotions. This network, taken at scale, is therefore a prime data set of annotated sentiment data. This paper uses millions of such reactions, derived…

Machine Learning · Computer Science 2022-08-04 Vihanga Jayawickrama , Gihan Weeraprameshwara , Nisansa de Silva , Yudhanjaya Wijeratne

This paper presents first benchmark corpus of Sanskrit Pratyaya (suffix) and inflectional words (padas) formed due to suffixes along with neural network based approaches to process the formation and splitting of inflectional words.…

Computation and Language · Computer Science 2024-09-05 Arun Kumar Singh , Sushant Dave , Prathosh A. P. , Brejesh Lall , Shresth Mehta

Hausa language belongs to the Afroasiatic phylum, and with more first-language speakers than any other sub-Saharan African language. With a majority of its speakers residing in the Northern and Southern areas of Nigeria and the Republic of…

Computation and Language · Computer Science 2021-02-18 Isa Inuwa-Dutse

We present a novel, open-access dataset designed for semantic layout analysis, built to support document recreation workflows through mapping with the Text Encoding Initiative (TEI) standard. This dataset includes 7,254 annotated pages…

Computer Vision and Pattern Recognition · Computer Science 2024-11-18 Thibault Clérice , Juliette Janes , Hugo Scheithauer , Sarah Bénière , Florian Cafiero , Laurent Romary , Simon Gabay , Benoît Sagot

A Kannada OCR, named Lipi Gnani, has been designed and developed from scratch, with the motivation of it being able to convert printed text or poetry in Kannada script, without any restriction on vocabulary. The training and test sets have…

Computer Vision and Pattern Recognition · Computer Science 2019-01-03 Shiva Kumar H R , Ramakrishnan A G

Figures of Speech (FoS) consist of multi-word phrases that are deeply intertwined with culture. While Neural Machine Translation (NMT) performs relatively well with the figurative expressions of high-resource languages, it often faces…

Computation and Language · Computer Science 2026-02-11 Johan Sofalas , Dilushri Pavithra , Nevidu Jayatilleke , Ruvan Weerasinghe

We introduced the contemporary Amharic corpus, which is automatically tagged for morpho-syntactic information. Texts are collected from 25,199 documents from different domains and about 24 million orthographic words are tokenized. Since it…

Computation and Language · Computer Science 2021-06-15 Andargachew Mekonnen Gezmu , Binyam Ephrem Seyoum , Michael Gasser , Andreas Nürnberger

This research investigates the area of Music Information Retrieval (MIR) and Music Emotion Recognition (MER) in relation to Sinhala songs, an underexplored field in music studies. The purpose of this study is to analyze the behavior of…

Computation and Language · Computer Science 2025-02-03 W. M. Yomal De Mel , Nisansa de Silva

This paper introduces TRIDIS (Tria Digita Scribunt), an open-source corpus of medieval and early modern manuscripts. TRIDIS aggregates multiple legacy collections (all published under open licenses) and incorporates large metadata…

Computation and Language · Computer Science 2025-04-18 Sergio Torres Aguilar

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples,…

Computation and Language · Computer Science 2025-11-10 Haneen Al-Homoud , Asma Ibrahim , Murtadha Al-Jubran , Fahad Al-Otaibi , Yazeed Al-Harbi , Daulet Toibazar , Kesen Wang , Pedro J. Moreno

The ACL Anthology (AA) is a digital repository of tens of thousands of articles on Natural Language Processing (NLP). This paper examines the literature as a whole to identify broad trends in productivity, focus, and impact. It presents the…

Digital Libraries · Computer Science 2019-11-12 Saif M. Mohammad

This paper presents a pipeline designed to transform raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divided into two major phases. The first involves extracting and cleaning text from raw…

Computation and Language · Computer Science 2026-05-18 Mihailo Škorić , Cosimo Palma

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognized as scheduled languages in the Indian Constitution. Despite…

Understanding stories is a challenging reading comprehension problem for machines as it requires reading a large volume of text and following long-range dependencies. In this paper, we introduce the Shmoop Corpus: a dataset of 231 stories…

Computation and Language · Computer Science 2020-01-03 Atef Chaudhury , Makarand Tapaswi , Seung Wook Kim , Sanja Fidler