English
Related papers

Related papers: SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.…

200 papers

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples,…

Computation and Language · Computer Science 2025-11-10 Haneen Al-Homoud , Asma Ibrahim , Murtadha Al-Jubran , Fahad Al-Otaibi , Yazeed Al-Harbi , Daulet Toibazar , Kesen Wang , Pedro J. Moreno

In this work, we introduce the MOldavian and ROmanian Dialectal COrpus (MOROCO), which is freely available for download at https://github.com/butnaruandrei/MOROCO. The corpus contains 33564 samples of text (with over 10 million tokens)…

Computation and Language · Computer Science 2019-06-04 Andrei M. Butnaru , Radu Tudor Ionescu

This paper is devoted to the adaptation of generative large language models for the Tajik language, a low-resource language with Cyrillic script. To overcome the shortage of digital text resources, the author created and publicly released…

Computation and Language · Computer Science 2026-05-06 Mullosharaf K. Arabov

Due to the high impact of the fast-evolving fields of machine learning and deep learning, Natural Language Processing (NLP) tasks have further obtained comprehensive performances for highly resourced languages such as English and Chinese.…

Computation and Language · Computer Science 2020-11-17 Lahiru Senevirathne , Piyumal Demotte , Binod Karunanayake , Udyogi Munasinghe , Surangika Ranathunga

We present a novel, open-access dataset designed for semantic layout analysis, built to support document recreation workflows through mapping with the Text Encoding Initiative (TEI) standard. This dataset includes 7,254 annotated pages…

Computer Vision and Pattern Recognition · Computer Science 2024-11-18 Thibault Clérice , Juliette Janes , Hugo Scheithauer , Sarah Bénière , Florian Cafiero , Laurent Romary , Simon Gabay , Benoît Sagot

With the advent of Web 2.0, the development in social technology coupled with global communication systematically brought positive and negative impacts to society. Copyright claims and Author identification are deemed crucial as there has…

Computation and Language · Computer Science 2025-01-17 Nabeelah Faumi , Adeepa Gunathilake , Benura Wickramanayake , Deelaka Dias , TGDK Sumanathilaka

Chinese dynastic histories form a large continuous linguistic space of approximately 2000 years, from the 3rd century BCE to the 18th century CE. The histories are documented in Classical (Literary) Chinese in a corpus of over 20 million…

Computation and Language · Computer Science 2020-05-19 Sergey Zinin , Yang Xu

While machine translation is regarded as a "solved problem" for many high-resource languages, close analysis quickly reveals that this is not the case for content that shows challenges such as poetic language, philosophical concepts,…

Computation and Language · Computer Science 2026-01-13 Sebastian Nehrdich , David Allport , Sven Sellmer , Jivnesh Sandhan , Manoj Balaji Jagadeeshan , Pawan Goyal , Sujeet Kumar , Kurt Keutzer

We introduced the contemporary Amharic corpus, which is automatically tagged for morpho-syntactic information. Texts are collected from 25,199 documents from different domains and about 24 million orthographic words are tokenized. Since it…

Computation and Language · Computer Science 2021-06-15 Andargachew Mekonnen Gezmu , Binyam Ephrem Seyoum , Michael Gasser , Andreas Nürnberger

We describe NorDiaChange: the first diachronic semantic change dataset for Norwegian. NorDiaChange comprises two novel subsets, covering about 80 Norwegian nouns manually annotated with graded semantic change over time. Both datasets follow…

Computation and Language · Computer Science 2022-04-29 Andrey Kutuzov , Samia Touileb , Petter Mæhlum , Tita Ranveig Enstad , Alexandra Wittemann

This paper presents the first-ever Sinhala physical common sense reasoning dataset created as part of Global PIQA. It contains 110 human-created and verified data samples, where each sample consists of a prompt, the corresponding correct…

Computation and Language · Computer Science 2026-02-03 Nisansa de Silva , Surangika Ranathunga

This work introduces Itihasa, a large-scale translation dataset containing 93,000 pairs of Sanskrit shlokas and their English translations. The shlokas are extracted from two Indian epics viz., The Ramayana and The Mahabharata. We first…

Computation and Language · Computer Science 2021-10-07 Rahul Aralikatte , Miryam de Lhoneux , Anoop Kunchukuttan , Anders Søgaard

We present spINAch, a large diachronic corpus of French speech from radio and television archives, balanced by speakers' gender, age (20-95 years old), and spanning 60 years from 1955 to 2015. The dataset includes over 320 hours of…

The amount of scholarly data has been increasing dramatically over the last years. For newcomers to a particular science domain (e.g., IR, physics, NLP) it is often difficult to spot larger trends and to position the latest research in the…

Digital Libraries · Computer Science 2021-12-08 Naman Paharia , Muhammad Syafiq Mohd Pozi , Adam Jatowt

Low-resource languages such as Sinhala are often overlooked by open-source Large Language Models (LLMs). In this research, we extend an existing multilingual LLM (Llama-3-8B) to better serve Sinhala. We enhance the LLM tokenizer with…

Computation and Language · Computer Science 2025-11-11 H. W. K. Aravinda , Rashad Sirajudeen , Samith Karunathilake , Nisansa de Silva , Surangika Ranathunga , Rishemjit Kaur

Objective: To build a comprehensive corpus covering syntactic and semantic annotations of Chinese clinical texts with corresponding annotation guidelines and methods as well as to develop tools trained on the annotated corpus, which…

Computation and Language · Computer Science 2016-11-09 Bin He , Bin Dong , Yi Guan , Jinfeng Yang , Zhipeng Jiang , Qiubin Yu , Jianyi Cheng , Chunyan Qu

This paper introduces "Czech Text Document Corpus v 2.0", a collection of text documents for automatic document classification in Czech language. It is composed of the text documents provided by the Czech News Agency and is freely available…

Computation and Language · Computer Science 2018-02-01 Pavel Král , Ladislav Lenc

The Ramayana is among the most influential literary traditions of South and Southeast Asia, transmitted across numerous linguistic and cultural contexts over two millennia. Despite extensive scholarship on regional Ramayana traditions,…

Computation and Language · Computer Science 2026-04-16 Sumesh VP

Long-context question answering (QA) over literary texts poses significant challenges for modern large language models, particularly in low-resource languages. We address the scarcity of long-context QA resources for Indic languages by…

Computation and Language · Computer Science 2026-01-07 Aarya Khandelwal , Ritwik Mishra , Rajiv Ratn Shah

The Swa-bhasha Resource Hub provides a comprehensive collection of data resources and algorithms developed for Romanized Sinhala to Sinhala transliteration between 2020 and 2025. These resources have played a significant role in advancing…