English
Related papers

Related papers: SiDiaC: Sinhala Diachronic Corpus

200 papers

We present a collection of open, machine-readable document datasets covering parliamentary proceedings, legal judgments, government publications, news, and tourism statistics from Sri Lanka. The collection currently comprises of 269,194…

Computation and Language · Computer Science 2026-05-18 Nuwan I. Senaratna

The digitisation of classical Sanskrit literature is impeded by a scarcity of annotated resources, particularly for Named Entity Recognition. While recent methodologies utilise generic Large Language Models (LLMs) for data augmentation,…

Computation and Language · Computer Science 2026-04-30 Akhil Rajeev P , Annarao Kulkarni

Large Language Models (LLMs) demonstrate impressive general knowledge and reasoning abilities, yet their evaluation has predominantly focused on global or anglocentric subjects, often neglecting low-resource languages and culturally…

This paper presents the first-ever Sinhala physical common sense reasoning dataset created as part of Global PIQA. It contains 110 human-created and verified data samples, where each sample consists of a prompt, the corresponding correct…

Computation and Language · Computer Science 2026-02-03 Nisansa de Silva , Surangika Ranathunga

In this paper, a method for measuring synchronic corpus (dis-)similarity put forward by Kilgarriff (2001) is adapted and extended to identify trends and correlated changes in diachronic text data, using the Corpus of Historical American…

Computation and Language · Computer Science 2015-08-28 Alexander Koplenig

The history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remained largely unexplored…

Computation and Language · Computer Science 2026-05-04 Seyoung Song , Nawon Kim , Songeun Chae , Kiwoong Park , Jiho Jin , Haneul Yoo , Kyunghyun Cho , Alice Oh

Due to the high impact of the fast-evolving fields of machine learning and deep learning, Natural Language Processing (NLP) tasks have further obtained comprehensive performances for highly resourced languages such as English and Chinese.…

Computation and Language · Computer Science 2020-11-17 Lahiru Senevirathne , Piyumal Demotte , Binod Karunanayake , Udyogi Munasinghe , Surangika Ranathunga

With the advent of Web 2.0, the development in social technology coupled with global communication systematically brought positive and negative impacts to society. Copyright claims and Author identification are deemed crucial as there has…

Computation and Language · Computer Science 2025-01-17 Nabeelah Faumi , Adeepa Gunathilake , Benura Wickramanayake , Deelaka Dias , TGDK Sumanathilaka

While machine translation is regarded as a "solved problem" for many high-resource languages, close analysis quickly reveals that this is not the case for content that shows challenges such as poetic language, philosophical concepts,…

Computation and Language · Computer Science 2026-01-13 Sebastian Nehrdich , David Allport , Sven Sellmer , Jivnesh Sandhan , Manoj Balaji Jagadeeshan , Pawan Goyal , Sujeet Kumar , Kurt Keutzer

Low-resource languages such as Sinhala are often overlooked by open-source Large Language Models (LLMs). In this research, we extend an existing multilingual LLM (Llama-3-8B) to better serve Sinhala. We enhance the LLM tokenizer with…

Computation and Language · Computer Science 2025-11-11 H. W. K. Aravinda , Rashad Sirajudeen , Samith Karunathilake , Nisansa de Silva , Surangika Ranathunga , Rishemjit Kaur

Natural Language Processing (NLP) plays a pivotal role in the realm of Digital Humanities (DH) and serves as the cornerstone for advancing the structural analysis of historical and cultural heritage texts. This is particularly true for the…

Computation and Language · Computer Science 2024-04-23 Xuemei Tang , Zekun Deng , Qi Su , Hao Yang , Jun Wang

Machine Transliteration provides the ability to transliterate a basic language into different languages in a computational way. Transliteration is an important technical process that has caught the attention most recently. The Sinhala…

Computation and Language · Computer Science 2024-04-23 Maneesha U. Athukorala , Deshan K. Sumanathilaka

The rapidly growing volume of scientific publications offers an interesting challenge for research on methods for analyzing the authorship of documents with one or more authors. However, most existing datasets lack scientific documents or…

Computation and Language · Computer Science 2023-05-11 Janek Bevendorff , Philipp Sauer , Lukas Gienapp , Wolfgang Kircheis , Erik Körner , Benno Stein , Martin Potthast

This paper presents a collection of highly comparable web corpora of Slovenian, Croatian, Bosnian, Montenegrin, Serbian, Macedonian, and Bulgarian, covering thereby the whole spectrum of official languages in the South Slavic language…

Computation and Language · Computer Science 2024-05-28 Nikola Ljubešić , Taja Kuzman

Objective: To build a comprehensive corpus covering syntactic and semantic annotations of Chinese clinical texts with corresponding annotation guidelines and methods as well as to develop tools trained on the annotated corpus, which…

Computation and Language · Computer Science 2016-11-09 Bin He , Bin Dong , Yi Guan , Jinfeng Yang , Zhipeng Jiang , Qiubin Yu , Jianyi Cheng , Chunyan Qu

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We create news article…

Computation and Language · Computer Science 2020-05-04 Anoop Kunchukuttan , Divyanshu Kakwani , Satish Golla , Gokul N. C. , Avik Bhattacharyya , Mitesh M. Khapra , Pratyush Kumar

The amount of scholarly data has been increasing dramatically over the last years. For newcomers to a particular science domain (e.g., IR, physics, NLP) it is often difficult to spot larger trends and to position the latest research in the…

Digital Libraries · Computer Science 2021-12-08 Naman Paharia , Muhammad Syafiq Mohd Pozi , Adam Jatowt

Chinese dynastic histories form a large continuous linguistic space of approximately 2000 years, from the 3rd century BCE to the 18th century CE. The histories are documented in Classical (Literary) Chinese in a corpus of over 20 million…

Computation and Language · Computer Science 2020-05-19 Sergey Zinin , Yang Xu

The Swa-bhasha Resource Hub provides a comprehensive collection of data resources and algorithms developed for Romanized Sinhala to Sinhala transliteration between 2020 and 2025. These resources have played a significant role in advancing…

This paper presents a semi-automatic approach to create a diachronic corpus of voices balanced for speaker's age, gender, and recording period, according to 32 categories (2 genders, 4 age ranges and 4 recording periods). Corpora were…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-29 Rémi Uro , David Doukhan , Albert Rilliard , Laëtitia Larcher , Anissa-Claire Adgharouamane , Marie Tahon , Antoine Laurent