中文
相关论文

相关论文: SiDiaC: Sinhala Diachronic Corpus

200 篇论文

In this work, we introduce the MOldavian and ROmanian Dialectal COrpus (MOROCO), which is freely available for download at https://github.com/butnaruandrei/MOROCO. The corpus contains 33564 samples of text (with over 10 million tokens)…

计算与语言 · 计算机科学 2019-06-04 Andrei M. Butnaru , Radu Tudor Ionescu

This study demonstrates how hybrid neural-symbolic methods can yield significant new insights into the evolution of a morphologically rich, low-resource language. We challenge the naive assumption that linguistic change is simplification by…

计算与语言 · 计算机科学 2025-12-08 Ananth Hariharan , David Mortensen

Kurdish is a less-resourced language consisting of different dialects written in various scripts. Approximately 30 million people in different countries speak the language. The lack of corpora is one of the main obstacles in Kurdish…

计算与语言 · 计算机科学 2019-09-26 Roshna Omer Abdulrahman , Hossein Hassani , Sina Ahmadi

Dyslexia in adults remains an under-researched and under-served area, particularly in non-English-speaking contexts, despite its significant impact on personal and professional lives. This work addresses that gap by focusing on Sinhala, a…

计算与语言 · 计算机科学 2025-10-07 Peshala Perera , Deshan Sumanathilaka

We present spINAch, a large diachronic corpus of French speech from radio and television archives, balanced by speakers' gender, age (20-95 years old), and spanning 60 years from 1955 to 2015. The dataset includes over 320 hours of…

The Ramayana is among the most influential literary traditions of South and Southeast Asia, transmitted across numerous linguistic and cultural contexts over two millennia. Despite extensive scholarship on regional Ramayana traditions,…

计算与语言 · 计算机科学 2026-04-16 Sumesh VP

The Facebook network allows its users to record their reactions to text via a typology of emotions. This network, taken at scale, is therefore a prime data set of annotated sentiment data. This paper uses millions of such reactions, derived…

机器学习 · 计算机科学 2022-08-04 Vihanga Jayawickrama , Gihan Weeraprameshwara , Nisansa de Silva , Yudhanjaya Wijeratne

This paper presents first benchmark corpus of Sanskrit Pratyaya (suffix) and inflectional words (padas) formed due to suffixes along with neural network based approaches to process the formation and splitting of inflectional words.…

计算与语言 · 计算机科学 2024-09-05 Arun Kumar Singh , Sushant Dave , Prathosh A. P. , Brejesh Lall , Shresth Mehta

Hausa language belongs to the Afroasiatic phylum, and with more first-language speakers than any other sub-Saharan African language. With a majority of its speakers residing in the Northern and Southern areas of Nigeria and the Republic of…

计算与语言 · 计算机科学 2021-02-18 Isa Inuwa-Dutse

We present a novel, open-access dataset designed for semantic layout analysis, built to support document recreation workflows through mapping with the Text Encoding Initiative (TEI) standard. This dataset includes 7,254 annotated pages…

计算机视觉与模式识别 · 计算机科学 2024-11-18 Thibault Clérice , Juliette Janes , Hugo Scheithauer , Sarah Bénière , Florian Cafiero , Laurent Romary , Simon Gabay , Benoît Sagot

A Kannada OCR, named Lipi Gnani, has been designed and developed from scratch, with the motivation of it being able to convert printed text or poetry in Kannada script, without any restriction on vocabulary. The training and test sets have…

计算机视觉与模式识别 · 计算机科学 2019-01-03 Shiva Kumar H R , Ramakrishnan A G

Figures of Speech (FoS) consist of multi-word phrases that are deeply intertwined with culture. While Neural Machine Translation (NMT) performs relatively well with the figurative expressions of high-resource languages, it often faces…

计算与语言 · 计算机科学 2026-02-11 Johan Sofalas , Dilushri Pavithra , Nevidu Jayatilleke , Ruvan Weerasinghe

We introduced the contemporary Amharic corpus, which is automatically tagged for morpho-syntactic information. Texts are collected from 25,199 documents from different domains and about 24 million orthographic words are tokenized. Since it…

计算与语言 · 计算机科学 2021-06-15 Andargachew Mekonnen Gezmu , Binyam Ephrem Seyoum , Michael Gasser , Andreas Nürnberger

This research investigates the area of Music Information Retrieval (MIR) and Music Emotion Recognition (MER) in relation to Sinhala songs, an underexplored field in music studies. The purpose of this study is to analyze the behavior of…

计算与语言 · 计算机科学 2025-02-03 W. M. Yomal De Mel , Nisansa de Silva

This paper introduces TRIDIS (Tria Digita Scribunt), an open-source corpus of medieval and early modern manuscripts. TRIDIS aggregates multiple legacy collections (all published under open licenses) and incorporates large metadata…

计算与语言 · 计算机科学 2025-04-18 Sergio Torres Aguilar

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples,…

The ACL Anthology (AA) is a digital repository of tens of thousands of articles on Natural Language Processing (NLP). This paper examines the literature as a whole to identify broad trends in productivity, focus, and impact. It presents the…

数字图书馆 · 计算机科学 2019-11-12 Saif M. Mohammad

This paper presents a pipeline designed to transform raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divided into two major phases. The first involves extracting and cleaning text from raw…

计算与语言 · 计算机科学 2026-05-18 Mihailo Škorić , Cosimo Palma

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognized as scheduled languages in the Indian Constitution. Despite…

Understanding stories is a challenging reading comprehension problem for machines as it requires reading a large volume of text and following long-range dependencies. In this paper, we introduce the Shmoop Corpus: a dataset of 231 stories…

计算与语言 · 计算机科学 2020-01-03 Atef Chaudhury , Makarand Tapaswi , Seung Wook Kim , Sanja Fidler