中文
相关论文

相关论文: Sinhala Language Corpora and Stopwords from a Deca…

200 篇论文

SinhaLegal introduces a Sinhala legislative text corpus containing approximately 2 million words across 1,206 legal documents. The dataset includes two types of legal documents: 1,065 Acts dated from 1981 to 2014 and 141 Bills from 2010 to…

计算与语言 · 计算机科学 2026-03-06 Minduli Lasandi , Nevidu Jayatilleke

We present a collection of open, machine-readable document datasets covering parliamentary proceedings, legal judgments, government publications, news, and tourism statistics from Sri Lanka. The collection currently comprises of 269,194…

计算与语言 · 计算机科学 2026-05-18 Nuwan I. Senaratna

SiDiaC-v.2.0 is the largest comprehensive Sinhala Diachronic Corpus to date, covering a period from 1800 CE to 1955 CE in terms of publication dates, and a historical span from the 5th to the 20th century CE in terms of written dates. The…

SiDiaC, the first comprehensive Sinhala Diachronic Corpus, covers a historical span from the 5th to the 20th century CE. SiDiaC comprises 58k words across 46 literary works, annotated carefully based on the written date, after filtering…

计算与语言 · 计算机科学 2026-05-19 Nevidu Jayatilleke , Nisansa de Silva

This research investigates the area of Music Information Retrieval (MIR) and Music Emotion Recognition (MER) in relation to Sinhala songs, an underexplored field in music studies. The purpose of this study is to analyze the behavior of…

计算与语言 · 计算机科学 2025-02-03 W. M. Yomal De Mel , Nisansa de Silva

Sinhala is the native language of the Sinhalese people who make up the largest ethnic group of Sri Lanka. The language belongs to the globe-spanning language tree, Indo-European. However, due to poverty in both linguistic and economic…

计算与语言 · 计算机科学 2026-01-13 Nisansa de Silva

The introduction of large language models (LLMs) has advanced natural language processing (NLP), but their effectiveness is largely dependent on pre-training resources. This is especially evident in low-resource languages, such as Sinhala,…

计算与语言 · 计算机科学 2024-03-26 Hansi Hettiarachchi , Damith Premasiri , Lasitha Uyangodage , Tharindu Ranasinghe

The Facebook network allows its users to record their reactions to text via a typology of emotions. This network, taken at scale, is therefore a prime data set of annotated sentiment data. This paper uses millions of such reactions, derived…

机器学习 · 计算机科学 2022-08-04 Vihanga Jayawickrama , Gihan Weeraprameshwara , Nisansa de Silva , Yudhanjaya Wijeratne

Due to the high impact of the fast-evolving fields of machine learning and deep learning, Natural Language Processing (NLP) tasks have further obtained comprehensive performances for highly resourced languages such as English and Chinese.…

计算与语言 · 计算机科学 2020-11-17 Lahiru Senevirathne , Piyumal Demotte , Binod Karunanayake , Udyogi Munasinghe , Surangika Ranathunga

SiPaKosa is a comprehensive corpus of Sinhala and Pali doctrinal texts comprising approximately 786K sentences and 9.25M words, incorporating 16 copyright-cleared historical Buddhist documents alongside the complete web-scraped Tripitaka…

计算与语言 · 计算机科学 2026-04-01 Ranidu Gurusinghe , Nevidu Jayatilleke

This paper presents the first-ever Sinhala physical common sense reasoning dataset created as part of Global PIQA. It contains 110 human-created and verified data samples, where each sample consists of a prompt, the corresponding correct…

计算与语言 · 计算机科学 2026-02-03 Nisansa de Silva , Surangika Ranathunga

Crawling national top-level domains has proven to be highly effective for collecting texts in less-resourced languages. This approach has been recently used for South Slavic languages and resulted in the largest general corpora for this…

计算与语言 · 计算机科学 2026-03-02 Taja Kuzman Pungeršek , Peter Rupnik , Vít Suchomel , Nikola Ljubešić

Large Language Models (LLMs) demonstrate impressive general knowledge and reasoning abilities, yet their evaluation has predominantly focused on global or anglocentric subjects, often neglecting low-resource languages and culturally…

The Swa-bhasha Resource Hub provides a comprehensive collection of data resources and algorithms developed for Romanized Sinhala to Sinhala transliteration between 2020 and 2025. These resources have played a significant role in advancing…

Figures of Speech (FoS) consist of multi-word phrases that are deeply intertwined with culture. While Neural Machine Translation (NMT) performs relatively well with the figurative expressions of high-resource languages, it often faces…

计算与语言 · 计算机科学 2026-02-11 Johan Sofalas , Dilushri Pavithra , Nevidu Jayatilleke , Ruvan Weerasinghe

The relationship between Facebook posts and the corresponding reaction feature is an interesting subject to explore and understand. To achieve this end, we test state-of-the-art Sinhala sentiment analysis models against a data set…

计算与语言 · 计算机科学 2022-08-04 Gihan Weeraprameshwara , Vihanga Jayawickrama , Nisansa de Silva , Yudhanjaya Wijeratne

This paper presents a collection of highly comparable web corpora of Slovenian, Croatian, Bosnian, Montenegrin, Serbian, Macedonian, and Bulgarian, covering thereby the whole spectrum of official languages in the South Slavic language…

计算与语言 · 计算机科学 2024-05-28 Nikola Ljubešić , Taja Kuzman

Low-resource languages such as Sinhala are often overlooked by open-source Large Language Models (LLMs). In this research, we extend an existing multilingual LLM (Llama-3-8B) to better serve Sinhala. We enhance the LLM tokenizer with…

We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource. The corpora are…

计算与语言 · 计算机科学 2020-07-16 Shashank Siripragada , Jerin Philip , Vinay P. Namboodiri , C V Jawahar

The widespread of offensive content online, such as hate speech and cyber-bullying, is a global phenomenon. This has sparked interest in the artificial intelligence (AI) and natural language processing (NLP) communities, motivating the…

‹ 上一页 1 2 3 10 下一页 ›