中文
相关论文

相关论文: Sri Lanka Document Datasets: A Large-Scale, Multil…

200 篇论文

SinhaLegal introduces a Sinhala legislative text corpus containing approximately 2 million words across 1,206 legal documents. The dataset includes two types of legal documents: 1,065 Acts dated from 1981 to 2014 and 141 Bills from 2010 to…

计算与语言 · 计算机科学 2026-03-06 Minduli Lasandi , Nevidu Jayatilleke

Sinhala is the native language of the Sinhalese people who make up the largest ethnic group of Sri Lanka. The language belongs to the globe-spanning language tree, Indo-European. However, due to poverty in both linguistic and economic…

计算与语言 · 计算机科学 2026-01-13 Nisansa de Silva

This paper presents two colloquial Sinhala language corpora from the language efforts of the Data, Analysis and Policy team of LIRNEasia, as well as a list of algorithmically derived stopwords. The larger of the two corpora spans 2010 to…

计算与语言 · 计算机科学 2020-07-16 Yudhanjaya Wijeratne , Nisansa de Silva

This paper presents the first-ever Sinhala physical common sense reasoning dataset created as part of Global PIQA. It contains 110 human-created and verified data samples, where each sample consists of a prompt, the corresponding correct…

计算与语言 · 计算机科学 2026-02-03 Nisansa de Silva , Surangika Ranathunga

The introduction of large language models (LLMs) has advanced natural language processing (NLP), but their effectiveness is largely dependent on pre-training resources. This is especially evident in low-resource languages, such as Sinhala,…

计算与语言 · 计算机科学 2024-03-26 Hansi Hettiarachchi , Damith Premasiri , Lasitha Uyangodage , Tharindu Ranasinghe

Text Simplification is a task that has been minimally explored for low-resource languages. Consequently, there are only a few manually curated datasets. In this paper, we present a human curated sentence-level text simplification dataset…

Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, long-context…

计算与语言 · 计算机科学 2025-10-01 Dayyán O'Brien , Bhavitvya Malik , Ona de Gibert , Pinzhen Chen , Barry Haddow , Jörg Tiedemann

SiPaKosa is a comprehensive corpus of Sinhala and Pali doctrinal texts comprising approximately 786K sentences and 9.25M words, incorporating 16 copyright-cleared historical Buddhist documents alongside the complete web-scraped Tripitaka…

计算与语言 · 计算机科学 2026-04-01 Ranidu Gurusinghe , Nevidu Jayatilleke

Large Language Models (LLMs) demonstrate impressive general knowledge and reasoning abilities, yet their evaluation has predominantly focused on global or anglocentric subjects, often neglecting low-resource languages and culturally…

Large, high-quality datasets are crucial for training Large Language Models (LLMs). However, so far, there are few datasets available for specialized critical domains such as law and the available ones are often only for the English…

计算与语言 · 计算机科学 2024-05-21 Joel Niklaus , Veton Matoshi , Matthias Stürmer , Ilias Chalkidis , Daniel E. Ho

The Swa-bhasha Resource Hub provides a comprehensive collection of data resources and algorithms developed for Romanized Sinhala to Sinhala transliteration between 2020 and 2025. These resources have played a significant role in advancing…

The widespread of offensive content online, such as hate speech and cyber-bullying, is a global phenomenon. This has sparked interest in the artificial intelligence (AI) and natural language processing (NLP) communities, motivating the…

The rapid evolution in the fields of computer science, data science, and artificial intelligence has significantly transformed the utilisation of data for decision-making. Data visualisation plays a critical role in any work that involves…

其他统计学 · 统计学 2024-01-08 Thiyanga S. Talagala

This study examines the application of artificial intelligence (AI) and large language models (LLM) to improve access to legal texts in Senegal's judicial system. The emphasis is on the difficulties of extracting and organizing legal…

计算与语言 · 计算机科学 2026-03-10 Oumar Kane , Mouhamad M. Allaya , Dame Samb , Mamadou Bousso

This paper introduces a cross-lingual statutory article retrieval (SAR) dataset designed to enhance legal information retrieval in multilingual settings. Our dataset features spoken-language-style legal inquiries in English, paired with…

计算与语言 · 计算机科学 2024-10-16 Yen-Hsiang Wang , Feng-Dian Su , Tzu-Yu Yeh , Yao-Chung Fan

The scale, variety, and quantity of publicly-available NLP datasets has grown rapidly as researchers propose new tasks, larger models, and novel benchmarks. Datasets is a community library for contemporary NLP designed to support this…

This paper presents a high-quality multilingual dataset for the documentation domain to advance research on localization of structured text. Unlike widely-used datasets for translation of plain text, we collect XML-structured parallel text…

计算与语言 · 计算机科学 2020-06-25 Kazuma Hashimoto , Raffaella Buschiazzo , James Bradbury , Teresa Marshall , Richard Socher , Caiming Xiong

Due to the high impact of the fast-evolving fields of machine learning and deep learning, Natural Language Processing (NLP) tasks have further obtained comprehensive performances for highly resourced languages such as English and Chinese.…

计算与语言 · 计算机科学 2020-11-17 Lahiru Senevirathne , Piyumal Demotte , Binod Karunanayake , Udyogi Munasinghe , Surangika Ranathunga

Parallel datasets are vital for performing and evaluating any kind of multilingual task. However, in the cases where one of the considered language pairs is a low-resource language, the existing top-down parallel data such as corpora are…

计算与语言 · 计算机科学 2023-09-26 Kasun Wickramasinghe , Nisansa de Silva
‹ 上一页 1 2 3 10 下一页 ›