中文
相关论文

相关论文: SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.…

200 篇论文

Machine Transliteration provides the ability to transliterate a basic language into different languages in a computational way. Transliteration is an important technical process that has caught the attention most recently. The Sinhala…

计算与语言 · 计算机科学 2024-04-23 Maneesha U. Athukorala , Deshan K. Sumanathilaka

Parallel datasets are vital for performing and evaluating any kind of multilingual task. However, in the cases where one of the considered language pairs is a low-resource language, the existing top-down parallel data such as corpora are…

计算与语言 · 计算机科学 2023-09-26 Kasun Wickramasinghe , Nisansa de Silva

This research investigates the area of Music Information Retrieval (MIR) and Music Emotion Recognition (MER) in relation to Sinhala songs, an underexplored field in music studies. The purpose of this study is to analyze the behavior of…

计算与语言 · 计算机科学 2025-02-03 W. M. Yomal De Mel , Nisansa de Silva

Slovak remains a low-resource language for automatic speech recognition (ASR), with fewer than 100 hours of publicly available training data. We present SloPal, a comprehensive Slovak parliamentary corpus comprising 330,000…

计算与语言 · 计算机科学 2026-03-17 Erik Božík , Marek Šuppa

Natural Language Processing (NLP) plays a pivotal role in the realm of Digital Humanities (DH) and serves as the cornerstone for advancing the structural analysis of historical and cultural heritage texts. This is particularly true for the…

计算与语言 · 计算机科学 2024-04-23 Xuemei Tang , Zekun Deng , Qi Su , Hao Yang , Jun Wang

This technical report presents the 600K-KS-OCR Dataset, a large-scale synthetic corpus comprising approximately 602,000 word-level segmented images designed for training and evaluating optical character recognition systems targeting…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Haq Nawaz Malik

In this paper, we present a scientific corpus of abstracts of academic papers in English -- Leicester Scientific Corpus (LSC). The LSC contains 1,673,824 abstracts of research articles and proceeding papers indexed by Web of Science (WoS)…

计算与语言 · 计算机科学 2019-12-17 Neslihan Suzen , Evgeny M. Mirkes , Alexander N. Gorban

Solving the problem of Optical Character Recognition (OCR) on printed text for Latin and its derivative scripts can now be considered settled due to the volumes of research done on English and other High-Resourced Languages (HRL). However,…

计算与语言 · 计算机科学 2025-08-26 Nevidu Jayatilleke , Nisansa de Silva

In this paper, we introduce the French-YMCA corpus, a new linguistic resource specifically tailored for children and adolescents. The motivation for building this corpus is clear: children have unique language requirements, as their…

计算与语言 · 计算机科学 2026-04-08 Cherifa Ben Khelil , Jean-Yves Antoine , Anaïs Halftermeyer , Frédéric Rayar , Mathieu Thebaud

We present SciDMT, an enhanced and expanded corpus for scientific mention detection, offering a significant advancement over existing related resources. SciDMT contains annotated scientific documents for datasets (D), methods (M), and tasks…

人工智能 · 计算机科学 2024-06-24 Huitong Pan , Qi Zhang , Cornelia Caragea , Eduard Dragut , Longin Jan Latecki

This study demonstrates how hybrid neural-symbolic methods can yield significant new insights into the evolution of a morphologically rich, low-resource language. We challenge the naive assumption that linguistic change is simplification by…

计算与语言 · 计算机科学 2025-12-08 Ananth Hariharan , David Mortensen

In this article, the beta version 0.1.0 of Opera Graeca Adnotata (OGA), the largest open-access multilayer corpus for Ancient Greek (AG) is presented. OGA consists of 1,687 literary works and 34M+ tokens coming from the PerseusDL and…

计算与语言 · 计算机科学 2024-04-02 Giuseppe G. A. Celano

We present the Patrologia Graeca Corpus, the first large-scale open OCR and linguistic resource for nineteenthcentury editions of Ancient Greek. The collection covers the remaining undigitized volumes of the Patrologia Graeca (PG), printed…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Chahan Vidal-Gorène , Bastien Kindt

We introduce the Speak & Improve Corpus 2025, a dataset of L2 learner English data with holistic scores and language error annotation, collected from open (spontaneous) speaking tests on the Speak & Improve learning platform. The aim of the…

计算与语言 · 计算机科学 2024-12-18 Kate Knill , Diane Nicholls , Mark J. F. Gales , Mengjie Qian , Pawel Stroinski

This paper presents a semi-automatic approach to create a diachronic corpus of voices balanced for speaker's age, gender, and recording period, according to 32 categories (2 genders, 4 age ranges and 4 recording periods). Corpora were…

音频与语音处理 · 电气工程与系统科学 2024-04-29 Rémi Uro , David Doukhan , Albert Rilliard , Laëtitia Larcher , Anissa-Claire Adgharouamane , Marie Tahon , Antoine Laurent

Progress in summarizing long texts is inhibited by the lack of appropriate evaluation frameworks. When a long summary must be produced to appropriately cover the facets of that text, that summary needs to present a coherent narrative to be…

计算与语言 · 计算机科学 2022-10-31 Tanya Goyal , Junyi Jessy Li , Greg Durrett

Sign language is a vital communication medium for the hearing-impaired community, enabling effective interaction and self-expression. To help bridge the communication gap between hearing and hearing-impaired individuals, a text-to-sign…

人机交互 · 计算机科学 2025-11-24 MD. Ashikul Islam , Prato Dewan , Md Fuadul Islam , Md. Ataullha , M. Shahidur Rahman

Many populous countries including India are burdened with a considerable backlog of legal cases. Development of automated systems that could process legal documents and augment legal practitioners can mitigate this. However, there is a…

This paper presents first benchmark corpus of Sanskrit Pratyaya (suffix) and inflectional words (padas) formed due to suffixes along with neural network based approaches to process the formation and splitting of inflectional words.…

计算与语言 · 计算机科学 2024-09-05 Arun Kumar Singh , Sushant Dave , Prathosh A. P. , Brejesh Lall , Shresth Mehta

We present the SAMER Corpus, the first manually annotated Arabic parallel corpus for text simplification targeting school-aged learners. Our corpus comprises texts of 159K words selected from 15 publicly available Arabic fiction novels most…

计算与语言 · 计算机科学 2024-04-30 Bashar Alhafni , Reem Hazim , Juan Piñeros Liberato , Muhamed Al Khalil , Nizar Habash