中文
相关论文

相关论文: SiDiaC: Sinhala Diachronic Corpus

200 篇论文

Word meaning is notoriously difficult to capture, both synchronically and diachronically. In this paper, we describe the creation of the largest resource of graded contextualized, diachronic word meaning annotation in four different…

计算与语言 · 计算机科学 2024-07-09 Dominik Schlechtweg , Nina Tahmasebi , Simon Hengchen , Haim Dubossarsky , Barbara McGillivray

Knowledge is central to human and scientific developments. Natural Language Processing (NLP) allows automated analysis and creation of knowledge. Data is a crucial NLP and machine learning ingredient. The scarcity of open datasets is a…

计算与语言 · 计算机科学 2022-10-19 Istiak Ahmad , Fahad AlQurashi , Rashid Mehmood

Solving the problem of Optical Character Recognition (OCR) on printed text for Latin and its derivative scripts can now be considered settled due to the volumes of research done on English and other High-Resourced Languages (HRL). However,…

计算与语言 · 计算机科学 2025-08-26 Nevidu Jayatilleke , Nisansa de Silva

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is…

音频与语音处理 · 电气工程与系统科学 2021-02-09 Tingwei Guo , Cheng Wen , Dongwei Jiang , Ne Luo , Ruixiong Zhang , Shuaijiang Zhao , Wubo Li , Cheng Gong , Wei Zou , Kun Han , Xiangang Li

Progress in summarizing long texts is inhibited by the lack of appropriate evaluation frameworks. When a long summary must be produced to appropriately cover the facets of that text, that summary needs to present a coherent narrative to be…

计算与语言 · 计算机科学 2022-10-31 Tanya Goyal , Junyi Jessy Li , Greg Durrett

This report presents the annotation guideline for LST20, a large-scale corpus with multiple layers of linguistic annotation for Thai language processing. Our guideline consists of five layers of linguistic annotation: word segmentation, POS…

This work introduces Itihasa, a large-scale translation dataset containing 93,000 pairs of Sanskrit shlokas and their English translations. The shlokas are extracted from two Indian epics viz., The Ramayana and The Mahabharata. We first…

计算与语言 · 计算机科学 2021-10-07 Rahul Aralikatte , Miryam de Lhoneux , Anoop Kunchukuttan , Anders Søgaard

In this paper, we introduce the first and largest Hindi text corpus, named BHAAV, which means emotions in Hindi, for analyzing emotions that a writer expresses through his characters in a story, as perceived by a narrator/reader. The corpus…

计算与语言 · 计算机科学 2019-10-10 Yaman Kumar , Debanjan Mahata , Sagar Aggarwal , Anmol Chugh , Rajat Maheshwari , Rajiv Ratn Shah

We present an overview of the SCIDOCA 2025 Shared Task, which focuses on citation discovery and prediction in scientific documents. The task is divided into three subtasks: (1) Citation Discovery, where systems must identify relevant…

数字图书馆 · 计算机科学 2025-09-30 An Dao , Vu Tran , Le-Minh Nguyen , Yuji Matsumoto

MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual dataset we have built for the WSDM 2023 Cup challenge that focuses on ad hoc retrieval across 18 different languages, which collectively encompass…

We present the Multilingual Cloud Corpus, the first national-scale, parallel, multimodal linguistic dataset of Bangladesh's ethnic and indigenous languages. Despite being home to approximately 40 minority languages spanning four language…

计算与语言 · 计算机科学 2026-03-09 Mohammad Mamun Or Rashid

This paper is an effort to complement the contributions made by researchers working toward the inclusion of non-English languages in natural language processing studies. Two novel Hindi language resources have been created and released for…

计算与语言 · 计算机科学 2022-08-02 Gayatri Venugopal-Wairagade , Jatinderkumar R. Saini , Dhanya Pramod

In this work, we present to the NLP community, and to the wider research community as a whole, an application for the diachronic analysis of research corpora. We open source an easy-to-use tool coined: DRIFT, which allows researchers to…

计算与语言 · 计算机科学 2021-09-13 Abheesht Sharma , Gunjan Chhablani , Harshit Pandey , Rajaswa Patil

This article presents a hybrid methodology for building a multilingual corpus designed to support the study of emerging concepts in the humanities and social sciences (HSS), illustrated here through the case of ``non-technological…

计算与语言 · 计算机科学 2025-12-09 Revekka Kyriakoglou , Anna Pappa

Many languages have vast amounts of handwritten texts, such as ancient scripts about folktale stories and historical narratives or contemporary documents and letters. Digitization of those texts has various applications, such as daily…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Ameer Majeed , Hossein Hassani

Despite 230 million speakers, Urdu remains critically under-resourced in speech technology. We introduce UrduSpeech: a large high-fidelity Urdu corpus comprising 156 hours of audio with 12-dimension paralinguistic metadata, encompassing…

音频与语音处理 · 电气工程与系统科学 2026-05-19 Attia Nafees ul Haq , Zeyu Zhu , Jingbin Hu , ChunJiang He , Lei Xie

It is now a common practice to compare models of human language processing by predicting participant reactions (such as reading times) to corpora consisting of rich naturalistic linguistic materials. However, many of the corpora used in…

In this paper, we present a scientific corpus of abstracts of academic papers in English -- Leicester Scientific Corpus (LSC). The LSC contains 1,673,824 abstracts of research articles and proceeding papers indexed by Web of Science (WoS)…

计算与语言 · 计算机科学 2019-12-17 Neslihan Suzen , Evgeny M. Mirkes , Alexander N. Gorban

Handwritten character recognition is a challenging research in the field of document image analysis over many decades due to numerous reasons such as large writing styles variation, inherent noise in data, expansive applications it offers,…

计算机视觉与模式识别 · 计算机科学 2021-07-21 Noushath Shaffi , Faizal Hajamohideen

Vietnamese exhibits extensive dialectal variation, posing challenges for NLP systems trained predominantly on standard Vietnamese. Such systems often underperform on dialectal inputs, especially from underrepresented Central and Southern…

计算与语言 · 计算机科学 2026-03-12 Khoa Anh Ta , Nguyen Van Dinh , Kiet Van Nguyen