English
Related papers

Related papers: Developing a Fine-Grained Corpus for a Less-resour…

200 papers

This article describes the MyST corpus developed as part of the My Science Tutor project -- one of the largest collections of children's conversational speech comprising approximately 400 hours, spanning some 230K utterances across about…

Computation and Language · Computer Science 2023-09-26 Sameer S. Pradhan , Ronald A. Cole , Wayne H. Ward

The present paper aims at presenting a lemmatization and a word-level error correction system for Sorani Kurdish. We propose a hybrid approach based on the morphological rules and a n-gram language model. We have called our lemmatization…

Computation and Language · Computer Science 2018-10-01 Shahin Salavati , Sina Ahmadi

The recognition of cursive script is regarded as a subtle task in optical character recognition due to its varied representation. Every cursive script has different nature and associated challenges. As Urdu is one of cursive language that…

Computer Vision and Pattern Recognition · Computer Science 2017-05-17 Saad Bin Ahmed , Saeeda Naz , Salahuddin Swati , Muhammad Imran Razzak

The number of open source language models that can produce Turkish is increasing day by day, as in other languages. In order to create the basic versions of such models, the training of multilingual models is usually continued with Turkish…

Computation and Language · Computer Science 2024-04-29 H. Toprak Kesgin , M. Kaan Yuce , Eren Dogan , M. Egemen Uzun , Atahan Uz , H. Emre Seyrek , Ahmed Zeer , M. Fatih Amasyali

This paper describes the design and development of CUCHILD, a large-scale Cantonese corpus of child speech. The corpus contains spoken words collected from 1,986 child speakers aged from 3 to 6 years old. The speech materials include 130…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-10 Si-Ioi Ng , Cymie Wing-Yee Ng , Jiarui Wang , Tan Lee , Kathy Yuet-Sheung Lee , Michael Chi-Fai Tong

Neural Machine Translation (NMT) models are typically trained on datasets with limited exposure to Scientific, Technical and Educational domains. Translation models thus, in general, struggle with tasks that involve scientific understanding…

Computation and Language · Computer Science 2024-12-13 Advait Joglekar , Srinivasan Umesh

Pronoun resolution is a challenging subset of an essential field in natural language processing called coreference resolution. Coreference resolution is about finding all entities in the text that refers to the same real-world entity. This…

Computation and Language · Computer Science 2022-11-14 Hassan Haji Mohammadi , Alireza Talebpour , Ahmad Mahmoudi Aznaveh , Samaneh Yazdani

The primary obstacle to developing technologies for low-resource languages is the lack of representative, usable data. In this paper, we report the deployment of technology-driven data collection methods for creating a corpus of more than…

All natural language processing systems (such as parsers, generators, taggers) need to have access to a lexicon about the words in the language. This thesis presents a lexicon architecture for natural language processing in Turkish. Given a…

cmp-lg · Computer Science 2008-02-03 Abdullah Kurtulus Yorulmaz

Automatic Speech Recognition (ASR) technology has witnessed significant advancements in recent years, revolutionizing human-computer interactions. While major languages have benefited from these developments, lesser-resourced languages like…

Computation and Language · Computer Science 2024-11-25 Muhammad Sharif , Zeeshan Abbas , Jiangyan Yi , Chenglin Liu

The development of speech technologies for languages with limited digital representation poses significant challenges, primarily due to the scarcity of available data. This issue is exacerbated in the era of large, data-intensive models.…

Computation and Language · Computer Science 2024-06-24 Georgios Paraskevopoulos , Chara Tsoukala , Athanasios Katsamanis , Vassilis Katsouros

The Serbian language is a Slavic language spoken by over 12 million speakers and well understood by over 15 million people. In the area of natural language processing, it can be considered a low-resourced language. Also, Serbian is…

Computation and Language · Computer Science 2023-04-13 Ulfeta A. Marovac , Aldina R. Avdić , Nikola Lj. Milošević

This study aims to determine the appropriate size of the Mongolian general corpus. This study used the Heaps function and Type Token Ratio to determine the appropriate size of the Mongolian general corpus. The sample corpus of 906,064…

Computation and Language · Computer Science 2023-07-13 Sunsoo Choi , Ganbat Tsend

In this article, we have introduced the first parallel corpus of Persian with more than 10 other European languages. This article describes primary steps toward preparing a Basic Language Resources Kit (BLARK) for Persian. Up to now, we…

Computation and Language · Computer Science 2014-04-18 Behrang Qasemizadeh , Saeed Rahimi , Behrooz Mahmoodi Bakhtiari

For any deep computational processing of language we need evidences, and one such set of evidences is corpus. This paper describes the development of a text-based corpus for the Bishnupriya Manipuri language. A Corpus is considered as a…

Computation and Language · Computer Science 2013-12-12 Nayan Jyoti Kalita , Navanath Saharia , Smriti Kumar Sinha

The research of knowledge-driven conversational systems is largely limited due to the lack of dialog data which consist of multi-turn conversations on multiple topics and with knowledge annotations. In this paper, we propose a Chinese…

Computation and Language · Computer Science 2020-04-09 Hao Zhou , Chujie Zheng , Kaili Huang , Minlie Huang , Xiaoyan Zhu

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other…

The quality and accessibility of multilingual datasets are crucial for advancing machine translation. However, previous corpora built from United Nations documents have suffered from issues such as opaque process, difficulty of…

Computation and Language · Computer Science 2025-09-22 Qiuyang Lu , Fangjian Shen , Zhengkai Tang , Qiang Liu , Hexuan Cheng , Hui Liu , Wushao Wen

We compiled a new sentence splitting corpus that is composed of 203K pairs of aligned complex source and simplified target sentences. Contrary to previously proposed text simplification corpora, which contain only a small number of split…

Computation and Language · Computer Science 2019-09-27 Christina Niklaus , Andre Freitas , Siegfried Handschuh

Large Language Models (LLMs) demonstrate exceptional zero-shot capabilities in various NLP tasks, significantly enhancing user experience and efficiency. However, this advantage is primarily limited to resource-rich languages. For the…

Computation and Language · Computer Science 2025-09-23 Wenhao Zhuang , Yuan Sun