English
Related papers

Related papers: Developing a Fine-Grained Corpus for a Less-resour…

200 papers

Depression is a common mental health condition that can lead to hopelessness, loss of interest, self-harm, and even suicide. Early detection is challenging due to individuals not self-reporting or seeking timely clinical help. With the rise…

Computation and Language · Computer Science 2025-08-25 Idrees Mohammed , Hossein Hassani

Automatic Speech Recognition (ASR) for low-resource languages remains a challenging task due to limited training data. This paper introduces a comprehensive study exploring the effectiveness of Whisper, a pre-trained ASR model, for Northern…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-23 Abdulhady Abas Abdullah , Shima Tabibian , Hadi Veisi , Aso Mahmudi , Tarik Rashid

The rapid spread of misinformation through social media platforms has raised concerns regarding its impact on public opinion. While misinformation is prevalent in other languages, the majority of research in this field has concentrated on…

Computation and Language · Computer Science 2024-03-25 Recep Firat Cekinel , Pinar Karagoz , Cagri Coltekin

Urdu is a challenging language because of, first, its Perso-Arabic script and second, its morphological system having inherent grammatical forms and vocabulary of Arabic, Persian and the native languages of South Asia. This paper describes…

Computation and Language · Computer Science 2022-04-08 Muhammad Humayoun , Harald Hammarström , Aarne Ranta

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Tingwei Guo , Cheng Wen , Dongwei Jiang , Ne Luo , Ruixiong Zhang , Shuaijiang Zhao , Wubo Li , Cheng Gong , Wei Zou , Kun Han , Xiangang Li

Semantic knowledge can be a great asset to natural language processing systems, but it is usually hand-coded for each application. Although some semantic information is available in general-purpose knowledge bases such as WordNet and Cyc,…

cmp-lg · Computer Science 2008-02-03 Ellen Riloff , Jessica Shepherd

The history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remained largely unexplored…

Computation and Language · Computer Science 2026-05-04 Seyoung Song , Nawon Kim , Songeun Chae , Kiwoong Park , Jiho Jin , Haneul Yoo , Kyunghyun Cho , Alice Oh

Many European languages possess rich biblical translation histories, yet existing corpora - in prioritizing linguistic breadth - often fail to capture this depth. To address this gap, we introduce a multilingual corpus of 651 New Testament…

Computation and Language · Computer Science 2026-05-14 Maciej Rapacz , Aleksander Smywiński-Pohl

We present the Knesset Corpus, a corpus of Hebrew parliamentary proceedings containing over 30 million sentences (over 384 million tokens) from all the (plenary and committee) protocols held in the Israeli parliament between 1998 and 2022.…

Computation and Language · Computer Science 2025-06-02 Gili Goldin , Nick Howell , Noam Ordan , Ella Rabinovich , Shuly Wintner

Document parsing is now widely used in applications, such as large-scale document digitization, retrieval-augmented generation, and domain-specific pipelines in healthcare and education. Benchmarking these models is crucial for assessing…

Computation and Language · Computer Science 2026-02-04 Deniz Yılmaz , Evren Ayberk Munis , Çağrı Toraman , Süha Kağan Köse , Burak Aktaş , Mehmet Can Baytekin , Bilge Kaan Görür

Tokenization shapes how language models perceive morphology and meaning in NLP, yet widely used frequency-driven subword tokenizers (e.g., Byte Pair Encoding and WordPiece) can fragment morphologically rich and agglutinative languages in…

Computation and Language · Computer Science 2026-04-01 M. Ali Bayram , Ali Arda Fincan , Ahmet Semih Gümüş , Sercan Karakaş , Banu Diri , Savaş Yıldırım , Demircan Çelik

Most low-resource languages do not have the necessary resources to create even a substantial monolingual corpus. These languages may often be found in government proceedings but mainly in Portable Document Format (PDF) that contains legacy…

Computation and Language · Computer Science 2022-12-19 Charangan Vasantharajan , Laksika Tharmalingam , Uthayasanker Thayasivam

Representing words and phrases into dense vectors of real numbers which encode semantic and syntactic properties is a vital constituent in natural language processing (NLP). The success of neural network (NN) models in NLP largely rely on…

Computation and Language · Computer Science 2021-01-01 Wazir Ali , Jay Kumar , Junyu Lu , Zenglin Xu

Arabic Optical Character Recognition (OCR) is essential for converting vast amounts of Arabic print media into digital formats. However, training modern OCR models, especially powerful vision-language models, is hampered by the lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Omer Nacar , Yasser Al-Habashi , Serry Sibaee , Adel Ammar , Wadii Boulila

The primary obstacle to developing technologies for low-resource languages is the lack of usable data. In this paper, we report the adoption and deployment of 4 technology-driven methods of data collection for Gondi, a low-resource…

Toxic language detection is crucial for creating safer online environments and limiting the spread of harmful content. While toxic language detection has been under-explored in Persian, the current work compares different methods for this…

Computation and Language · Computer Science 2025-06-05 Zahra Bokaei , Walid Magdy , Bonnie Webber

Keyphrases provide an extremely dense summary of a text. Such information can be used in many Natural Language Processing tasks, such as information retrieval and text summarization. Since previous studies on Persian keyword or keyphrase…

Computation and Language · Computer Science 2020-09-28 Ehsan Doostmohammadi , Mohammad Hadi Bokaei , Hossein Sameti

To enhance the reliability and robustness of language identification (LID) and language diarization (LD) systems for heterogeneous populations and scenarios, there is a need for speech processing models to be trained on datasets that…

Folktales are linguistically very rich and culturally significant in understanding the source language. Historically, only human translation has been used for translating folklore. Therefore, the number of translated texts is very sparse,…

Computation and Language · Computer Science 2024-10-15 Olena Burda-Lassen