中文
相关论文

相关论文: Developing a Fine-Grained Corpus for a Less-resour…

200 篇论文

Despite having a large number of speakers, the Kurdish language is among the less-resourced languages. In this work we highlight the challenges and problems in providing the required tools and techniques for processing texts written in…

信息检索 · 计算机科学 2012-12-04 Kyumars Sheykh Esmaili

Kurdish, an Indo-European language spoken by over 30 million speakers, is considered a dialect continuum and known for its diversity in language varieties. Previous studies addressing language and speech technology for Kurdish handle it in…

计算与语言 · 计算机科学 2024-03-05 Sina Ahmadi , Daban Q. Jaff , Md Mahfuz Ibn Alam , Antonios Anastasopoulos

Machine translation has been a major motivation of development in natural language processing. Despite the burgeoning achievements in creating more efficient machine translation systems thanks to deep learning methods, parallel corpora have…

计算与语言 · 计算机科学 2020-10-06 Sina Ahmadi , Hossein Hassani , Daban Q. Jaff

Machine translation is the task of translating texts from one language to another using computers. It has been one of the major tasks in natural language processing and computational linguistics and has been motivating to facilitate human…

计算与语言 · 计算机科学 2020-10-14 Sina Ahmadi , Mariam Masoud

One of the major challenges that under-represented and endangered language communities face in language technology is the lack or paucity of language data. This is also the case of the Southern varieties of the Kurdish and Laki languages…

计算与语言 · 计算机科学 2023-04-05 Sina Ahmadi , Zahra Azin , Sara Belelli , Antonios Anastasopoulos

We present KUTED, a speech-to-text translation (S2TT) dataset for Central Kurdish, derived from TED and TEDx talks. The corpus comprises 91,000 sentence pairs, including 170 hours of English audio, 1.65 million English tokens, and 1.40…

计算与语言 · 计算机科学 2026-04-02 Mohammad Mohammadamini , Daban Q. Jaff , Josep Crego , Marie Tahon , Antoine Laurent

Handwriting recognition is one of the active and challenging areas of research in the field of image processing and pattern recognition. It has many applications that include: a reading aid for visual impairment, automated reading and…

计算机视觉与模式识别 · 计算机科学 2022-10-26 Rebin M. Ahmed , Tarik A. Rashid , Polla Fattah , Abeer Alsadoon , Nebojsa Bacanin , Seyedali Mirjalili , S. Vimal , Amit Chhabra

Named Entity Recognition (NER) is one of the essential applications of Natural Language Processing (NLP). It is also an instrument that plays a significant role in many other NLP applications, such as Machine Translation (MT), Information…

计算与语言 · 计算机科学 2023-01-13 Sazan Salar , Hossein Hassani

This work contributes towards balancing the inclusivity and global applicability of natural language processing techniques by proposing the first 'name entity recognition' dataset for Kurdish Sorani, a low-resource and under-represented…

计算与语言 · 计算机科学 2025-12-01 Bakhtawar Abdalla , Rebwar Mala Nabi , Hassan Eshkiki , Fabio Caraffini

Segmentation is a fundamental step for most Natural Language Processing tasks. The Kurdish language is a multi-dialect, under-resourced language which is written in different scripts. The lack of various segmented corpora is one of the…

计算与语言 · 计算机科学 2020-05-01 Roshna Omer Abdulrahman , Hossein Hassani

Access to Kurdish medicine brochures is limited, depriving Kurdish-speaking communities of critical health information. To address this problem, we developed a specialized Machine Translation (MT) model to translate English medicine…

计算与语言 · 计算机科学 2025-01-24 Mariam Shamal , Hossein Hassani

The processing of the Arabic language is a complex field of research. This is due to many factors, including the complex and rich morphology of Arabic, its high degree of ambiguity, and the presence of several regional varieties that need…

计算与语言 · 计算机科学 2022-05-20 Karim El Haff , Mustafa Jarrar , Tymaa Hammouda , Fadi Zaraket

In this paper, we introduce the first large vocabulary speech recognition system (LVSR) for the Central Kurdish language, named Jira. The Kurdish language is an Indo-European language spoken by more than 30 million people in several…

人工智能 · 计算机科学 2021-02-16 Hadi Veisi , Hawre Hosseini , Mohammad Mohammadamini , Wirya Fathy , Aso Mahmudi

Tagged corpora play a crucial role in a wide range of Natural Language Processing. The Part of Speech Tagging (POST) is essential in developing tagged corpora. It is time-and-effort-consuming and costly, and therefore, it could be more…

计算与语言 · 计算机科学 2022-02-01 Hossein Hassani

Semantic Textual Similarity (STS) measures the degree of meaning overlap between two texts and underpins many NLP tasks. While extensive resources exist for high-resource languages, low-resource languages such as Kurdish remain underserved.…

计算与语言 · 计算机科学 2025-12-01 Abdulhady Abas Abdullah , Hadi Veisi , Hussein M. Al

Morphological analysis is the study of the formation and structure of words. It plays a crucial role in various tasks in Natural Language Processing (NLP) and Computational Linguistics (CL) such as machine translation and text and speech…

计算与语言 · 计算机科学 2020-05-22 Sina Ahmadi , Hossein Hassani

In this article, we present a rule-based approach for transliterating two mostly used orthographies in Sorani Kurdish. Our work consists of detecting a character in a word by removing the possible ambiguities and mapping it into the target…

计算与语言 · 计算机科学 2018-11-27 Sina Ahmadi

We present an experimental dataset, Basic Dataset for Sorani Kurdish Automatic Speech Recognition (BD-4SK-ASR), which we used in the first attempt in developing an automatic speech recognition for Sorani Kurdish. The objective of the…

计算与语言 · 计算机科学 2019-12-03 Akam Qader , Hossein Hassani

Idiom detection using Natural Language Processing (NLP) is the computerized process of recognizing figurative expressions within a text that convey meanings beyond the literal interpretation of the words. While idiom detection has seen…

计算与语言 · 计算机科学 2025-08-19 Skala Kamaran Omer , Hossein Hassani

We present an open-source speech corpus for the Kazakh language. The Kazakh speech corpus (KSC) contains around 332 hours of transcribed audio comprising over 153,000 utterances spoken by participants from different regions and age groups,…

音频与语音处理 · 电气工程与系统科学 2021-07-22 Yerbolat Khassanov , Saida Mussakhojayeva , Almas Mirzakhmetov , Alen Adiyev , Mukhamet Nurpeiissov , Huseyin Atakan Varol
‹ 上一页 1 2 3 10 下一页 ›