中文
相关论文

相关论文: Using Punkt for Sentence Segmentation in non-Latin…

200 篇论文

Idiom detection using Natural Language Processing (NLP) is the computerized process of recognizing figurative expressions within a text that convey meanings beyond the literal interpretation of the words. While idiom detection has seen…

计算与语言 · 计算机科学 2025-08-19 Skala Kamaran Omer , Hossein Hassani

Machine translation is the task of translating texts from one language to another using computers. It has been one of the major tasks in natural language processing and computational linguistics and has been motivating to facilitate human…

计算与语言 · 计算机科学 2020-10-14 Sina Ahmadi , Mariam Masoud

In this article, we present a rule-based approach for transliterating two mostly used orthographies in Sorani Kurdish. Our work consists of detecting a character in a word by removing the possible ambiguities and mapping it into the target…

计算与语言 · 计算机科学 2018-11-27 Sina Ahmadi

Many NLP pipelines split text into sentences as one of the crucial preprocessing steps. Prior sentence segmentation tools either rely on punctuation or require a considerable amount of sentence-segmented training data: both central…

计算与语言 · 计算机科学 2023-05-31 Benjamin Minixhofer , Jonas Pfeiffer , Ivan Vulić

Digit, letter and word recognition for a particular script has various applications in todays commercial contexts. Nevertheless, only a limited number of relevant studies have dealt with Persian scripts. In this paper, deep neural networks…

计算机视觉与模式识别 · 计算机科学 2020-11-17 Mehdi Bonyani , Simindokht Jahangard , Morteza Daneshmand

Handwriting recognition is one of the active and challenging areas of research in the field of image processing and pattern recognition. It has many applications that include: a reading aid for visual impairment, automated reading and…

计算机视觉与模式识别 · 计算机科学 2022-10-26 Rebin M. Ahmed , Tarik A. Rashid , Polla Fattah , Abeer Alsadoon , Nebojsa Bacanin , Seyedali Mirjalili , S. Vimal , Amit Chhabra

Words are properly segmented in the Persian writing system; in practice, however, these writing rules are often neglected, resulting in single words being written disjointedly and multiple words written without any white spaces between…

计算与语言 · 计算机科学 2020-10-29 Ehsan Doostmohammadi , Minoo Nassajian , Adel Rahimi

Kurdish is a less-resourced language consisting of different dialects written in various scripts. Approximately 30 million people in different countries speak the language. The lack of corpora is one of the main obstacles in Kurdish…

计算与语言 · 计算机科学 2019-09-26 Roshna Omer Abdulrahman , Hossein Hassani , Sina Ahmadi

We present an experimental dataset, Basic Dataset for Sorani Kurdish Automatic Speech Recognition (BD-4SK-ASR), which we used in the first attempt in developing an automatic speech recognition for Sorani Kurdish. The objective of the…

计算与语言 · 计算机科学 2019-12-03 Akam Qader , Hossein Hassani

This paper enhances the study of sentiment analysis for the Central Kurdish language by integrating the Bidirectional Encoder Representations from Transformers (BERT) into Natural Language Processing techniques. Kurdish is a low-resourced…

计算与语言 · 计算机科学 2025-09-23 Kozhin muhealddin Awlla , Hadi Veisi , Abdulhady Abas Abdullah

Semantic Textual Similarity (STS) measures the degree of meaning overlap between two texts and underpins many NLP tasks. While extensive resources exist for high-resource languages, low-resource languages such as Kurdish remain underserved.…

计算与语言 · 计算机科学 2025-12-01 Abdulhady Abas Abdullah , Hadi Veisi , Hussein M. Al

Tagged corpora play a crucial role in a wide range of Natural Language Processing. The Part of Speech Tagging (POST) is essential in developing tagged corpora. It is time-and-effort-consuming and costly, and therefore, it could be more…

计算与语言 · 计算机科学 2022-02-01 Hossein Hassani

This work contributes towards balancing the inclusivity and global applicability of natural language processing techniques by proposing the first 'name entity recognition' dataset for Kurdish Sorani, a low-resource and under-represented…

计算与语言 · 计算机科学 2025-12-01 Bakhtawar Abdalla , Rebwar Mala Nabi , Hassan Eshkiki , Fabio Caraffini

Hawrami, a dialect of Kurdish, is classified as an endangered language as it suffers from the scarcity of data and the gradual loss of its speakers. Natural Language Processing projects can be used to partially compensate for data…

计算与语言 · 计算机科学 2024-09-26 Aram Khaksar , Hossein Hassani

Word segmentation plays a pivotal role in improving any Arabic NLP application. Therefore, a lot of research has been spent in improving its accuracy. Off-the-shelf tools, however, are: i) complicated to use and ii) domain/dialect…

计算与语言 · 计算机科学 2017-09-05 Hassan Sajjad , Fahim Dalvi , Nadir Durrani , Ahmed Abdelali , Yonatan Belinkov , Stephan Vogel

Kurdish Sign Language (KuSL) is the natural language of the Kurdish Deaf people. We work on automatic translation between spoken Kurdish and KuSL. Sign languages evolve rapidly and follow grammatical rules that differ from spoken languages.…

计算与语言 · 计算机科学 2023-05-12 Zina Kamal , Hossein Hassani

Extracting concise information from scientific documents aids learners, researchers, and practitioners. Automatic Text Summarization (ATS), a key Natural Language Processing (NLP) application, automates this process. While ATS methods exist…

计算与语言 · 计算机科学 2025-04-22 Rondik Hadi Abdulrahman , Hossein Hassani

Urdu is a cursive script language and has similarities with Arabic and many other South Asian languages. Urdu is difficult to classify due to its complex geometrical and morphological structure. Character classification can be processed…

计算机视觉与模式识别 · 计算机科学 2025-08-20 Sumaiya Fazal , Sheeraz Ahmed

This research introduces a state-of-the-art Persian spelling correction system that seamlessly integrates deep learning techniques with phonetic analysis, significantly enhancing the accuracy and efficiency of natural language processing…

计算与语言 · 计算机科学 2024-07-23 Seyed Mohammad Sadegh Dashti , Amid Khatibi Bardsiri , Mehdi Jafari Shahbazzadeh

State-of-the-art Natural Language Processing algorithms rely heavily on efficient word segmentation. Urdu is amongst languages for which word segmentation is a complex task as it exhibits space omission as well as space insertion issues.…

计算与语言 · 计算机科学 2018-06-15 Haris Bin Zia , Agha Ali Raza , Awais Athar
‹ 上一页 1 2 3 10 下一页 ›