English
Related papers

Related papers: Evaluating Persian Tokenizers

200 papers

In this paper, we present a novel method for detecting negative words in Persian. We first used an algorithm to an exceptions list which was later modified by hand. We then used the mentioned lists and a Persian polarity corpus in our rule…

Computation and Language · Computer Science 2017-08-23 Reza Takhshid , Adel Rahimi

Lexers and parsers are typically defined separately and connected by a token stream. This separate definition is important for modularity and reduces the potential for parsing ambiguity. However, materializing tokens as data structures and…

Programming Languages · Computer Science 2023-04-17 Jeremy Yallop , Ningning Xie , Neel Krishnaswami

Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech-language modelling, and low-resource speech processing. We introduce ParsVoice, the largest…

Sound · Computer Science 2026-05-27 Mohammad Javad Ranjbar Kalahroodi , Heshaam Faili , Azadeh Shakery

In this work, we provide a systematic and comprehensive empirical comparison of pretrained multilingual language models versus their monolingual counterparts with regard to their monolingual task performance. We study a set of nine…

Computation and Language · Computer Science 2021-06-03 Phillip Rust , Jonas Pfeiffer , Ivan Vulić , Sebastian Ruder , Iryna Gurevych

Large Language Models (LLMs) have achieved remarkable performance on a wide range of Natural Language Processing (NLP) benchmarks, often surpassing human-level accuracy. However, their reliability in high-stakes domains such as medicine,…

TTS, or text-to-speech, is a complicated process that can be accomplished through appropriate modeling using deep learning methods. In order to implement deep learning models, a suitable dataset is required. Since there is a scarce amount…

Computation and Language · Computer Science 2023-04-10 Mohammd Hasan Shamgholi , Vahid Saeedi , Javad Peymanfard , Leila Alhabib , Hossein Zeinali

The surge of pre-trained language models has begun a new era in the field of Natural Language Processing (NLP) by allowing us to build powerful language models. Among these models, Transformer-based models such as BERT have become…

Computation and Language · Computer Science 2021-10-12 Mehrdad Farahani , Mohammad Gharachorloo , Marzieh Farahani , Mohammad Manthouri

Large language models (LLMs) have shown superior capabilities in translating figurative language compared to neural machine translation (NMT) systems. However, the impact of different prompting methods and LLM-NMT combinations on idiom…

Computation and Language · Computer Science 2025-02-25 Sara Rezaeimanesh , Faezeh Hosseini , Yadollah Yaghoobzadeh

This work contributes towards balancing the inclusivity and global applicability of natural language processing techniques by proposing the first 'name entity recognition' dataset for Kurdish Sorani, a low-resource and under-represented…

Computation and Language · Computer Science 2025-12-01 Bakhtawar Abdalla , Rebwar Mala Nabi , Hassan Eshkiki , Fabio Caraffini

Speech tokenization is crucial in digital speech processing, converting continuous speech signals into discrete units for various computational tasks. This paper introduces a novel speech tokenizer with broad applicability across downstream…

Machine Learning · Computer Science 2025-07-10 Wonjin Jung , Sungil Kang , Dong-Yeon Cho

Part-of-speech (POS) tagging is considered as one of the basic but necessary tools which are required for many Natural Language Processing (NLP) applications such as word sense disambiguation, information retrieval, information processing,…

Computation and Language · Computer Science 2020-01-13 Ibrahim Gashaw , H L. Shashirekha

Language recognition has been significantly advanced in recent years by means of modern machine learning methods such as deep learning and benchmarks with rich annotations. However, research is still limited in low-resource formal…

Computation and Language · Computer Science 2020-06-03 Hadi Abdi Khojasteh , Ebrahim Ansari , Mahdi Bohlouli

This research introduces the first large-scale, well-balanced Persian social media text classification dataset, specifically designed to address the lack of comprehensive resources in this domain. The dataset comprises 36,000 posts across…

Computation and Language · Computer Science 2026-05-26 Isun Chehreh , Ebrahim Ansari

Given Farsi's speaker base of over 127 million people and the growing availability of digital text, including more than 1.3 million articles on Wikipedia, it is considered a middle-resource language. However, this label quickly crumbles…

Computation and Language · Computer Science 2025-09-09 Donya Rooein , Flor Miriam Plaza-del-Arco , Debora Nozza , Dirk Hovy

Detection of semantic contradictory sentences is one of the most challenging and fundamental issues for NLP applications such as recognition of textual entailments. Contradiction in this study includes different types of semantic…

Computation and Language · Computer Science 2021-07-06 Zeinab Rahimi , Mehrnoush ShamsFard

Sign Language Recognition (SLR) is a fast-growing field that aims to fill the communication gaps between the hearing-impaired and people without hearing loss. Existing solutions for Persian Sign Language (PSL) are limited to word-level…

Human-Computer Interaction · Computer Science 2024-06-25 Amirparsa Salmankhah , Amirreza Rajabi , Negin Kheirmand , Ali Fadaeimanesh , Amirreza Tarabkhah , Amirreza Kazemzadeh , Hamed Farbeh

This research paper presents a part-of-speech (POS) annotated dataset and tagger tool for the low-resource Uzbek language. The dataset includes 12 tags, which were used to develop a rule-based POS-tagger tool. The corpus text used in the…

Computation and Language · Computer Science 2023-03-02 Maksud Sharipov , Elmurod Kuriyozov , Ollabergan Yuldashev , Ogabek Sobirov

Spell checking and morphological analysis are two fundamental tasks in text and natural language processing and are addressed in the early stages of the development of language technology. Despite the previous efforts, there is no progress…

Computation and Language · Computer Science 2021-09-15 Sina Ahmadi

Today, Social networks such as Twitter are the most widely used platforms for communication of people. Analyzing this data has useful information to recognize the opinion of people in tweets. Sentiment analysis plays a vital role in NLP,…

Computation and Language · Computer Science 2026-02-25 Seyed Himan Ghaderi , Saeed Sarbazi Azad , Mohammad Mehdi Jaziriyan , Ahmad Akbari

Sentiment analysis refers to the use of natural language processing to identify and extract subjective information from textual resources. One approach for sentiment extraction is using a sentiment lexicon. A sentiment lexicon is a set of…

Computation and Language · Computer Science 2019-11-14 Behnam Sabeti , Pedram Hosseini , Gholamreza Ghassem-Sani , Seyed Abolghasem Mirroshandel