中文
相关论文

相关论文: Evaluating Persian Tokenizers

200 篇论文

In this paper, we present a novel method for detecting negative words in Persian. We first used an algorithm to an exceptions list which was later modified by hand. We then used the mentioned lists and a Persian polarity corpus in our rule…

计算与语言 · 计算机科学 2017-08-23 Reza Takhshid , Adel Rahimi

Lexers and parsers are typically defined separately and connected by a token stream. This separate definition is important for modularity and reduces the potential for parsing ambiguity. However, materializing tokens as data structures and…

编程语言 · 计算机科学 2023-04-17 Jeremy Yallop , Ningning Xie , Neel Krishnaswami

Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech-language modelling, and low-resource speech processing. We introduce ParsVoice, the largest…

声音 · 计算机科学 2026-05-27 Mohammad Javad Ranjbar Kalahroodi , Heshaam Faili , Azadeh Shakery

In this work, we provide a systematic and comprehensive empirical comparison of pretrained multilingual language models versus their monolingual counterparts with regard to their monolingual task performance. We study a set of nine…

计算与语言 · 计算机科学 2021-06-03 Phillip Rust , Jonas Pfeiffer , Ivan Vulić , Sebastian Ruder , Iryna Gurevych

Large Language Models (LLMs) have achieved remarkable performance on a wide range of Natural Language Processing (NLP) benchmarks, often surpassing human-level accuracy. However, their reliability in high-stakes domains such as medicine,…

TTS, or text-to-speech, is a complicated process that can be accomplished through appropriate modeling using deep learning methods. In order to implement deep learning models, a suitable dataset is required. Since there is a scarce amount…

计算与语言 · 计算机科学 2023-04-10 Mohammd Hasan Shamgholi , Vahid Saeedi , Javad Peymanfard , Leila Alhabib , Hossein Zeinali

The surge of pre-trained language models has begun a new era in the field of Natural Language Processing (NLP) by allowing us to build powerful language models. Among these models, Transformer-based models such as BERT have become…

计算与语言 · 计算机科学 2021-10-12 Mehrdad Farahani , Mohammad Gharachorloo , Marzieh Farahani , Mohammad Manthouri

Large language models (LLMs) have shown superior capabilities in translating figurative language compared to neural machine translation (NMT) systems. However, the impact of different prompting methods and LLM-NMT combinations on idiom…

计算与语言 · 计算机科学 2025-02-25 Sara Rezaeimanesh , Faezeh Hosseini , Yadollah Yaghoobzadeh

This work contributes towards balancing the inclusivity and global applicability of natural language processing techniques by proposing the first 'name entity recognition' dataset for Kurdish Sorani, a low-resource and under-represented…

计算与语言 · 计算机科学 2025-12-01 Bakhtawar Abdalla , Rebwar Mala Nabi , Hassan Eshkiki , Fabio Caraffini

Speech tokenization is crucial in digital speech processing, converting continuous speech signals into discrete units for various computational tasks. This paper introduces a novel speech tokenizer with broad applicability across downstream…

机器学习 · 计算机科学 2025-07-10 Wonjin Jung , Sungil Kang , Dong-Yeon Cho

Part-of-speech (POS) tagging is considered as one of the basic but necessary tools which are required for many Natural Language Processing (NLP) applications such as word sense disambiguation, information retrieval, information processing,…

计算与语言 · 计算机科学 2020-01-13 Ibrahim Gashaw , H L. Shashirekha

Language recognition has been significantly advanced in recent years by means of modern machine learning methods such as deep learning and benchmarks with rich annotations. However, research is still limited in low-resource formal…

计算与语言 · 计算机科学 2020-06-03 Hadi Abdi Khojasteh , Ebrahim Ansari , Mahdi Bohlouli

This research introduces the first large-scale, well-balanced Persian social media text classification dataset, specifically designed to address the lack of comprehensive resources in this domain. The dataset comprises 36,000 posts across…

计算与语言 · 计算机科学 2026-05-26 Isun Chehreh , Ebrahim Ansari

Given Farsi's speaker base of over 127 million people and the growing availability of digital text, including more than 1.3 million articles on Wikipedia, it is considered a middle-resource language. However, this label quickly crumbles…

计算与语言 · 计算机科学 2025-09-09 Donya Rooein , Flor Miriam Plaza-del-Arco , Debora Nozza , Dirk Hovy

Detection of semantic contradictory sentences is one of the most challenging and fundamental issues for NLP applications such as recognition of textual entailments. Contradiction in this study includes different types of semantic…

计算与语言 · 计算机科学 2021-07-06 Zeinab Rahimi , Mehrnoush ShamsFard

Sign Language Recognition (SLR) is a fast-growing field that aims to fill the communication gaps between the hearing-impaired and people without hearing loss. Existing solutions for Persian Sign Language (PSL) are limited to word-level…

This research paper presents a part-of-speech (POS) annotated dataset and tagger tool for the low-resource Uzbek language. The dataset includes 12 tags, which were used to develop a rule-based POS-tagger tool. The corpus text used in the…

计算与语言 · 计算机科学 2023-03-02 Maksud Sharipov , Elmurod Kuriyozov , Ollabergan Yuldashev , Ogabek Sobirov

Spell checking and morphological analysis are two fundamental tasks in text and natural language processing and are addressed in the early stages of the development of language technology. Despite the previous efforts, there is no progress…

计算与语言 · 计算机科学 2021-09-15 Sina Ahmadi

Today, Social networks such as Twitter are the most widely used platforms for communication of people. Analyzing this data has useful information to recognize the opinion of people in tweets. Sentiment analysis plays a vital role in NLP,…

计算与语言 · 计算机科学 2026-02-25 Seyed Himan Ghaderi , Saeed Sarbazi Azad , Mohammad Mehdi Jaziriyan , Ahmad Akbari

Sentiment analysis refers to the use of natural language processing to identify and extract subjective information from textual resources. One approach for sentiment extraction is using a sentiment lexicon. A sentiment lexicon is a set of…

计算与语言 · 计算机科学 2019-11-14 Behnam Sabeti , Pedram Hosseini , Gholamreza Ghassem-Sani , Seyed Abolghasem Mirroshandel