中文
相关论文

相关论文: Evaluating Persian Tokenizers

200 篇论文

Persian poses unique audio understanding challenges through its classical poetry, traditional music, and pervasive code-switching - none captured by existing benchmarks. We introduce PARSA-Bench (Persian Audio Reasoning and Speech…

计算与语言 · 计算机科学 2026-03-17 Mohammad Javad Ranjbar Kalahroodi , Mohammad Amini , Parmis Bathayan , Heshaam Faili , Azadeh Shakery

The rapid growth in data on the internet requires a data mining process to reach a decision to support insight. The Persian language has strong potential for deep research in any aspect of natural language processing, especially sentimental…

信息检索 · 计算机科学 2020-02-26 Mohammad Heydari

Topic detection is a complex process and depends on language because it somehow needs to analyze text. There have been few studies on topic detection in Persian, and the existing algorithms are not remarkable. Therefore, we aimed to study…

计算与语言 · 计算机科学 2024-03-18 Elnaz Zafarani-Moattar , Mohammad Reza Kangavari , Amir Masoud Rahmani

In this paper, we introduce a comprehensive benchmark for Persian (Farsi) text embeddings, built upon the Massive Text Embedding Benchmark (MTEB). Our benchmark includes 63 datasets spanning seven different tasks: classification,…

计算与语言 · 计算机科学 2025-05-20 Erfan Zinvandi , Morteza Alikhani , Mehran Sarmadi , Zahra Pourbahman , Sepehr Arvin , Reza Kazemi , Arash Amini

Text processing is one of the sub-branches of natural language processing. Recently, the use of machine learning and neural networks methods has been given greater consideration. For this reason, the representation of words has become very…

计算与语言 · 计算机科学 2017-12-20 Siamak Sarmady , Erfan Rahmani

Grapheme-to-phoneme (G2P) conversion for Persian presents unique challenges due to its complex phonological features, particularly homographs and Ezafe, which exist in formal and informal language contexts. This paper introduces an…

计算与语言 · 计算机科学 2025-05-13 Abbas Bertina , Shahab Beirami , Hossein Biniazian , Elham Esmaeilnia , Soheil Shahi , Mahdi Pirnia

Background: The accuracy of spelling in Electronic Health Records (EHRs) is a critical factor for efficient clinical care, research, and ensuring patient safety. The Persian language, with its abundant vocabulary and complex…

计算与语言 · 计算机科学 2024-08-08 Seyed Mohammad Sadegh Dashti , Seyedeh Fatemeh Dashti

Pronoun resolution is a challenging subset of an essential field in natural language processing called coreference resolution. Coreference resolution is about finding all entities in the text that refers to the same real-world entity. This…

计算与语言 · 计算机科学 2022-11-14 Hassan Haji Mohammadi , Alireza Talebpour , Ahmad Mahmoudi Aznaveh , Samaneh Yazdani

Despite the progress made in recent years in addressing natural language understanding (NLU) challenges, the majority of this progress remains to be concentrated on resource-rich languages like English. This work focuses on Persian…

Social media hold valuable, vast and unstructured information on public opinion that can be utilized to improve products and services. The automatic analysis of such data, however, requires a deep understanding of natural language. Current…

计算与语言 · 计算机科学 2019-10-01 Kia Dashtipour , Mandar Gogate , Jingpeng Li , Fengling Jiang , Bin Kong , Amir Hussain

Ezafe is a grammatical particle in some Iranian languages that links two words together. Regardless of the important information it conveys, it is almost always not indicated in Persian script, resulting in mistakes in reading complex…

计算与语言 · 计算机科学 2020-10-06 Ehsan Doostmohammadi , Minoo Nassajian , Adel Rahimi

One fundamental task for NLP is to determine the similarity between two texts and evaluate the extent of their likeness. The previous methods for the Persian language have low accuracy and are unable to comprehend the structure and meaning…

计算与语言 · 计算机科学 2024-07-30 Seyed Mojtaba Sadjadi , Zeinab Rajabi , Leila Rabiei , Mohammad-Shahram Moin

The first step in any NLP pipeline is to split the text into individual tokens. The most obvious and straightforward approach is to use words as tokens. However, given a large text corpus, representing all the words is not efficient in…

计算与语言 · 计算机科学 2021-09-30 Zaid Alyafeai , Maged S. Al-shaibani , Mustafa Ghaleb , Irfan Ahmad

In the recent decade, with the enormous growth of digital content in internet and databases, sentiment analysis has received more and more attention between information retrieval and natural language processing researchers. Sentiment…

计算与语言 · 计算机科学 2014-12-30 Ayoub Bagheri , Mohamad Saraee

Undoubtedly, one of the most important issues in computer science is intelligent speech recognition. In these systems, computers try to detect and respond to the speeches they are listening to, like humans. In this research, presenting of a…

声音 · 计算机科学 2019-01-16 Saber Malekzadeh

Tokenization is a fundamental preprocessing step in NLP, directly impacting large language models' (LLMs) ability to capture syntactic, morphosyntactic, and semantic structures. This paper introduces a novel framework for systematically…

计算与语言 · 计算机科学 2025-07-22 M. Ali Bayram , Ali Arda Fincan , Ahmet Semih Gümüş , Sercan Karakaş , Banu Diri , Savaş Yıldırım

In this paper, a novel hierarchical Persian stemming approach based on the Part-Of-Speech of the word in a sentence is presented. The implemented stemmer includes hash tables and several deterministic finite automata in its different levels…

计算与语言 · 计算机科学 2014-03-13 Ayshe Rashidi , Mina Zolfy Lighvan

Coreference resolution, critical for identifying textual entities referencing the same entity, faces challenges in pronoun resolution, particularly identifying pronoun antecedents. Existing methods often treat pronoun resolution as a…

计算与语言 · 计算机科学 2024-05-20 Hassan Haji Mohammadi , Alireza Talebpour , Ahmad Mahmoudi Aznaveh , Samaneh Yazdani

In this paper we explore the problem of document summarization in Persian language from two distinct angles. In our first approach, we modify a popular and widely cited Persian document summarization framework to see how it works on a…

计算与语言 · 计算机科学 2016-06-13 Saeid Parvandeh , Shibamouli Lahiri , Fahimeh Boroumand

Toxic language detection is crucial for creating safer online environments and limiting the spread of harmful content. While toxic language detection has been under-explored in Persian, the current work compares different methods for this…

计算与语言 · 计算机科学 2025-06-05 Zahra Bokaei , Walid Magdy , Bonnie Webber