English
Related papers

Related papers: PARSI: Persian Authorship Recognition via Stylomet…

200 papers

This Paper presents a method for lexicon reduction of Printed Farsi subwords based on their holistic shape features. Because of the large number of Persian subwords variously shaped from a simple letter to a complex combination of several…

Computer Vision and Pattern Recognition · Computer Science 2016-01-26 Homa Davoudi , Ehsanollah Kabir

In this paper, we propose a novel approach for measuring the degree of similarity between categories of two pieces of Persian text, which were published as descriptions of two separate advertisements. We built an appropriate dataset for…

Computation and Language · Computer Science 2019-09-27 Hossein Keshavarz , Shohreh Tabatabayi Seifi , Mohammad Izadi

Tokenization plays a significant role in the process of lexical analysis. Tokens become the input for other natural language processing tasks, like semantic parsing and language modeling. Natural Language Processing in Persian is…

Computation and Language · Computer Science 2022-02-23 Danial Kamali , Behrooz Janfada , Mohammad Ebrahim Shenasa , Behrouz Minaei-Bidgoli

This study addresses automatic transliteration from Tajik (Cyrillic script) to Persian (Perso-Arabic script). We present a curated, lexicographically verified parallel corpus of 52,152 Tajik--Persian words and short phrases, compiled from…

Computation and Language · Computer Science 2026-05-12 Mullosharaf K. Arabov

Recent advances in language models (LMs), have demonstrated significant efficacy in tasks related to the arts and humanities. While LMs have exhibited exceptional performance across a wide range of natural language processing tasks, there…

Computation and Language · Computer Science 2023-12-07 Amir Panahandeh , Hanie Asemi , Esmaeil Nourani

In this paper we introduce PerPaDa, a Persian paraphrase dataset that is collected from users' input in a plagiarism detection system. As an implicit crowdsourcing experience, we have gathered a large collection of original and paraphrased…

Computation and Language · Computer Science 2022-09-14 Salar Mohtaj , Fatemeh Tavakkoli , Habibollah Asghari

Recognizing a piece of writing as a poem or prose is usually easy for the majority of people; however, only specialists can determine which meter a poem belongs to. In this paper, we build Recurrent Neural Network (RNN) models that can…

Computation and Language · Computer Science 2019-05-15 Waleed A. Yousef , Omar M. Ibrahime , Taha M. Madbouly , Moustafa A. Mahmoud

We introduced PerCoR (Persian Commonsense Reasoning), the first large-scale Persian benchmark for commonsense reasoning. PerCoR contains 106K multiple-choice sentence-completion problems drawn from more than forty news, cultural, and other…

Computation and Language · Computer Science 2026-01-19 Morteza Alikhani , Mohammadtaha Bagherifard , Erfan Zinvandi , Mehran Sarmadi

Language recognition has been significantly advanced in recent years by means of modern machine learning methods such as deep learning and benchmarks with rich annotations. However, research is still limited in low-resource formal…

Computation and Language · Computer Science 2020-06-03 Hadi Abdi Khojasteh , Ebrahim Ansari , Mahdi Bohlouli

Stylometry, the science of inferring characteristics of the author from the characteristics of documents written by that author, is a problem with a long history and belongs to the core task of Text categorization that involves authorship…

Computation and Language · Computer Science 2012-10-16 Tanmoy Chakraborty , Sivaji Bandyopadhyay

In this paper, we introduce a comprehensive benchmark for Persian (Farsi) text embeddings, built upon the Massive Text Embedding Benchmark (MTEB). Our benchmark includes 63 datasets spanning seven different tasks: classification,…

Computation and Language · Computer Science 2025-05-20 Erfan Zinvandi , Morteza Alikhani , Mehran Sarmadi , Zahra Pourbahman , Sepehr Arvin , Reza Kazemi , Arash Amini

Automatic spelling correction stands as a pivotal challenge within the ambit of natural language processing (NLP), demanding nuanced solutions. Traditional spelling correction techniques are typically only capable of detecting and…

Computation and Language · Computer Science 2024-07-23 Seyed Mohammad Sadegh Dashti , Amid Khatibi Bardsiri , Mehdi Jafari Shahbazzadeh

Stylistic analysis of text is a key task in research areas ranging from authorship attribution to forensic analysis and personality profiling. The existing approaches for stylistic analysis are plagued by issues like topic influence, lack…

Computation and Language · Computer Science 2023-12-07 Ronald Wilson , Avanti Bhandarkar , Damon Woodard

In this article we describe an application of Machine Learning (ML) and Linguistic Modeling to generate persian poems. In fact we teach machine by reading and learning persian poems to generate fake poems in the same style of the original…

Computation and Language · Computer Science 2018-10-17 Mehdi Hosseini Moghadam , Bardia Panahbehagh

Classification Ensemble, which uses the weighed polling of outputs, is the art of combining a set of basic classifiers for generating high-performance, robust and more stable results. This study aims to improve the results of identifying…

Computer Vision and Pattern Recognition · Computer Science 2016-04-27 Maziar Kazemi , Muhammad Yousefnezhad , Saber Nourian

Automatic recognition of Urdu handwritten digits and characters, is a challenging task. It has applications in postal address reading, bank's cheque processing, and digitization and preservation of handwritten manuscripts from old ages.…

Computer Vision and Pattern Recognition · Computer Science 2019-12-18 Hazrat Ali , Ahsan Ullah , Talha Iqbal , Shahid Khattak

Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech-language modelling, and low-resource speech processing. We introduce ParsVoice, the largest…

Sound · Computer Science 2026-05-27 Mohammad Javad Ranjbar Kalahroodi , Heshaam Faili , Azadeh Shakery

In order to provide benchmark performance for Urdu text document classification, the contribution of this paper is manifold. First, it pro-vides a publicly available benchmark dataset manually tagged against 6 classes. Second, it…

Computation and Language · Computer Science 2020-03-04 Muhammad Nabeel Asim , Muhammad Usman Ghani , Muhammad Ali Ibrahim , Sheraz Ahmad , Waqar Mahmood , Andreas Dengel

Social media hold valuable, vast and unstructured information on public opinion that can be utilized to improve products and services. The automatic analysis of such data, however, requires a deep understanding of natural language. Current…

Computation and Language · Computer Science 2019-10-01 Kia Dashtipour , Mandar Gogate , Jingpeng Li , Fengling Jiang , Bin Kong , Amir Hussain

Pronoun resolution is a challenging subset of an essential field in natural language processing called coreference resolution. Coreference resolution is about finding all entities in the text that refers to the same real-world entity. This…

Computation and Language · Computer Science 2022-11-14 Hassan Haji Mohammadi , Alireza Talebpour , Ahmad Mahmoudi Aznaveh , Samaneh Yazdani