English
Related papers

Related papers: Persian Rhetorical Structure Theory

200 papers

Large language models (LLMs) have shown remarkable progress in reasoning abilities and general natural language processing (NLP) tasks, yet their performance on Arabic data, characterized by rich morphology, diverse dialects, and complex…

Computation and Language · Computer Science 2025-12-16 Ahmed Hasanaath , Aisha Alansari , Ahmed Ashraf , Chafik Salmane , Hamzah Luqman , Saad Ezzini

Discourse-annotated corpora are an important resource for the community, but they are often annotated according to different frameworks. This makes comparison of the annotations difficult, thereby also preventing researchers from searching…

Computation and Language · Computer Science 2018-03-16 Vera Demberg , Fatemeh Torabi Asr , Merel Scholman

Despite impressive multilingual capabilities, large language models (LLMs) remain poorly evaluated on literary knowledge in non-English languages. We introduce PersLitEval, a benchmark of 4,514 Persian literature multiple-choice questions…

Computation and Language · Computer Science 2026-05-27 Ruhallah Niazi , Faeze Ghorbanpour , Alexander Fraser

Document and discourse segmentation are two fundamental NLP tasks pertaining to breaking up text into constituents, which are commonly used to help downstream tasks such as information retrieval or text summarization. In this work, we…

Computation and Language · Computer Science 2020-12-08 Michal Lukasik , Boris Dadachev , Gonçalo Simões , Kishore Papineni

The patterns in which the syntax of different languages converges and diverges are often used to inform work on cross-lingual transfer. Nevertheless, little empirical work has been done on quantifying the prevalence of different syntactic…

Computation and Language · Computer Science 2020-07-14 Dmitry Nikolaev , Ofir Arviv , Taelin Karidi , Neta Kenneth , Veronika Mitnik , Lilja Maria Saeboe , Omri Abend

The article is focused on automatic development and ranking of a large corpus for Russian paraphrase generation which proves to be the first corpus of such type in Russian computational linguistics. Existing manually annotated paraphrase…

Computation and Language · Computer Science 2020-06-18 Vadim Gudkov , Olga Mitrofanova , Elizaveta Filippskikh

Recently, there has been a growing interest in the use of deep learning techniques for tasks in natural language processing (NLP), with sentiment analysis being one of the most challenging areas, particularly in the Persian language. The…

Computation and Language · Computer Science 2024-03-19 Mohammad Heydari , Mohsen Khazeni , Mohammad Ali Soltanshahi

The proliferation of hate speech and offensive comments on social media has become increasingly prevalent due to user activities. Such comments can have detrimental effects on individuals' psychological well-being and social behavior. While…

This paper introduces the Balanced Arabic Readability Evaluation Corpus (BAREC), a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 69,441 sentences spanning 1+ million words, carefully curated to cover…

Computation and Language · Computer Science 2025-06-17 Khalid N. Elmadani , Nizar Habash , Hanada Taha-Thomure

We present PashtoCorp, a 1.25-billion-word corpus for Pashto, a language spoken by 60 million people that remains severely underrepresented in NLP. The corpus is assembled from 39 sources spanning seven HuggingFace datasets and 32…

Computation and Language · Computer Science 2026-03-18 Hanif Rahman

This research paper presents a part-of-speech (POS) annotated dataset and tagger tool for the low-resource Uzbek language. The dataset includes 12 tags, which were used to develop a rule-based POS-tagger tool. The corpus text used in the…

Computation and Language · Computer Science 2023-03-02 Maksud Sharipov , Elmurod Kuriyozov , Ollabergan Yuldashev , Ogabek Sobirov

In this paper we introduce PerPaDa, a Persian paraphrase dataset that is collected from users' input in a plagiarism detection system. As an implicit crowdsourcing experience, we have gathered a large collection of original and paraphrased…

Computation and Language · Computer Science 2022-09-14 Salar Mohtaj , Fatemeh Tavakkoli , Habibollah Asghari

This research introduces the first large-scale, well-balanced Persian social media text classification dataset, specifically designed to address the lack of comprehensive resources in this domain. The dataset comprises 36,000 posts across…

Computation and Language · Computer Science 2026-05-26 Isun Chehreh , Ebrahim Ansari

As large language models (LLMs) become increasingly embedded in our daily lives, evaluating their quality and reliability across diverse contexts has become essential. While comprehensive benchmarks exist for assessing LLM performance in…

Sign Language Recognition (SLR) is a fast-growing field that aims to fill the communication gaps between the hearing-impaired and people without hearing loss. Existing solutions for Persian Sign Language (PSL) are limited to word-level…

Human-Computer Interaction · Computer Science 2024-06-25 Amirparsa Salmankhah , Amirreza Rajabi , Negin Kheirmand , Ali Fadaeimanesh , Amirreza Tarabkhah , Amirreza Kazemzadeh , Hamed Farbeh

Retrieval augmented generation (RAG) models, which integrate large-scale pre-trained generative models with external retrieval mechanisms, have shown significant success in various natural language processing (NLP) tasks. However, applying…

Computation and Language · Computer Science 2024-11-07 Hossein Hosseini , Mohammad Sobhan Zare , Amir Hossein Mohammadi , Arefeh Kazemi , Zahra Zojaji , Mohammad Ali Nematbakhsh

We have developed a full discourse parser in the Penn Discourse Treebank (PDTB) style. Our trained parser first identifies all discourse and non-discourse relations, locates and labels their arguments, and then classifies their relation…

Computation and Language · Computer Science 2016-08-30 Ziheng Lin , Hwee Tou Ng , Min-Yen Kan

The Iranian Persian language has two varieties: standard and colloquial. Most natural language processing tools for Persian assume that the text is in standard form: this assumption is wrong in many real applications especially web content.…

Computation and Language · Computer Science 2020-12-11 Mohammad Sadegh Rasooli , Farzane Bakhtyari , Fatemeh Shafiei , Mahsa Ravanbakhsh , Chris Callison-Burch

In this work, we employ a semi-automatic method based on back translation to generate a sentential paraphrase corpus for the Armenian language. The initial collection of sentences is translated from Armenian to English and back twice,…

Computation and Language · Computer Science 2020-09-29 Arthur Malajyan , Karen Avetisyan , Tsolak Ghukasyan

We present an expanded version of our previously released Kazakh text-to-speech (KazakhTTS) synthesis corpus. In the new KazakhTTS2 corpus, the overall size has increased from 93 hours to 271 hours, the number of speakers has risen from two…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-21 Saida Mussakhojayeva , Yerbolat Khassanov , Huseyin Atakan Varol