English
Related papers

Related papers: HmBlogs: A big general Persian corpus

200 papers

Text corpora are essential for training models used in tasks like summarization, translation, and large language models (LLMs). While various efforts have been made to collect monolingual and multilingual datasets in many languages, Persian…

Computation and Language · Computer Science 2025-02-14 Sara Bourbour Hosseinbeigi , Fatemeh Taherinezhad , Heshaam Faili , Hamed Baghbani , Fatemeh Nadi , Mostafa Amiri

Introduction: Microblogging websites have massed rich data sources for sentiment analysis and opinion mining. In this regard, sentiment classification has frequently proven inefficient because microblog posts typically lack syntactically…

Computation and Language · Computer Science 2024-03-08 Mojtaba Mazoochi , Leila Rabiei , Farzaneh Rahmani , Zeinab Rajabi

We introduce FaBERT, a Persian BERT-base model pre-trained on the HmBlogs corpus, encompassing both informal and formal Persian texts. FaBERT is designed to excel in traditional Natural Language Understanding (NLU) tasks, addressing the…

Computation and Language · Computer Science 2024-02-12 Mostafa Masumi , Seyed Soroush Majd , Mehrnoush Shamsfard , Hamid Beigy

Social networks analysis and exploring is important for researchers, sociologists, academics, and various businesses due to their information potential. Because of the large volume, diversity, and the data growth rate in web 2.0, some…

Social and Information Networks · Computer Science 2014-09-02 Zeinab Borhani-fard , Leila Esmaeili , Behrouz Minaei-Bidgoli , Mehdi Nasiri

Despite the widespread use of the Persian language by millions globally, limited efforts have been made in natural language processing for this language. The use of large language models as effective tools in various natural language…

Computation and Language · Computer Science 2023-12-27 Mohammad Amin Abbasi , Arash Ghafouri , Mahdi Firouzmandi , Hassan Naderi , Behrouz Minaei Bidgoli

Language recognition has been significantly advanced in recent years by means of modern machine learning methods such as deep learning and benchmarks with rich annotations. However, research is still limited in low-resource formal…

Computation and Language · Computer Science 2020-06-03 Hadi Abdi Khojasteh , Ebrahim Ansari , Mahdi Bohlouli

In this article, we have introduced the first parallel corpus of Persian with more than 10 other European languages. This article describes primary steps toward preparing a Basic Language Resources Kit (BLARK) for Persian. Up to now, we…

Computation and Language · Computer Science 2014-04-18 Behrang Qasemizadeh , Saeed Rahimi , Behrooz Mahmoodi Bakhtiari

Large language models (LLMs) have demonstrated notable creative abilities in generating literary texts, including poetry and short stories. However, prior research has primarily centered on English, with limited exploration of non-English…

Computation and Language · Computer Science 2026-02-16 Armin Tourajmehr , Mohammad Reza Modarres , Yadollah Yaghoobzadeh

Keyphrases provide an extremely dense summary of a text. Such information can be used in many Natural Language Processing tasks, such as information retrieval and text summarization. Since previous studies on Persian keyword or keyphrase…

Computation and Language · Computer Science 2020-09-28 Ehsan Doostmohammadi , Mohammad Hadi Bokaei , Hossein Sameti

Informal language is a style of spoken or written language frequently used in casual conversations, social media, weblogs, emails and text messages. In informal writing, the language faces some lexical and/or syntactic changes varying among…

Computation and Language · Computer Science 2023-08-11 Vahide Tajalli , Fateme Kalantari , Mehrnoush Shamsfard

Introduction: Part-of-Speech (POS) Tagging, the process of classifying words into their respective parts of speech (e.g., verb or noun), is essential in various natural language processing applications. POS tagging is a crucial…

Computation and Language · Computer Science 2023-10-03 Leyla Rabiei , Farzaneh Rahmani , Mohammad Khansari , Zeinab Rajabi , Moein Salimi

One of the most major and essential tasks in natural language processing is machine translation that is now highly dependent upon multilingual parallel corpora. Through this paper, we introduce the biggest Persian-English parallel corpus…

Computation and Language · Computer Science 2020-02-03 Omid Kashefi

Tagged corpora play a crucial role in a wide range of Natural Language Processing. The Part of Speech Tagging (POST) is essential in developing tagged corpora. It is time-and-effort-consuming and costly, and therefore, it could be more…

Computation and Language · Computer Science 2022-02-01 Hossein Hassani

Over the past years, interest in discourse analysis and discourse parsing has steadily grown, and many discourse-annotated corpora and, as a result, discourse parsers have been built. In this paper, we present a discourse-annotated corpus…

Computation and Language · Computer Science 2021-06-29 Sara Shahmohammadi , Hadi Veisi , Ali Darzi

Sentiment Analysis (SA) is a major field of study in natural language processing, computational linguistics and information retrieval. Interest in SA has been constantly growing in both academia and industry over the recent years. Moreover,…

Computation and Language · Computer Science 2021-01-05 Pedram Hosseini , Ali Ahmadian Ramaki , Hassan Maleki , Mansoureh Anvari , Seyed Abolghasem Mirroshandel

This research introduces the first large-scale, well-balanced Persian social media text classification dataset, specifically designed to address the lack of comprehensive resources in this domain. The dataset comprises 36,000 posts across…

Computation and Language · Computer Science 2026-05-26 Isun Chehreh , Ebrahim Ansari

An automated approach to text readability assessment is essential to a language and can be a powerful tool for improving the understandability of texts written and published in that language. However, the Persian language, which is spoken…

Computation and Language · Computer Science 2020-04-23 Hamid Mohammadi , Seyed Hossein Khasteh

DeepMine is a speech database in Persian and English designed to build and evaluate text-dependent, text-prompted, and text-independent speaker verification, as well as Persian speech recognition systems. It contains more than 1850 speakers…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-10 Hossein Zeinali , Lukáš Burget , Jan "Honza'' Černocký

Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech-language modelling, and low-resource speech processing. We introduce ParsVoice, the largest…

Sound · Computer Science 2026-05-27 Mohammad Javad Ranjbar Kalahroodi , Heshaam Faili , Azadeh Shakery

Over recent years a lot of research papers and studies have been published on the development of effective approaches that benefit from a large amount of user-generated content and build intelligent predictive models on top of them. This…

Computation and Language · Computer Science 2021-01-21 Mohammad Kasra Habib
‹ Prev 1 2 3 10 Next ›