English
Related papers

Related papers: HmBlogs: A big general Persian corpus

200 papers

We introduce the Self-Annotated Reddit Corpus (SARC), a large corpus for sarcasm research and for training and evaluating systems for sarcasm detection. The corpus has 1.3 million sarcastic statements -- 10 times more than any previous…

Computation and Language · Computer Science 2018-03-26 Mikhail Khodak , Nikunj Saunshi , Kiran Vodrahalli

We release a corpus of 43 million atomic edits across 8 languages. These edits are mined from Wikipedia edit history and consist of instances in which a human editor has inserted a single contiguous phrase into, or deleted a single…

Computation and Language · Computer Science 2018-08-29 Manaal Faruqui , Ellie Pavlick , Ian Tenney , Dipanjan Das

We present the QuranMorph corpus, a morphologically annotated corpus for the Quran (77,429 tokens). Each token in the QuranMorph was manually lemmatized and tagged with its part-of-speech by three expert linguists. The lemmatization process…

Computation and Language · Computer Science 2025-06-24 Diyam Akra , Tymaa Hammouda , Mustafa Jarrar

We present ASCAT (Arabic Scientific Corpus for Advanced Translation), a high-quality English-Arabic parallel benchmark corpus designed for scientific translation evaluation constructed through a systematic multi-engine translation and human…

Computation and Language · Computer Science 2026-04-02 Serry Sibaee , Khloud Al Jallad , Zineb Yousfi , Israa Elsayed Elhosiny , Yousra El-Ghawi , Batool Balah , Omer Nacar

Text mining is becoming vital as Web 2.0 offers collaborative content creation and sharing. Now Researchers have growing interest in text mining methods for discovering knowledge. Text mining researchers come from variety of areas like:…

Artificial Intelligence · Computer Science 2012-01-11 Mehwish Aziz , Muhammad Rafi

Enhancing reasoning capabilities in small language models is critical for specialized applications such as medical question answering, particularly in underrepresented languages like Persian. In this study, we employ Reinforcement Learning…

Computation and Language · Computer Science 2026-01-07 Mehrdad Ghassabi , Sadra Hakim , Hamidreza Baradaran Kashani , Pedram Rostami

Large Language Models (LLMs) demonstrate remarkable fluency across high-resource languages yet consistently fail to generate coherent text in Kashmiri, a language spoken by approximately seven million people. This performance disparity…

Computation and Language · Computer Science 2026-01-06 Haq Nawaz Malik

We describe here an experimental framework and the results obtained on microblogs retrieval. We study the contribution one popular approach, i.e., words embeddings, and investigate the impact of the training set on the learned embedding. We…

Information Retrieval · Computer Science 2019-11-19 Philippe Mulhem , Lorraine Goeuriot , Massih-Reza Amini , Nayanika Dogra

We present a hybrid methodology for generating large-scale semantic relationship datasets in low-resource languages, demonstrated through a comprehensive Turkish semantic relations corpus. Our approach integrates three phases: (1) FastText…

Computation and Language · Computer Science 2026-01-21 Ebubekir Tosun , Mehmet Emin Buldur , Özay Ezerceli , Mahmoud ElHussieni

The availability of corpora is a major factor in building natural language processing applications. However, the costs of acquiring corpora can prevent some researchers from going further in their endeavours. The ease of access to freely…

Computation and Language · Computer Science 2017-02-28 Wajdi Zaghouani

Currently there are lots of plagiarism detection approaches. But few of them implemented and adapted for Persian languages. In this paper, our work on designing and implementation of a plagiarism detection system based on pre-processing and…

Artificial Intelligence · Computer Science 2014-03-10 Maryam Mahmoodi , Mohammad Mahmoodi Varnamkhasti

The lack of large-scale datasets has been a major hindrance to the development of NLP tasks such as spelling correction and grammatical error correction (GEC). As a complementary new resource for these tasks, we present the GitHub Typo…

Computation and Language · Computer Science 2019-12-02 Masato Hagiwara , Masato Mita

One fundamental task for NLP is to determine the similarity between two texts and evaluate the extent of their likeness. The previous methods for the Persian language have low accuracy and are unable to comprehend the structure and meaning…

Computation and Language · Computer Science 2024-07-30 Seyed Mojtaba Sadjadi , Zeinab Rajabi , Leila Rabiei , Mohammad-Shahram Moin

The development of speech technologies for languages with limited digital representation poses significant challenges, primarily due to the scarcity of available data. This issue is exacerbated in the era of large, data-intensive models.…

Computation and Language · Computer Science 2024-06-24 Georgios Paraskevopoulos , Chara Tsoukala , Athanasios Katsamanis , Vassilis Katsouros

We present KS-PRET-5M, the largest publicly available pretraining dataset for the Kashmiri language, comprising 5,090,244 (5.09M) words, 27,692,959 (27.6M) characters, and a vocabulary of 295,433 (295.4K) unique word types. We assembled the…

Computation and Language · Computer Science 2026-04-14 Haq Nawaz Malik , Nahfid Nissar

The research in the field of music is rapidly growing, and this trend emphasizes the need for comprehensive data. Though researchers have made an effort to contribute their own datasets, many data collections lack the requisite inclusivity…

Sound · Computer Science 2023-11-21 Parsa Rasouli , Azam Bastanfard

We present a new open-source parallel corpus consisting of news articles collected from the Bianet magazine, an online newspaper that publishes Turkish news, often along with their translations in English and Kurdish. In this paper, we…

Computation and Language · Computer Science 2018-05-15 Duygu Ataman

This paper presents the first comprehensive comparative analysis of modern machine learning architectures for transliteration between Tajik (Cyrillic script) and Persian (Arabic script). A key contribution is the creation and validation of…

Computation and Language · Computer Science 2026-05-05 Mullosharaf K. Arabov

The democratization of AI is currently hindered by the immense computational costs required to train Large Language Models (LLMs) for low-resource languages. This paper presents Persian-Phi, a 3.8B parameter model that challenges the…

Computation and Language · Computer Science 2025-12-09 Amir Mohammad Akhlaghi , Amirhossein Shabani , Mostafa Abdolmaleki , Saeed Reza Kheradpisheh

Persian Poetry has consistently expressed its philosophy, wisdom, speech, and rationale on the basis of its couplets, making it an enigmatic language on its own to both native and non-native speakers. Nevertheless, the notice able gap…

Computation and Language · Computer Science 2022-09-22 Reza Khanmohammadi , Mitra Sadat Mirshafiee , Yazdan Rezaee Jouryabi , Seyed Abolghasem Mirroshandel