English
Related papers

Related papers: PerSoMed: A Large-Scale Balanced Dataset for Persi…

200 papers

We present Persian Question Answering Dataset (PQuAD), a crowdsourced reading comprehension dataset on Persian Wikipedia articles. It includes 80,000 questions along with their answers, with 25% of the questions being adversarially…

Computation and Language · Computer Science 2023-02-22 Kasra Darvishi , Newsha Shahbodagh , Zahra Abbasiantaeb , Saeedeh Momtazi

This research introduces a state-of-the-art Persian spelling correction system that seamlessly integrates deep learning techniques with phonetic analysis, significantly enhancing the accuracy and efficiency of natural language processing…

Computation and Language · Computer Science 2024-07-23 Seyed Mohammad Sadegh Dashti , Amid Khatibi Bardsiri , Mehdi Jafari Shahbazzadeh

A user-generated text on social media enables health workers to keep track of information, identify possible outbreaks, forecast disease trends, monitor emergency cases, and ascertain disease awareness and response to official health…

Computation and Language · Computer Science 2022-04-12 Usman Naseem , Byoung Chan Lee , Matloob Khushi , Jinman Kim , Adam G. Dunn

With the surge of pretrained language models, a new pathway has been opened to incorporate Persian text contextual information. Meanwhile, as many other countries, including Iran, are fighting against COVID-19, a plethora of COVID-19…

Information Retrieval · Computer Science 2021-07-15 Reza Khanmohammadi , Mitra Sadat Mirshafiee , Mehdi Allahyari

The growing importance of culturally-aware natural language processing systems has led to an increasing demand for resources that capture sociopragmatic phenomena across diverse languages. Nevertheless, Arabic-language resources for…

Social media aids disaster response but suffers from noise, hindering accurate impact assessment and decision making for resilient cities, which few studies considered. To address the problem, this study proposes the first domain-specific…

Computation and Language · Computer Science 2024-06-03 Jin Han , Zhe Zheng , Xin-Zheng Lu , Ke-Yin Chen , Jia-Rui Lin

Persian names present unique challenges for natural language processing applications, particularly in gender detection and digital identity creation, due to transliteration inconsistencies and cultural-specific naming patterns. Existing…

Machine Learning · Computer Science 2025-09-16 Farbod Bijary , Mohsen Ebadpour , Amirhosein Tajbakhsh

In this paper, we propose a methodology for task 10 of SemEval23, focusing on detecting and classifying online sexism in social media posts. The task is tackling a serious issue, as detecting harmful content on social media platforms is…

Computation and Language · Computer Science 2023-04-26 Sana Sabah Al-Azzawi , György Kovács , Filip Nilsson , Tosin Adewumi , Marcus Liwicki

While emerging Persian NLP benchmarks have expanded into pragmatics and politeness, they rarely distinguish between memorized cultural facts and the ability to reason about implicit social norms. We introduce DivanBench, a diagnostic…

Computation and Language · Computer Science 2026-02-20 Alireza Sakhaeirad , Ali Ma'manpoosh , Arshia Hemmat

Musical instrument classification is essential for music information retrieval (MIR) and generative music systems. However, research on non-Western traditions, particularly Persian music, remains limited. We address this gap by introducing…

Large Language Models (LLMs), trained on extensive datasets using advanced deep learning architectures, have demonstrated remarkable performance across a wide range of language tasks, becoming a cornerstone of modern AI technologies.…

Semantic role labeling is a crucial task in natural language processing, enabling better comprehension of natural language. However, the lack of annotated data in multiple languages has posed a challenge for researchers. To address this, a…

Computation and Language · Computer Science 2024-08-29 Mohammad Ebrahimi , Behrouz Minaei Bidgoli , Nasim Khozouei

In the current environment, psychological issues are prevalent and widespread, with social media serving as a key outlet for individuals to share their feelings. This results in the generation of vast quantities of data daily, where…

Computation and Language · Computer Science 2024-06-13 Wei Zhai , Hongzhi Qi , Qing Zhao , Jianqiang Li , Ziqi Wang , Han Wang , Bing Xiang Yang , Guanghui Fu

Turkish is one of the most popular languages in the world. Wide us of this language on social media platforms such as Twitter, Instagram, or Tiktok and strategic position of the country in the world politics makes it appealing for the…

Computation and Language · Computer Science 2023-12-01 Ali Najafi , Onur Varol

Recently, pre-trained transformer-based architectures have proven to be very efficient at language modeling and understanding, given that they are trained on a large enough corpus. Applications in language generation for Arabic are still…

Computation and Language · Computer Science 2021-03-09 Wissam Antoun , Fady Baly , Hazem Hajj

With the rapid expansion of academic literature and the proliferation of preprints, researchers face growing challenges in manually organizing and labeling large volumes of articles. The NSLP 2024 FoRC Shared Task I addresses this challenge…

Artificial Intelligence · Computer Science 2024-08-16 Benjamin Wolff , Eva Seidlmayer , Konrad U. Förstner

A public dataset, with a variety of properties suitable for sentiment analysis [1], event prediction, trend detection and other text mining applications, is needed in order to be able to successfully perform analysis studies. The vast…

Computation and Language · Computer Science 2018-02-01 Semiha Makinist , Ibrahim Riza Hallac , Betul Ay Karakus , Galip Aydin

The performance of learning models heavily relies on the availability and adequacy of training data. To address the dataset adequacy issue, researchers have extensively explored data augmentation (DA) as a promising approach. DA generates…

Computation and Language · Computer Science 2023-08-22 Dania Refai , Saleh Abo-Soud , Mohammad Abdel-Rahman

The Arabic language suffers from a great shortage of datasets suitable for training deep learning models, and the existing ones include general non-specialized classifications. In this work, we introduce a new Arab medical dataset, which…

Computation and Language · Computer Science 2021-07-06 Jaafar Hammoud , Aleksandra Vatian , Natalia Dobrenko , Nikolai Vedernikov , Anatoly Shalyto , Natalia Gusarova

This study presents EgyBERT, an Arabic language model pretrained on 10.4 GB of Egyptian dialectal texts. We evaluated EgyBERT's performance by comparing it with five other multidialect Arabic language models across 10 evaluation datasets.…

Computation and Language · Computer Science 2024-08-08 Faisal Qarah
‹ Prev 1 4 5 6 7 8 10 Next ›