English
Related papers

Related papers: SaudiBERT: A Large Language Model Pretrained on Sa…

200 papers

This study aims at investigating the effect of applying single learner machine learning approach and ensemble machine learning approach for offensive language detection on Arabic language. Classifying Arabic social media text is a very…

Computation and Language · Computer Science 2020-05-20 Fatemah Husain

Modern Arabic ASR systems such as wav2vec 2.0 excel at word- and sentence-level transcription, yet struggle to classify isolated letters. In this study, we show that this phoneme-level task, crucial for language learning, speech therapy,…

Computation and Language · Computer Science 2025-08-28 Hadi Zaatiti , Hatem Hajri , Osama Abdullah , Nader Masmoudi

The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Arabic texts,…

The Arabic language is characterized by a rich tapestry of regional dialects that differ substantially in phonetics and lexicon, reflecting the geographic and cultural diversity of its speakers. Despite the availability of many…

This paper proposes a methodology to prepare corpora in Arabic language from online social network (OSN) and review site for Sentiment Analysis (SA) task. The paper also proposes a methodology for generating a stopword list from the…

Computation and Language · Computer Science 2014-10-07 Walaa Medhat , Ahmed H. Yousef , Hoda Korashy

Event-argument extraction is a challenging task, particularly in Arabic due to sparse linguistic resources. To fill this gap, we introduce the \hadath corpus ($550$k tokens) as an extension of Wojood, enriched with event-argument…

Computation and Language · Computer Science 2024-08-01 Alaa Aljabari , Lina Duaibes , Mustafa Jarrar , Mohammed Khalilia

Given the impact of language models on the field of Natural Language Processing, a number of Spanish encoder-only masked language models (aka BERTs) have been trained and released. These models were developed either within large projects…

Computation and Language · Computer Science 2023-09-25 Rodrigo Agerri , Eneko Agirre

We describe an Arabic-Hebrew parallel corpus of TED talks built upon WIT3, the Web inventory that repurposes the original content of the TED website in a way which is more convenient for MT researchers. The benchmark consists of about 2,000…

Computation and Language · Computer Science 2016-10-04 Mauro Cettolo

Social media platforms have a vital role in the modern world, serving as conduits for communication, the exchange of ideas, and the establishment of networks. However, the misuse of these platforms through toxic comments, which can range…

Computation and Language · Computer Science 2025-06-24 Mukaffi Bin Moin , Pronay Debnath , Usafa Akther Rifa , Rijeet Bin Anis

With the growth of social media platform influence, the effect of their misuse becomes more and more impactful. The importance of automatic detection of threatening and abusive language can not be overestimated. However, most of the…

Computation and Language · Computer Science 2022-07-15 Maaz Amjad , Alisa Zhila , Grigori Sidorov , Andrey Labunets , Sabur Butta , Hamza Imam Amjad , Oxana Vitman , Alexander Gelbukh

Transformer-based models have advanced NLP, yet Hebrew still lacks a large-scale RoBERTa encoder which is extensively trained. Existing models such as HeBERT, AlephBERT, and HeRo are limited by corpus size, vocabulary, or training depth. We…

Computation and Language · Computer Science 2025-10-27 Raphael Scheible-Schmitt

Since BERT appeared, Transformer language models and transfer learning have become state-of-the-art for Natural Language Understanding tasks. Recently, some works geared towards pre-training specially-crafted models for particular domains,…

Computation and Language · Computer Science 2022-05-05 Juan Manuel Pérez , Damián A. Furman , Laura Alonso Alemany , Franco Luque

This study presents a comprehensive comparative evaluation of four state-of-the-art Large Language Models (LLMs)--Claude 3.7 Sonnet, DeepSeek-V3, Gemini 2.0 Flash, and GPT-4o--for sentiment analysis and emotion detection in Persian social…

Computation and Language · Computer Science 2025-09-19 Kian Tohidi , Kia Dashtipour , Simone Rebora , Sevda Pourfaramarz

Training on multiple modalities of input can augment the capabilities of a language model. Here, we ask whether such a training regime can improve the quality and efficiency of these systems as well. We focus on text--audio and introduce…

Computation and Language · Computer Science 2023-12-08 Lukas Wolf , Greta Tuckute , Klemen Kotar , Eghbal Hosseini , Tamar Regev , Ethan Wilcox , Alex Warstadt

We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm ideas for a certain…

Computation and Language · Computer Science 2024-03-28 Injy Hamed , Fadhl Eryani , David Palfreyman , Nizar Habash

Pretrained language models based on the Transformer architecture have achieved state-of-the-art results in various natural language processing tasks such as part-of-speech tagging, named entity recognition, and question answering. However,…

Computation and Language · Computer Science 2021-08-24 B. Mansurov , A. Mansurov

Social media is heading towards more and more personalization, where individuals reveal their beliefs, interests, habits, and activities, simply offering glimpses into their personality traits. This study, explores the correlation between…

Computation and Language · Computer Science 2024-07-24 Mokhaiber Dandash , Masoud Asadpour

Social media currently provide a window on our lives, making it possible to learn how people from different places, with different backgrounds, ages, and genders use language. In this work we exploit a newly-created Arabic dataset with…

Computation and Language · Computer Science 2019-11-05 Muhammad Abdul-Mageed , Chiyu Zhang , Arun Rajendran , AbdelRahim Elmadany , Michael Przystupa , Lyle Ungar

This paper describes the training process of the first Czech monolingual language representation models based on BERT and ALBERT architectures. We pre-train our models on more than 340K of sentences, which is 50 times more than multilingual…

Computation and Language · Computer Science 2021-08-23 Jakub Sido , Ondřej Pražák , Pavel Přibáň , Jan Pašek , Michal Seják , Miloslav Konopík

We train a bilingual Arabic-Hebrew language model using a transliterated version of Arabic texts in Hebrew, to ensure both languages are represented in the same script. Given the morphological, structural similarities, and the extensive…

Computation and Language · Computer Science 2024-02-27 Aviad Rom , Kfir Bar