English
Related papers

Related papers: An Arabic-Hebrew parallel corpus of TED talks

200 papers

We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration,…

Computation and Language · Computer Science 2025-05-28 Hawau Olamide Toyin , Rufael Marew , Humaid Alblooshi , Samar M. Magdy , Hanan Aldarmaki

The CMU Wilderness Multilingual Speech Dataset (Black, 2019) is a newly published multilingual speech dataset based on recorded readings of the New Testament. It provides data to build Automatic Speech Recognition (ASR) and Text-to-Speech…

Computation and Language · Computer Science 2020-02-27 Marcely Zanon Boito , William N. Havard , Mahault Garnerin , Éric Le Ferrand , Laurent Besacier

We introduce ALHD, the first large-scale comprehensive Arabic dataset explicitly designed to distinguish between human- and LLM-generated texts. ALHD spans three genres (news, social media, reviews), covering both MSA and dialectal Arabic,…

Computation and Language · Computer Science 2025-10-23 Ali Khairallah , Arkaitz Zubiaga

The rapid advancements in Large Language Models (LLMs) have led to significant improvements in various natural language processing tasks. However, the evaluation of LLMs' legal knowledge, particularly in non-English languages such as…

This paper describes the ON-TRAC Consortium translation systems developed for two challenge tracks featured in the Evaluation Campaign of IWSLT 2022: low-resource and dialect speech translation. For the Tunisian Arabic-English dataset…

Large language models are increasingly consulted for Islamic knowledge, yet no comprehensive benchmark evaluates their performance across core Islamic disciplines. We introduce IslamicMMLU, a benchmark of 10,013 multiple-choice questions…

Computation and Language · Computer Science 2026-04-06 Ali Abdelaal , Mohammed Nader Al Haffar , Mahmoud Fawzi , Walid Magdy

This paper presents a novel Dialectal Sound and Vowelization Recovery framework, designed to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages, that extends beyond its standard orthographic…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-06 Yassine El Kheir , Hamdy Mubarak , Ahmed Ali , Shammur Absar Chowdhury

Pre-trained Language Models (PLMs) are integral to many modern natural language processing (NLP) systems. Although multilingual models cover a wide range of languages, they often grapple with challenges like high inference costs and a lack…

Computation and Language · Computer Science 2024-07-19 Murtadha Ahmed , Saghir Alfasly , Bo Wen , Jamaal Qasem , Mohammed Ahmed , Yunfeng Liu

With the rise of generative text-to-speech models, distinguishing between real and synthetic speech has become challenging, especially for Arabic that have received limited research attention. Most spoof detection efforts have focused on…

Computation and Language · Computer Science 2025-09-30 Mohamed Maged , Alhassan Ehab , Ali Mekky , Besher Hassan , Shady Shehata

Matching texts in highly inflected languages such as Arabic by simple stemming strategy is unlikely to perform well. In this paper, we present a strategy for automatic text matching technique for for inflectional languages, using Arabic as…

Computation and Language · Computer Science 2014-03-25 Tarek El-Shishtawy , Fatma El-Ghannam

Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10000 Arabic tweets and…

Computation and Language · Computer Science 2025-05-26 Wajdi Zaghouani , Md. Rafiul Biswas

Transcribed speech and user-generated text in Arabic typically contain a mixture of Modern Standard Arabic (MSA), the standardized language taught in schools, and Dialectal Arabic (DA), used in daily communications. To handle this…

Computation and Language · Computer Science 2023-10-24 Amr Keleg , Sharon Goldwater , Walid Magdy

Large Language Models (LLMs) have demonstrated significant promise for various applications in healthcare. However, their efficacy in the Arabic medical domain remains unexplored due to the lack of high-quality domain-specific datasets and…

Computation and Language · Computer Science 2025-08-25 Mouath Abu Daoud , Chaimae Abouzahir , Leen Kharouf , Walid Al-Eisawi , Nizar Habash , Farah E. Shamout

The detection of toxic language in the Arabic language has emerged as an active area of research in recent years, and reviewing the existing datasets employed for training the developed solutions has become a pressing need. This paper…

Computation and Language · Computer Science 2024-01-31 Imene Bensalem , Paolo Rosso , Hanane Zitouni

This paper introduces a new multi-modal dataset for visual and audio-visual speech recognition. It includes face tracks from over 400 hours of TED and TEDx videos, along with the corresponding subtitles and word alignment boundaries. The…

Computer Vision and Pattern Recognition · Computer Science 2018-10-30 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Using pre-trained transformer models such as BERT has proven to be effective in many NLP tasks. This paper presents our work to fine-tune BERT models for Arabic Word Sense Disambiguation (WSD). We treated the WSD task as a sentence-pair…

Computation and Language · Computer Science 2022-05-20 Moustafa Al-Hajj , Mustafa Jarrar

We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluation for Modern…

Despite the growing importance of Arabic as a global language, there is a notable lack of language models pre-trained exclusively on Arabic data. This shortage has led to limited benchmarks available for assessing language model performance…

Computation and Language · Computer Science 2024-07-02 Shahad Al-Khalifa , Hend Al-Khalifa

While large language models (LLMs) excel in various natural language tasks in English, their performance in lower-resourced languages like Hebrew, especially for generative tasks such as abstractive summarization, remains unclear. The high…

Computation and Language · Computer Science 2025-07-14 Tzuf Paz-Argaman , Itai Mondshine , Asaf Achi Mordechai , Reut Tsarfaty

Recent research in natural language processing (NLP) has achieved impressive performance in tasks such as machine translation (MT), news classification, and question-answering in high-resource languages. However, the performance of MT…

Computation and Language · Computer Science 2024-03-29 Atnafu Lambebo Tonja , Olga Kolesnikova , Alexander Gelbukh , Jugal Kalita