English
Related papers

Related papers: An Arabic-Hebrew parallel corpus of TED talks

200 papers

We present ASCAT (Arabic Scientific Corpus for Advanced Translation), a high-quality English-Arabic parallel benchmark corpus designed for scientific translation evaluation constructed through a systematic multi-engine translation and human…

Computation and Language · Computer Science 2026-04-02 Serry Sibaee , Khloud Al Jallad , Zineb Yousfi , Israa Elsayed Elhosiny , Yousra El-Ghawi , Batool Balah , Omer Nacar

In this paper, we introduce TEDxTN, the first publicly available Tunisian Arabic to English speech translation dataset. This work is in line with the ongoing effort to mitigate the data scarcity obstacle for a number of Arabic dialects. We…

Computation and Language · Computer Science 2025-11-17 Fethi Bougares , Salima Mdhaffar , Haroun Elleuch , Yannick Estève

We present KUTED, a speech-to-text translation (S2TT) dataset for Central Kurdish, derived from TED and TEDx talks. The corpus comprises 91,000 sentence pairs, including 170 hours of English audio, 1.65 million English tokens, and 1.40…

Computation and Language · Computer Science 2026-04-02 Mohammad Mohammadamini , Daban Q. Jaff , Josep Crego , Marie Tahon , Antoine Laurent

We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm ideas for a certain…

Computation and Language · Computer Science 2024-03-28 Injy Hamed , Fadhl Eryani , David Palfreyman , Nizar Habash

This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and specialized capabilities. We propose a taxonomy organizing…

We present an approach based on multilingual sentence embeddings to automatically extract parallel sentences from the content of Wikipedia articles in 85 languages, including several dialects or low-resource languages. We do not limit the…

Computation and Language · Computer Science 2019-07-17 Holger Schwenk , Vishrav Chaudhary , Shuo Sun , Hongyu Gong , Francisco Guzmán

We train a bilingual Arabic-Hebrew language model using a transliterated version of Arabic texts in Hebrew, to ensure both languages are represented in the same script. Given the morphological, structural similarities, and the extensive…

Computation and Language · Computer Science 2024-02-27 Aviad Rom , Kfir Bar

Machine translation between Arabic and Hebrew has so far been limited by a lack of parallel corpora, despite the political and cultural importance of this language pair. Previous work relied on manually-crafted grammars or pivoting via…

Computation and Language · Computer Science 2016-09-27 Yonatan Belinkov , James Glass

Text summarization has been intensively studied in many languages, and some languages have reached advanced stages. Yet, Arabic Text Summarization (ATS) is still in its developing stages. Existing ATS datasets are either small or lack…

Computation and Language · Computer Science 2022-10-26 Abdulaziz Alhamadani , Xuchao Zhang , Jianfeng He , Chang-Tien Lu

This paper introduces PEACH, a sentence-aligned parallel English-Arabic corpus of healthcare texts encompassing patient information leaflets and educational materials. The corpus contains 51,671 parallel sentences, totaling approximately…

Computation and Language · Computer Science 2025-08-11 Rania Al-Sabbagh

Arabic spans over 30 spoken varieties, yet no open-source text-to-speech system unifies them. Key barriers include substantial cross-dialect lexical and phonological divergence, scarce synthesis-grade data, and the absence of a standardized…

Computation and Language · Computer Science 2026-04-01 Yushen Chen , Junzhe Liu , Yujie Tu , Zhikang Niu , Yuzhe Liang , Chunyu Qiang , Chen Zhang , Kai Yu , Xie Chen

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source…

Computation and Language · Computer Science 2021-06-16 Elizabeth Salesky , Matthew Wiesner , Jacob Bremerman , Roldano Cattoni , Matteo Negri , Marco Turchi , Douglas W. Oard , Matt Post

With the development of large text-to-speech (TTS) models and scale-up of the training data, state-of-the-art TTS systems have achieved impressive performance. In this paper, we present WenetSpeech4TTS, a multi-domain Mandarin corpus…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-21 Linhan Ma , Dake Guo , Kun Song , Yuepeng Jiang , Shuai Wang , Liumeng Xue , Weiming Xu , Huan Zhao , Binbin Zhang , Lei Xie

We show that margin-based bitext mining in a multilingual sentence space can be applied to monolingual corpora of billions of sentences. We are using ten snapshots of a curated common crawl corpus (Wenzek et al., 2019) totalling 32.7…

Computation and Language · Computer Science 2020-05-04 Holger Schwenk , Guillaume Wenzek , Sergey Edunov , Edouard Grave , Armand Joulin

This study presents EgyBERT, an Arabic language model pretrained on 10.4 GB of Egyptian dialectal texts. We evaluated EgyBERT's performance by comparing it with five other multidialect Arabic language models across 10 evaluation datasets.…

Computation and Language · Computer Science 2024-08-08 Faisal Qarah

We present the Knesset Corpus, a corpus of Hebrew parliamentary proceedings containing over 30 million sentences (over 384 million tokens) from all the (plenary and committee) protocols held in the Israeli parliament between 1998 and 2022.…

Computation and Language · Computer Science 2025-06-02 Gili Goldin , Nick Howell , Noam Ordan , Ella Rabinovich , Shuly Wintner

Zero-shot multi-speaker text-to-speech (ZS-TTS) systems have advanced for English, however, it still lags behind due to insufficient resources. We address this gap for Arabic, a language of more than 450 million native speakers, by first…

Computation and Language · Computer Science 2024-07-09 Khai Duy Doan , Abdul Waheed , Muhammad Abdul-Mageed

Ramsa is a developing 41-hour speech corpus of Emirati Arabic designed to support sociolinguistic research and low-resource language technologies. It contains recordings from structured interviews with native speakers and episodes from…

Computation and Language · Computer Science 2026-03-10 Rania Al-Sabbagh

This paper proposes a novel approach to an automatic estimation of three speaker traits from Arabic speech: gender, emotion, and dialect. After showing promising results on different text classification tasks, the multi-task learning (MTL)…

Computation and Language · Computer Science 2020-12-15 Wael Farhan , Muhy Eddin Za'ter , Qusai Abu Obaidah , Hisham al Bataineh , Zyad Sober , Hussein T. Al-Natsheh

We propose a method for efficiently finding all parallel passages in a large corpus, even if the passages are not quite identical due to rephrasing and orthographic variation. The key ideas are the representation of each word in the corpus…

Computation and Language · Computer Science 2023-06-22 Avi Shmidman , Moshe Koppel , Ely Porat
‹ Prev 1 2 3 10 Next ›