English
Related papers

Related papers: Tadabur: A Large-Scale Quran Audio Dataset

200 papers

In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours in total. We collect…

Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10000 Arabic tweets and…

Computation and Language · Computer Science 2025-05-26 Wajdi Zaghouani , Md. Rafiul Biswas

In this study, we introduce ManaTTS, the most extensive publicly accessible single-speaker Persian corpus, and a comprehensive framework for collecting transcribed speech datasets for the Persian language. ManaTTS, released under the open…

Sound · Computer Science 2024-09-12 Mahta Fetrat Qharabagh , Zahra Dehghanian , Hamid R. Rabiee

In this study, we introduce YODAS (YouTube-Oriented Dataset for Audio and Speech), a large-scale, multilingual dataset comprising currently over 500k hours of speech data in more than 100 languages, sourced from both labeled and unlabeled…

Computation and Language · Computer Science 2024-06-04 Xinjian Li , Shinnosuke Takamichi , Takaaki Saeki , William Chen , Sayaka Shiota , Shinji Watanabe

Diacritization of Arabic text is both an interesting and a challenging problem at the same time with various applications ranging from speech synthesis to helping students learning the Arabic language. Like many other tasks or problems in…

Computation and Language · Computer Science 2019-05-07 Ali Fadel , Ibraheem Tuffaha , Bara' Al-Jawarneh , Mahmoud Al-Ayyoub

The lack of a publicly-available large-scale and diverse dataset has long been a significant bottleneck for singing voice applications like Singing Voice Synthesis (SVS) and Singing Voice Conversion (SVC). To tackle this problem, we present…

Sound · Computer Science 2025-05-15 Yicheng Gu , Chaoren Wang , Junan Zhang , Xueyao Zhang , Zihao Fang , Haorui He , Zhizheng Wu

This paper introduces a new multi-modal dataset for visual and audio-visual speech recognition. It includes face tracks from over 400 hours of TED and TEDx videos, along with the corresponding subtitles and word alignment boundaries. The…

Computer Vision and Pattern Recognition · Computer Science 2018-10-30 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Jiaben Chen , Zixin Wang , Ailing Zeng , Yang Fu , Xueyang Yu , Siyuan Cen , Julian Tanke , Yihang Chen , Koichi Saito , Yuki Mitsufuji , Chuang Gan

A comprehensive examination of data science vocabulary usage over the past 13 years in this work is conducted. The investigation commences with a dataset comprising 16,018 abstracts that feature the term "data science" in either the title,…

Applications · Statistics 2023-10-24 Igor Barahona

Most Arabic natural language processing tools and resources are developed to serve Modern Standard Arabic (MSA), which is the official written language in the Arab World. Some Dialectal Arabic varieties, notably Egyptian Arabic, have…

Computation and Language · Computer Science 2016-09-13 Salam Khalifa , Nizar Habash , Dana Abdulrahim , Sara Hassan

Text summarization has been intensively studied in many languages, and some languages have reached advanced stages. Yet, Arabic Text Summarization (ATS) is still in its developing stages. Existing ATS datasets are either small or lack…

Computation and Language · Computer Science 2022-10-26 Abdulaziz Alhamadani , Xuchao Zhang , Jianfeng He , Chang-Tien Lu

This paper presents BSTC (Baidu Speech Translation Corpus), a large-scale Chinese-English speech translation dataset. This dataset is constructed based on a collection of licensed videos of talks or lectures, including about 68 hours of…

Computation and Language · Computer Science 2021-04-28 Ruiqing Zhang , Xiyang Wang , Chuanqiang Zhang , Zhongjun He , Hua Wu , Zhi Li , Haifeng Wang , Ying Chen , Qinfei Li

Music source separation performance has greatly improved in recent years with the advent of approaches based on deep learning. Such methods typically require large amounts of labelled training data, which in the case of music consist of…

Sound · Computer Science 2019-09-19 Ethan Manilow , Gordon Wichern , Prem Seetharaman , Jonathan Le Roux

The goal of this paper is twofold. First, we introduce DALI, a large and rich multimodal dataset containing 5358 audio tracks with their time-aligned vocal melody notes and lyrics at four levels of granularity. The second goal is to explain…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-26 Gabriel Meseguer-Brocal , Alice Cohen-Hadria , Geoffroy Peeters

Question answering (QA) systems are now available through numerous commercial applications for a wide variety of domains, serving millions of users that interact with them via speech interfaces. However, current benchmarks in QA research do…

Computation and Language · Computer Science 2021-09-27 Fahim Faisal , Sharlina Keshava , Md Mahfuz ibn Alam , Antonios Anastasopoulos

Nowadays, one of the main challenges for Question Answering Systems is to answer complex questions using various sources of information. Multi-hop questions are a type of complex questions that require multi-step reasoning to answer. In…

Computation and Language · Computer Science 2023-04-25 Arash Ghafouri , Hasan Naderi , Mohammad Aghajani asl , Mahdi Firouzmandi

Speech recognition has received a less attention in Bengali literature due to the lack of a comprehensive dataset. In this paper, we describe the development process of the first comprehensive Bengali speech dataset on real numbers. It…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-28 Md Mahadi Hasan Nahid , Md. Ashraful Islam , Bishwajit Purkaystha , Md Saiful Islam

We introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models. The dataset consists of 1.1k human-written natural language descriptions…

Arabic is a widely-spoken language with a long and rich history, but existing corpora and language technology focus mostly on modern Arabic and its varieties. Therefore, studying the history of the language has so far been mostly limited to…

Computation and Language · Computer Science 2018-09-12 Yonatan Belinkov , Alexander Magidow , Alberto Barrón-Cedeño , Avi Shmidman , Maxim Romanov

The Arabic language is among the most popular languages in the world with a huge variety of dialects spoken in 22 countries. In this study, we address the problem of classifying 18 Arabic dialects of the QADI dataset of Arabic tweets. RNN…

Computation and Language · Computer Science 2025-07-01 Omar A. Essameldin , Ali O. Elbeih , Wael H. Gomaa , Wael F. Elsersy
‹ Prev 1 3 4 5 6 7 10 Next ›