English
Related papers

Related papers: Tadabur: A Large-Scale Quran Audio Dataset

200 papers

Speech translation for Indian languages remains a challenging task due to the scarcity of large-scale, publicly available datasets that capture the linguistic diversity and domain coverage essential for real-world applications. Existing…

Large Language Models (LLMs) have demonstrated significant promise for various applications in healthcare. However, their efficacy in the Arabic medical domain remains unexplored due to the lack of high-quality domain-specific datasets and…

Computation and Language · Computer Science 2025-08-25 Mouath Abu Daoud , Chaimae Abouzahir , Leen Kharouf , Walid Al-Eisawi , Nizar Habash , Farah E. Shamout

Query-based document summarization aims to extract or generate a summary of a document which directly answers or is relevant to the search query. It is an important technique that can be beneficial to a variety of applications such as…

Artificial Intelligence · Computer Science 2020-10-29 Mingjun Zhao , Shengli Yan , Bang Liu , Xinwang Zhong , Qian Hao , Haolan Chen , Di Niu , Bowei Long , Weidong Guo

This paper describes the Arabic MGB-3 Challenge - Arabic Speech Recognition in the Wild. Unlike last year's Arabic MGB-2 Challenge, for which the recognition task was based on more than 1,200 hours broadcast TV news recordings from…

Computation and Language · Computer Science 2017-09-22 Ahmed Ali , Stephan Vogel , Steve Renals

We present KUTED, a speech-to-text translation (S2TT) dataset for Central Kurdish, derived from TED and TEDx talks. The corpus comprises 91,000 sentence pairs, including 170 hours of English audio, 1.65 million English tokens, and 1.40…

Computation and Language · Computer Science 2026-04-02 Mohammad Mohammadamini , Daban Q. Jaff , Josep Crego , Marie Tahon , Antoine Laurent

Developing Automatic Speech Recognition (ASR) systems for Tunisian Arabic Dialect is challenging due to the dialect's linguistic complexity and the scarcity of annotated speech datasets. To address these challenges, we propose the LinTO…

Computation and Language · Computer Science 2025-04-04 Hedi Naouara , Jean-Pierre Lorré , Jérôme Louradour

The development of Urdu scene text detection, recognition, and Visual Question Answering (VQA) technologies is crucial for advancing accessibility, information retrieval, and linguistic diversity in digital content, facilitating better…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Hiba Maryam , Ling Fu , Jiajun Song , Tajrian ABM Shafayet , Qidi Luo , Xiang Bai , Yuliang Liu

This article presents morphologically-annotated Yemeni, Sudanese, Iraqi, and Libyan Arabic dialects Lisan corpora. Lisan features around 1.2 million tokens. We collected the content of the corpora from several social media platforms. The…

Computation and Language · Computer Science 2022-12-20 Mustafa Jarrar , Fadi A Zaraket , Tymaa Hammouda , Daanish Masood Alavi , Martin Waahlisch

In this paper, we release a largest ever medical Question Answering (QA) dataset with 26 million QA pairs. We benchmark many existing approaches in our dataset in terms of both retrieval and generation. Experimental results show that the…

Computation and Language · Computer Science 2023-05-03 Jianquan Li , Xidong Wang , Xiangbo Wu , Zhiyi Zhang , Xiaolong Xu , Jie Fu , Prayag Tiwari , Xiang Wan , Benyou Wang

In this paper, we introduce the Temporal Audio Source Counting Network (TaCNet), an innovative architecture that addresses limitations in audio source counting tasks. TaCNet operates directly on raw audio inputs, eliminating complex…

In this paper, we introduce the Extreme Metal Vocals Dataset, which comprises a collection of recordings of extreme vocal techniques performed within the realm of heavy metal music. The dataset consists of 760 audio excerpts of 1 second to…

Sound · Computer Science 2024-06-26 Modan Tailleur , Julien Pinquier , Laurent Millot , Corsin Vogel , Mathieu Lagrange

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real online one-to-one…

Computation and Language · Computer Science 2022-06-28 Chengfei Li , Shuhao Deng , Yaoping Wang , Guangjing Wang , Yaguang Gong , Changbin Chen , Jinfeng Bai

Accurate and contextually faithful responses are critical when applying large language models (LLMs) to sensitive and domain-specific tasks, such as answering queries related to quranic studies. General-purpose LLMs often struggle with…

Computation and Language · Computer Science 2025-03-24 Zahra Khalila , Arbi Haza Nasution , Winda Monika , Aytug Onan , Yohei Murakami , Yasir Bin Ismail Radi , Noor Mohammad Osmani

We introduce QuerYD, a new large-scale dataset for retrieval and event localisation in video. A unique feature of our dataset is the availability of two audio tracks for each video: the original audio, and a high-quality spoken description…

Computer Vision and Pattern Recognition · Computer Science 2021-02-18 Andreea-Maria Oncescu , João F. Henriques , Yang Liu , Andrew Zisserman , Samuel Albanie

Medical audio data is difficult to collect due to privacy regulations and high annotation costs arising from domain expertise. Thus, existing benchmarks tend to underrepresent complex medical audio scenarios. To address this challenge, we…

As speech generation technology advances, the risk of misuse through deepfake audio has become a pressing concern, which underscores the critical need for robust detection systems. However, many existing speech deepfake datasets are limited…

Sound · Computer Science 2025-07-30 Wen Huang , Yanmei Gu , Zhiming Wang , Huijia Zhu , Yanmin Qian

Human conversations are complicated and building a human-like dialogue agent is an extremely challenging task. With the rapid development of deep learning techniques, data-driven models become more and more prevalent which need a huge…

Computation and Language · Computer Science 2020-03-25 Meng Chen , Ruixue Liu , Lei Shen , Shaozu Yuan , Jingyan Zhou , Youzheng Wu , Xiaodong He , Bowen Zhou

Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, induce dependencies with various prediction models during dataset…

We propose a method for efficiently finding all parallel passages in a large corpus, even if the passages are not quite identical due to rephrasing and orthographic variation. The key ideas are the representation of each word in the corpus…

Computation and Language · Computer Science 2023-06-22 Avi Shmidman , Moshe Koppel , Ely Porat

The lack of impaired speech data hinders advancements in the development of inclusive speech technologies, particularly in low-resource languages such as Akan. To address this gap, this study presents a curated corpus of speech samples from…