English
Related papers

Related papers: AfriVoices-KE: A Multilingual Speech Dataset for K…

200 papers

Audio captioning is the novel task of general audio content description using free text. It is an intermodal translation task (not speech-to-text), where a system accepts as an input an audio signal and outputs the textual description (i.e.…

Sound · Computer Science 2019-10-22 Konstantinos Drossos , Samuel Lipping , Tuomas Virtanen

The global aging population faces considerable challenges, particularly in communication, due to the prevalence of hearing and speech impairments. To address these, we introduce the AVE speech, a comprehensive multi-modal dataset for speech…

Sound · Computer Science 2025-07-08 Dongliang Zhou , Yakun Zhang , Jinghan Wu , Xingyu Zhang , Liang Xie , Erwei Yin

Large-scale multilingual evaluations, such as MEGA, often include only a handful of African languages due to the scarcity of high-quality evaluation data and the limited discoverability of existing African datasets. This lack of…

Computation and Language · Computer Science 2025-06-10 Jessica Ojo , Odunayo Ogundepo , Akintunde Oladipo , Kelechi Ogueji , Jimmy Lin , Pontus Stenetorp , David Ifeoluwa Adelani

We present SpeakingFaces as a publicly-available large-scale multimodal dataset developed to support machine learning research in contexts that utilize a combination of thermal, visual, and audio data streams; examples include…

Human-Computer Interaction · Computer Science 2021-05-04 Madina Abdrakhmanova , Askat Kuzdeuov , Sheikh Jarju , Yerbolat Khassanov , Michael Lewis , Huseyin Atakan Varol

This paper presents a novel framework for speech transcription and synthesis, leveraging edge-cloud parallelism to enhance processing speed and accessibility for Kinyarwanda and Swahili speakers. It addresses the scarcity of powerful…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-10-21 Pacome Simon Mbonimpa , Diane Tuyizere , Azizuddin Ahmed Biyabani , Ozan K. Tonguz

Audio-language models (ALMs) generate linguistic descriptions of sound-producing events and scenes. Advances in dataset creation and computational power have led to significant progress in this domain. This paper surveys 69 datasets used to…

Sound · Computer Science 2025-02-10 Gijs Wijngaard , Elia Formisano , Michele Esposito , Michel Dumontier

This study focuses on recognizing Bangladeshi dialects and converting diverse Bengali accents into standardized formal Bengali speech. Dialects, often referred to as regional languages, are distinctive variations of a language spoken in a…

Computation and Language · Computer Science 2024-11-19 Md. Nazmus Sadat Samin , Jawad Ibn Ahad , Tanjila Ahmed Medha , Fuad Rahman , Mohammad Ruhul Amin , Nabeel Mohammed , Shafin Rahman

Automatic speech recognition (ASR) for African languages remains constrained by limited labeled data and the lack of systematic guidance on model selection, data scaling, and decoding strategies. Large pre-trained systems such as Whisper,…

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Tingwei Guo , Cheng Wen , Dongwei Jiang , Ne Luo , Ruixiong Zhang , Shuaijiang Zhao , Wubo Li , Cheng Gong , Wei Zou , Kun Han , Xiangang Li

This paper presents an augmentation of MSCOCO dataset where speech is added to image and text. Speech captions are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images.…

Computation and Language · Computer Science 2020-11-24 William Havard , Laurent Besacier , Olivier Rosec

We present Kunkado, a 160-hour Bambara ASR dataset compiled from Malian radio archives to capture present-day spontaneous speech across a wide range of topics. It includes code-switching, disfluencies, background noise, and overlapping…

Computation and Language · Computer Science 2025-12-23 Yacouba Diarra , Panga Azazia Kamate , Nouhoum Souleymane Coulibaly , Michael Leventhal

Advancements in audio deepfake technology offers benefits like AI assistants, better accessibility for speech impairments, and enhanced entertainment. However, it also poses significant risks to security, privacy, and trust in digital…

Sound · Computer Science 2025-06-27 Abhay Kumar , Kunal Verma , Omkar More

Code-switching, the alternation between two or more languages within communication, poses great challenges for Automatic Speech Recognition (ASR) systems. Existing models and datasets are limited in their ability to effectively handle these…

Sound · Computer Science 2025-11-14 Yupei Li , Zifan Wei , Heng Yu , Jiahao Xue , Huichi Zhou , Björn W. Schuller

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other…

The development of robust, multilingual speaker recognition systems is hindered by a lack of large-scale, publicly available and multilingual datasets, particularly for the read-speech style crucial for applications like anti-spoofing. To…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-26 Aref Farhadipour , Jan Marquenie , Srikanth Madikeri , Eleanor Chodroff

We introduce the largest transcribed Arabic speech corpus, QASR, collected from the broadcast domain. This multi-dialect speech dataset contains 2,000 hours of speech sampled at 16kHz crawled from Aljazeera news channel. The dataset is…

Computation and Language · Computer Science 2021-06-25 Hamdy Mubarak , Amir Hussein , Shammur Absar Chowdhury , Ahmed Ali

Deepfakes represent a growing concern across domains such as disinformation, fraud, and non-consensual media. In particular, the rise of video conference and identity-driven attacks in high-stakes scenarios--such as impostor hiring--demands…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Sarah Barrington , Maty Bohacek , Hany Farid

We present Ara-BEST-RQ, a family of self-supervised learning (SSL) models specifically designed for multi-dialectal Arabic speech processing. Leveraging 5,640 hours of crawled Creative Commons speech and combining it with publicly available…

Computation and Language · Computer Science 2026-03-24 Haroun Elleuch , Ryan Whetten , Salima Mdhaffar , Yannick Estève , Fethi Bougares

This paper presents the development of a prototype Automatic Speech Recognition (ASR) system specifically designed for Bengali biomedical data. Recent advancements in Bengali ASR are encouraging, but a lack of domain-specific data limits…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-21 Shariar Kabir , Nazmun Nahar , Shyamasree Saha , Mamunur Rashid

Recent efforts in Spoken Dialogue Modeling aim to synthesize spoken dialogue without the need for direct transcription, thereby preserving the wealth of non-textual information inherent in speech. However, this approach faces a challenge…

Computation and Language · Computer Science 2024-07-03 Yu-Kuan Fu , Cheng-Kuang Lee , Hsiu-Hsuan Wang , Hung-yi Lee
‹ Prev 1 3 4 5 6 7 10 Next ›