中文
相关论文

相关论文: Libri-adhoc40: A dataset collected from synchroniz…

200 篇论文

Speech separation has been shown effective for multi-talker speech recognition. Under the ad hoc microphone array setup where the array consists of spatially distributed asynchronous microphones, additional challenges must be overcome as…

声音 · 计算机科学 2021-03-04 Dongmei Wang , Takuya Yoshioka , Zhuo Chen , Xiaofei Wang , Tianyan Zhou , Zhong Meng

This paper introduces a new dataset, Libri-Adapt, to support unsupervised domain adaptation research on speech recognition models. Built on top of the LibriSpeech corpus, Libri-Adapt contains English speech recorded on mobile and…

音频与语音处理 · 电气工程与系统科学 2020-09-09 Akhil Mathur , Fahim Kawsar , Nadia Berthouze , Nicholas D. Lane

While deep-learning-based speaker localization has shown advantages in challenging acoustic environments, it often yields only direction-of-arrival (DOA) cues rather than precise two-dimensional (2D) coordinates. To address this, we propose…

音频与语音处理 · 电气工程与系统科学 2024-04-02 Shupei Liu , Linfeng Feng , Yijun Gong , Chengdong Liang , Chen Zhang , Xiao-Lei Zhang , Xuelong Li

Far-field speech processing is an important and challenging problem. In this paper, we propose \textit{deep ad-hoc beamforming}, a deep-learning-based multichannel speech enhancement framework based on ad-hoc microphone arrays, to address…

声音 · 计算机科学 2021-02-10 Xiao-Lei Zhang

Recently, ad-hoc microphone array has been widely studied. Unlike traditional microphone array settings, the spatial arrangement and number of microphones of ad-hoc microphone arrays are not known in advance, which hinders the adaptation of…

声音 · 计算机科学 2021-07-02 Chengdong Liang , Junqi Chen , Shanzheng Guan , Xiao-Lei Zhang

We introduce LibriConvo, a simulated multi-speaker conversational dataset based on speaker-aware conversation simulation (SASC), designed to support training and evaluation of speaker diarization and automatic speech recognition (ASR)…

音频与语音处理 · 电气工程与系统科学 2025-10-28 Máté Gedeon , Péter Mihajlik

Speech enhancement promises higher efficiency in ad-hoc microphone arrays than in constrained microphone arrays thanks to the wide spatial coverage of the devices in the acoustic scene. However, speech enhancement in ad-hoc microphone…

信号处理 · 电气工程与系统科学 2021-06-16 Nicolas Furnon , Romain Serizel , Slim Essid , Irina Illina

We propose a system that transcribes the conversation of a typical meeting scenario that is captured by a set of initially unsynchronized microphone arrays at unknown positions. It consists of subsystems for signal synchronization,…

音频与语音处理 · 电气工程与系统科学 2022-05-03 Tobias Gburrek , Christoph Boeddeker , Thilo von Neumann , Tobias Cord-Landwehr , Joerg Schmalenstroeer , Reinhold Haeb-Umbach

In this paper, we introduce Libriheavy, a large-scale ASR corpus consisting of 50,000 hours of read English speech derived from LibriVox. To the best of our knowledge, Libriheavy is the largest freely-available corpus of speech with…

音频与语音处理 · 电气工程与系统科学 2024-01-17 Wei Kang , Xiaoyu Yang , Zengwei Yao , Fangjun Kuang , Yifan Yang , Liyong Guo , Long Lin , Daniel Povey

Recent works in spoken language translation (SLT) have attempted to build end-to-end speech-to-text translation without using source language transcription during learning or decoding. However, while large quantities of parallel texts (such…

计算与语言 · 计算机科学 2018-02-12 Ali Can Kocabiyikoglu , Laurent Besacier , Olivier Kraif

This paper introduces a new speech corpus called "LibriTTS" designed for text-to-speech use. It is derived from the original audio and text materials of the LibriSpeech corpus, which has been used for training and evaluating automatic…

声音 · 计算机科学 2019-04-08 Heiga Zen , Viet Dang , Rob Clark , Yu Zhang , Ron J. Weiss , Ye Jia , Zhifeng Chen , Yonghui Wu

We present LibriWASN, a data set whose design follows closely the LibriCSS meeting recognition data set, with the marked difference that the data is recorded with devices that are randomly positioned on a meeting table and whose sampling…

声音 · 计算机科学 2023-08-22 Joerg Schmalenstroeer , Tobias Gburrek , Reinhold Haeb-Umbach

This paper introduces the Voices Obscured In Complex Environmental Settings (VOICES) corpus, a freely available dataset under Creative Commons BY 4.0. This dataset will promote speech and signal processing research of speech recorded by…

There is an emerging need for comparable data for multi-microphone processing, particularly in acoustic sensor networks. However, commonly available databases are often limited in the spatial diversity of the microphones or only allow for…

音频与语音处理 · 电气工程与系统科学 2026-02-11 Daniel Fejgin , Wiebke Middelberg , Simon Doclo

Speech enhancement in ad-hoc microphone arrays is often hindered by the asynchronization of the devices composing the microphone array. Asynchronization comes from sampling time offset and sampling rate offset which inevitably occur when…

音频与语音处理 · 电气工程与系统科学 2023-08-01 Nicolas Furnon , Romain Serizel , Slim Essid , Irina Illina

This paper presents a sensory fusion neuromorphic dataset collected with precise temporal synchronization using a set of Address-Event-Representation sensors and tools. The target application is the lip reading of several keywords for…

Speaker verification based on ad-hoc microphone arrays has the potential of reducing the error significantly in adverse acoustic environments. However, existing approaches extract utterance-level speaker embeddings from each channel of an…

声音 · 计算机科学 2022-03-29 Chengdong Liang , Yijiang Chen , Jiadi Yao , Xiao-Lei Zhang

Speech separation with several speakers is a challenging task because of the non-stationarity of the speech and the strong signal similarity between interferent sources. Current state-of-the-art solutions can separate well the different…

信号处理 · 电气工程与系统科学 2021-02-09 Nicolas Furnon , Romain Serizel , Irina Illina , Slim Essid

The evolving speech processing landscape is increasingly focused on complex scenarios like meetings or cocktail parties with multiple simultaneous speakers and far-field conditions. Existing methodologies for addressing these challenges…

Ad-hoc distributed microphone environments, where microphone locations and numbers are unpredictable, present a challenge to traditional deep learning models, which typically require fixed architectures. To tailor deep learning models to…

音频与语音处理 · 电气工程与系统科学 2024-06-17 Jihyun Kim , Stijn Kindt , Nilesh Madhu , Hong-Goo Kang
‹ 上一页 1 2 3 10 下一页 ›