中文
相关论文

相关论文: Zero-Shot Mono-to-Binaural Speech Synthesis

200 篇论文

In this work, we present an end-to-end binaural speech synthesis system that combines a low-bitrate audio codec with a powerful binaural decoder that is capable of accurate speech binauralization while faithfully reconstructing…

声音 · 计算机科学 2022-07-11 Wen Chin Huang , Dejan Markovic , Alexander Richard , Israel Dejene Gebru , Anjali Menon

This paper addresses the problem of localizing audio sources using binaural measurements. We propose a supervised formulation that simultaneously localizes multiple sources at different locations. The approach is intrinsically efficient…

声音 · 计算机科学 2016-04-18 Antoine Deleforge , Radu Horaud , Yoav Schechner , Laurent Girin

Binaural audio provides human listeners with an immersive spatial sound experience, but most existing videos lack binaural audio recordings. We propose an audio spatialization method that draws on visual information in videos to convert…

计算机视觉与模式识别 · 计算机科学 2021-11-23 Rishabh Garg , Ruohan Gao , Kristen Grauman

Audio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the need for extensive,…

The success of monocular depth estimation relies on large and diverse training sets. Due to the challenges associated with acquiring dense ground-truth depth across different environments at scale, a number of datasets with distinct…

计算机视觉与模式识别 · 计算机科学 2020-08-26 René Ranftl , Katrin Lasinger , David Hafner , Konrad Schindler , Vladlen Koltun

Spoken language understanding systems using audio-only data are gaining popularity, yet their ability to handle unseen intents remains limited. In this study, we propose a generalized zero-shot audio-to-intent classification framework with…

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

声音 · 计算机科学 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

Stereo matching serves as a cornerstone in 3D vision, aiming to establish pixel-wise correspondences between stereo image pairs for depth recovery. Despite remarkable progress driven by deep neural architectures, current models often…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Xianda Guo , Chenming Zhang , Youmin Zhang , Ruilin Wang , Dujun Nie , Wenzhao Zheng , Matteo Poggi , Hao Zhao , Mang Ye , Qin Zou , Long Chen

Conventional text-to-speech (TTS) research has predominantly focused on enhancing the quality of synthesized speech for speakers in the training dataset. The challenge of synthesizing lifelike speech for unseen, out-of-dataset speakers,…

声音 · 计算机科学 2024-04-30 Wenbin Wang , Yang Song , Sanjay Jha

Deep learning techniques for separating audio into different sound sources face several challenges. Standard architectures require training separate models for different types of audio sources. Although some universal separators employ a…

声音 · 计算机科学 2022-02-15 Ke Chen , Xingjian Du , Bilei Zhu , Zejun Ma , Taylor Berg-Kirkpatrick , Shlomo Dubnov

Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms…

音频与语音处理 · 电气工程与系统科学 2024-04-11 Ziyue Jiang , Jinglin Liu , Yi Ren , Jinzheng He , Zhenhui Ye , Shengpeng Ji , Qian Yang , Chen Zhang , Pengfei Wei , Chunfeng Wang , Xiang Yin , Zejun Ma , Zhou Zhao

Monaural speech enhancement has achieved remarkable progress recently. However, its performance has been constrained by the limited spatial cues available at a single microphone. To overcome this limitation, we introduce a strategy to map…

音频与语音处理 · 电气工程与系统科学 2024-03-05 Xinmeng Xu , Yuhong Yang , Weiping Tu

Voice Conversion research in recent times has increasingly focused on improving the zero-shot capabilities of existing methods. Despite remarkable advancements, current architectures still tend to struggle in zero-shot cross-lingual…

声音 · 计算机科学 2025-05-26 Advait Joglekar , Divyanshu Singh , Rooshil Rohit Bhatia , S. Umesh

In this paper, we study zero-shot learning in audio classification through factored linear and nonlinear acoustic-semantic projections between audio instances and sound classes. Zero-shot learning in audio classification refers to…

音频与语音处理 · 电气工程与系统科学 2021-02-03 Huang Xie , Okko Räsänen , Tuomas Virtanen

Audio-visual zero-shot learning aims to classify samples consisting of a pair of corresponding audio and video sequences from classes that are not present during training. An analysis of the audio-visual data reveals a large degree of…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Jie Hong , Zeeshan Hayder , Junlin Han , Pengfei Fang , Mehrtash Harandi , Lars Petersson

Recently, zero-shot TTS and VC methods have gained attention due to their practicality of being able to generate voices even unseen during training. Among these methods, zero-shot modifications of the VITS model have shown superior…

音频与语音处理 · 电气工程与系统科学 2023-05-29 Seongyeon Park , Bohyung Kim , Tae-hyun Oh

Data scarcity and the modality gap between the speech and text modalities are two major obstacles of end-to-end Speech Translation (ST) systems, thus hindering their performance. Prior work has attempted to mitigate these challenges by…

计算与语言 · 计算机科学 2024-06-07 Ioannis Tsiamas , Gerard I. Gállego , José A. R. Fonollosa , Marta R. Costa-jussà

We present a novel way of conditioning a pretrained denoising diffusion speech model to produce speech in the voice of a novel person unseen during training. The method requires a short (~3 seconds) sample from the target person, and…

声音 · 计算机科学 2022-06-23 Alon Levkovitch , Eliya Nachmani , Lior Wolf

Despite progress in video-to-audio generation, the field focuses predominantly on mono output, lacking spatial immersion. Existing binaural approaches remain constrained by a two-stage pipeline that first generates mono audio and then…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Mengchen Zhang , Qi Chen , Tong Wu , Zihan Liu , Dahua Lin

Binaural audio gives the listener the feeling of being in the recording place and enhances the immersive experience if coupled with AR/VR. But the problem with binaural audio recording is that it requires a specialized setup which is not…

声音 · 计算机科学 2021-08-12 Kranti Kumar Parida , Siddharth Srivastava , Neeraj Matiyali , Gaurav Sharma