English
Related papers

Related papers: MSR-NV: Neural Vocoder Using Multiple Sampling Rat…

200 papers

Deep learning has dramatically improved the performance of sounds recognition. However, learning acoustic models directly from the raw waveform is still challenging. Current waveform-based models generally use time-domain convolutional…

Sound · Computer Science 2018-03-29 Boqing Zhu , Changjian Wang , Feng Liu , Jin Lei , Zengquan Lu , Yuxing Peng

Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often yield suboptimal…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Chihyun Liu , Jiaxuan Fan , Mingtung Sun , Michael Anthony , Mingsian R. Bai , Yu Tsao

Singing voice synthesis (SVS) has advanced significantly, enabling models to generate vocals with accurate pitch and consistent style. As these capabilities improve, the need for reliable evaluation and optimization becomes increasingly…

Sound · Computer Science 2025-12-03 Xueyan Li , Yuxin Wang , Mengjie Jiang , Qingzi Zhu , Jiang Zhang , Zoey Kim , Yazhe Niu

Non-verbal Vocalizations (NVs), such as laughter and sighs, are vital for conveying emotion and intention in human speech, yet most existing speech systems neglect them, which severely compromises communicative richness and emotional…

Sound · Computer Science 2026-01-14 Runchuan Ye , Yixuan Zhou , Renjie Yu , Zijian Lin , Kehan Li , Xiang Li , Xin Liu , Guoyang Zeng , Zhiyong Wu

Most state-of-the-art Text-to-Speech systems use the mel-spectrogram as an intermediate representation, to decompose the task into acoustic modelling and waveform generation. A mel-spectrogram is extracted from the waveform by a simple,…

We present NeRF-SR, a solution for high-resolution (HR) novel view synthesis with mostly low-resolution (LR) inputs. Our method is built upon Neural Radiance Fields (NeRF) that predicts per-point density and color with a multi-layer…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Chen Wang , Xian Wu , Yuan-Chen Guo , Song-Hai Zhang , Yu-Wing Tai , Shi-Min Hu

Autoregressive neural vocoders have achieved outstanding performance in speech synthesis tasks such as text-to-speech and voice conversion. An autoregressive vocoder predicts a sample at some time step conditioned on those at previous time…

Sound · Computer Science 2024-06-06 Po-chun Hsu , Da-rong Liu , Andy T. Liu , Hung-yi Lee

Source separation can improve automatic speech recognition (ASR) under multi-party meeting scenarios by extracting single-speaker signals from overlapped speech. Despite the success of self-supervised learning models in single-channel…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-04 Yuang Li , Xianrui Zheng , Philip C. Woodland

Recent progress in deep generative models has improved the quality of neural vocoders in speech domain. However, generating a high-quality singing voice remains challenging due to a wider variety of musical expressions in pitch, loudness,…

Sound · Computer Science 2022-10-19 Naoya Takahashi , Mayank Kumar , Singh , Yuki Mitsufuji

It is challenging to accelerate the training process while ensuring both high-quality generated voices and acceptable inference speed. In this paper, we propose a novel neural vocoder called InstructSing, which can converge much faster…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-11 Chang Zeng , Chunhui Wang , Xiaoxiao Miao , Jian Zhao , Zhonglin Jiang , Yong Chen

Developing microphone array technologies for a small number of microphones is important due to the constraints of many devices. One direction to address this situation consists of virtually augmenting the number of microphone signals, e.g.,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-13 Tsubasa Ochiai , Marc Delcroix , Tomohiro Nakatani , Rintaro Ikeshita , Keisuke Kinoshita , Shoko Araki

Audio-Visual Speech Recognition (AVSR) achieves robust speech recognition in noisy environments by combining auditory and visual information. However, recent Large Language Model (LLM) based AVSR systems incur high computational costs due…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Jeong Hun Yeo , Hyeongseop Rha , Se Jin Park , Yong Man Ro

Recently, diffusion models (DMs) have been increasingly used in audio processing tasks, including speech super-resolution (SR), which aims to restore high-frequency content given low-resolution speech utterances. This is commonly achieved…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-22 Chin-Yun Yu , Sung-Lin Yeh , György Fazekas , Hao Tang

Neural Radiance Fields (NeRF) methods have proved effective as compact, high-quality and versatile representations for 3D scenes, and enable downstream tasks such as editing, retrieval, navigation, etc. Various neural architectures are…

Computer Vision and Pattern Recognition · Computer Science 2023-01-04 Shuangkang Fang , Weixin Xu , Heng Wang , Yi Yang , Yufeng Wang , Shuchang Zhou

In recent years, various flow-based generative models have been proposed to generate high-fidelity waveforms in real-time. However, these models require either a well-trained teacher network or a number of flow steps making them…

Sound · Computer Science 2020-07-06 Hyeongju Kim , Hyeonseung Lee , Woo Hyun Kang , Sung Jun Cheon , Byoung Jin Choi , Nam Soo Kim

We are interested in a novel task, singing voice beautifying (SVB). Given the singing voice of an amateur singer, SVB aims to improve the intonation and vocal tone of the voice, while keeping the content and vocal timbre. Current automatic…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-03 Jinglin Liu , Chengxi Li , Yi Ren , Zhiying Zhu , Zhou Zhao

In recent years, deep learning-based approaches have significantly improved the performance of single-channel speech enhancement. However, due to the limitation of training data and computational complexity, real-time enhancement of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-16 Zehua Zhang , Lu Zhang , Xuyi Zhuang , Yukun Qian , Heng Li , Mingjiang Wang

While the spatial directivity of multichannel speech enhancement algorithms improves with the number of microphones, fitting large capture arrays into real-world edge devices is typically limited by physical constraints. To overcome this…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-08 Dongheon Lee , Ashutosh Pandey , Sanjeel Parekh , Daniel Wong , Jacob Donley , Buye Xu , Juan Azcarreta

Neural vocoders, used for converting the spectral representations of an audio signal to the waveforms, are a commonly used component in speech synthesis pipelines. It focuses on synthesizing waveforms from low-dimensional representation,…

Sound · Computer Science 2021-12-07 Ehab A. AlBadawy , Andrew Gibiansky , Qing He , Jilong Wu , Ming-Ching Chang , Siwei Lyu

Recently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance over the conventional MVDR by replacing the matrix inversion…

Sound · Computer Science 2021-04-27 Xiyun Li , Yong Xu , Meng Yu , Shi-Xiong Zhang , Jiaming Xu , Bo Xu , Dong Yu