English
Related papers

Related papers: MusicHiFi: Fast High-Fidelity Stereo Vocoding

200 papers

Neural vocoders are central to speech synthesis; despite their success, most still suffer from limited prosody modeling and inaccurate phase reconstruction. We propose a vocoder that introduces prosody-guided harmonic attention to enhance…

Sound · Computer Science 2026-01-22 Mohammed Salah Al-Radhi , Riad Larbi , Mátyás Bartalis , Géza Németh

In recent years, with the realistic generation results and a wide range of personalized applications, diffusion-based generative models gain huge attention in both visual and audio generation areas. Compared to the considerable advancements…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Shiqi Yang , Zhi Zhong , Mengjie Zhao , Shusuke Takahashi , Masato Ishii , Takashi Shibuya , Yuki Mitsufuji

Neural autoencoders underpin generative models. Practical, large-scale use of neural autoencoders for generative modeling necessitates fast encoding, low latent rates, and a single model across representations. Existing approaches are…

Sound · Computer Science 2026-02-23 Jonah Casebeer , Ge Zhu , Zhepei Wang , Nicholas J. Bryan

The advent of Large Models marks a new era in machine learning, significantly outperforming smaller models by leveraging vast datasets to capture and synthesize complex patterns. Despite these advancements, the exploration into scaling,…

Sound · Computer Science 2024-02-05 Shijia Liao , Shiyi Lan , Arun George Zachariah

Audio super-resolution is challenging owing to its ill-posed nature. Recently, the application of diffusion models in audio super-resolution has shown promising results in alleviating this challenge. However, diffusion-based models have…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-12 Jun-Hak Yun , Seung-Bin Kim , Seong-Whan Lee

Stereophonic audio is an indispensable ingredient to enhance human auditory experience. Recent research has explored the usage of visual information as guidance to generate binaural or ambisonic audio from mono ones with stereo supervision.…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Hang Zhou , Xudong Xu , Dahua Lin , Xiaogang Wang , Ziwei Liu

Generative Adversarial Network (GAN) based vocoders are superior in inference speed and synthesis quality when reconstructing an audible waveform from an acoustic representation. This study focuses on improving the discriminator to promote…

Sound · Computer Science 2023-11-28 Yicheng Gu , Xueyao Zhang , Liumeng Xue , Zhizheng Wu

It is challenging to accelerate the training process while ensuring both high-quality generated voices and acceptable inference speed. In this paper, we propose a novel neural vocoder called InstructSing, which can converge much faster…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-11 Chang Zeng , Chunhui Wang , Xiaoxiao Miao , Jian Zhao , Zhonglin Jiang , Yong Chen

Binaural stereo audio is recorded by imitating the way the human ear receives sound, which provides people with an immersive listening experience. Existing approaches leverage autoencoders and directly exploit visual spatial information to…

Sound · Computer Science 2023-11-15 Zhaojian Li , Bin Zhao , Yuan Yuan

Most neural vocoders employ band-limited mel-spectrograms to generate waveforms. If full-band spectral features are used as the input, the vocoder can be provided with as much acoustic information as possible. However, in some models…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Won Jang , Dan Lim , Jaesam Yoon , Bongwan Kim , Juntae Kim

Current two-stage TTS framework typically integrates an acoustic model with a vocoder -- the acoustic model predicts a low resolution intermediate representation such as Mel-spectrum while the vocoder generates waveform from the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-23 Jian Cong , Shan Yang , Lei Xie , Dan Su

We present a deep convolutional GAN which leverages techniques from MP3/Vorbis audio compression to produce long, high-quality audio samples with long-range coherence. The model uses a Modified Discrete Cosine Transform (MDCT) data…

Sound · Computer Science 2021-01-14 Korneel van den Broek

The task of Mel vocoding, i.e., the inversion of a Mel magnitude spectrogram to an audio waveform, is still a key component in many text-to-speech (TTS) systems today. Based on generative flow matching, our prior work on generative STFT…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-19 Simon Welker , Tal Peer , Timo Gerkmann

Recently, mainstream mel-spectrogram-based neural vocoders rely on generative adversarial network (GAN) for high-fidelity speech generation, e.g., HiFi-GAN and BigVGAN. However, the use of GAN restricts training efficiency and model…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-12 Hui-Peng Du , Yang Ai , Rui-Chen Zheng , Ye-Xin Lu , Zhen-Hua Ling

Neural network-based methods have recently demonstrated state-of-the-art results on image synthesis and super-resolution tasks, in particular by using variants of generative adversarial networks (GANs) with supervised feature losses.…

Sound · Computer Science 2019-03-22 Sung Kim , Visvesh Sathe

We introduce a state-of-the-art real-time, high-fidelity, audio codec leveraging neural networks. It consists in a streaming encoder-decoder architecture with quantized latent space trained in an end-to-end fashion. We simplify and speed-up…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-25 Alexandre Défossez , Jade Copet , Gabriel Synnaeve , Yossi Adi

Audio super-resolution is a challenging task of recovering the missing high-resolution features from a low-resolution signal. To address this, generative adversarial networks (GAN) have been used to achieve promising results by training the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-12 Shichao Hu , Bin Zhang , Beici Liang , Ethan Zhao , Simon Lui

Recently, GAN vocoders have seen rapid progress in speech synthesis, starting to outperform autoregressive models in perceptual quality with much higher generation speed. However, autoregressive vocoders are still the common choice for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-10 Ahmed Mustafa , Jan Büthe , Srikanth Korse , Kishan Gupta , Guillaume Fuchs , Nicola Pia

Audio coding is an essential module in the real-time communication system. Neural audio codecs can compress audio samples with a low bitrate due to the strong modeling and generative capabilities of deep neural networks. To address the poor…

Sound · Computer Science 2023-10-18 Wenzhe Liu , Wei Xiao , Meng Wang , Shan Yang , Yupeng Shi , Yuyong Kang , Dan Su , Shidong Shang , Dong Yu

Audio restoration has become increasingly significant in modern society, not only due to the demand for high-quality auditory experiences enabled by advanced playback devices, but also because the growing capabilities of generative audio…

Sound · Computer Science 2025-01-08 Kai Li , Yi Luo