English
Related papers

Related papers: Unified Source-Filter GAN with Harmonic-plus-Noise…

200 papers

Pre-trained models for automatic speech recognition (ASR) and speech enhancement (SE) have exhibited remarkable capabilities under matched noise and channel conditions. However, these models often suffer from severe performance degradation…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-03 Chien-Chun Wang , Hung-Shin Lee , Hsin-Min Wang , Berlin Chen

The traditional vocoders have the advantages of high synthesis efficiency, strong interpretability, and speech editability, while the neural vocoders have the advantage of high synthesis quality. To combine the advantages of two vocoders,…

Sound · Computer Science 2022-03-08 Tao Wang , Ruibo Fu , Jiangyan Yi , Jianhua Tao , Zhengqi Wen

Neural vocoders are now being used in a wide range of speech processing applications. In many of those applications, the vocoder can be the most complex component, so finding lower complexity algorithms can lead to significant practical…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-06 Jean-Marc Valin , Ahmed Mustafa , Jan Büthe

Among the major remaining challenges for generative adversarial networks (GANs) is the capacity to synthesize globally and locally coherent images with object shapes and textures indistinguishable from real images. To target this issue we…

Computer Vision and Pattern Recognition · Computer Science 2021-03-23 Edgar Schönfeld , Bernt Schiele , Anna Khoreva

The application of generative adversarial networks (GANs) has recently advanced speech super-resolution (SR) based on intermediate representations like mel-spectrograms. However, existing SR methods that typically rely on independently…

Sound · Computer Science 2025-01-20 Shengkui Zhao , Kun Zhou , Zexu Pan , Yukun Ma , Chong Zhang , Bin Ma

With the emergence of GAN-based vocoders, the discriminator, as a crucial component, has been developed recently. In our work, we focus on improving the time-frequency based discriminator. Particularly, Short-Time Fourier Transform (STFT)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-04 Nan Xu , Zhaolong Huang , Xiao Zeng

The prevailing method for neural speech enhancement predominantly utilizes fully-supervised deep learning with simulated pairs of far-field noisy-reverberant speech and clean speech. Nonetheless, these models frequently demonstrate…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-16 Tong Lei , Qinwen Hu , Ziyao Lin , Andong Li , Rilin Chen , Meng Yu , Dong Yu , Jing Lu

In this paper, we propose a novel controllable text-to-image generative adversarial network (ControlGAN), which can effectively synthesise high-quality images and also control parts of the image generation according to natural language…

Computer Vision and Pattern Recognition · Computer Science 2019-12-20 Bowen Li , Xiaojuan Qi , Thomas Lukasiewicz , Philip H. S. Torr

This paper describes a method for decomposing steady-state instrument data into excitation and formant filter components. The input data, taken from several series of recordings of acoustical instruments is analyzed in the frequency domain,…

Sound · Computer Science 2007-05-23 Ilia Bisnovatyi , Michael J. O'Donnell

Incorporating prior knowledge like lexical constraints into the model's output to generate meaningful and coherent sentences has many applications in dialogue system, machine translation, image captioning, etc. However, existing RNN-based…

Computation and Language · Computer Science 2019-11-20 Dayiheng Liu , Jie Fu , Qian Qu , Jiancheng Lv

We propose a generative adversarial network for point cloud upsampling, which can not only make the upsampled points evenly distributed on the underlying surface but also efficiently generate clean high frequency regions. The generator of…

Computer Vision and Pattern Recognition · Computer Science 2022-12-14 Hao Liu , Hui Yuan , Junhui Hou , Raouf Hamzaoui , Wei Gao

Single-image generative adversarial networks learn from the internal distribution of a single training example to generate variations of it, removing the need of a large dataset. In this paper we introduce SpecSinGAN, an unconditional…

Sound · Computer Science 2022-04-06 Adrián Barahona-Ríos , Tom Collins

Real-world audio recordings are often degraded by factors such as noise, reverberation, and equalization distortion. This paper introduces HiFi-GAN, a deep learning method to transform recorded speech to sound as though it had been recorded…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-23 Jiaqi Su , Zeyu Jin , Adam Finkelstein

In practical application of speech codecs, a multitude of factors such as the quality of the radio connection, limiting hardware or required user experience necessitate trade-offs between achievable perceptual quality, engendered bitrate…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-25 Kishan Gupta , Srikanth Korse , Andreas Brendel , Nicola Pia , Guillaume Fuchs

Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generation highly challenging. Existing audio-video generation models often fail to maintain…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Shihao Cheng , Jiaxu Zhang , Quanyue Song , Shansong Liu , Zhizhi Guo , Xiaolei Zhang , Chi Zhang , Xuelong Li , Zhigang Tu

We propose a novel approach to generative adversarial networks (GANs) in which the standard i.i.d. Gaussian latent prior is replaced or hybridized with a quantum-correlated prior derived from measurements of a 16-qubit entangling circuit.…

Quantum Physics · Physics 2025-07-03 Hongni Jin , Kenneth M. Merz

We propose a multi-stage framework for universal speech enhancement, designed for the Interspeech 2025 URGENT Challenge. Our system first employs a Sparse Compression Network to robustly separate sources and extract an initial clean speech…

Sound · Computer Science 2025-06-03 Nabarun Goswami , Tatsuya Harada

Neural vocoders are central to speech synthesis; despite their success, most still suffer from limited prosody modeling and inaccurate phase reconstruction. We propose a vocoder that introduces prosody-guided harmonic attention to enhance…

Sound · Computer Science 2026-01-22 Mohammed Salah Al-Radhi , Riad Larbi , Mátyás Bartalis , Géza Németh

Voice conversion is to generate a new speech with the source content and a target voice style. In this paper, we focus on one general setting, i.e., non-parallel many-to-many voice conversion, which is close to the real-world scenario. As…

Sound · Computer Science 2022-07-28 Jian Ma , Zhedong Zheng , Hao Fei , Feng Zheng , Tat-seng Chua , Yi Yang

This paper introduces the Multi-Band Excited WaveNet a neural vocoder for speaking and singing voices. It aims to advance the state of the art towards an universal neural vocoder, which is a model that can generate voice signals from…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-08 Axel Roebel , Frederik Bous
‹ Prev 1 4 5 6 7 8 10 Next ›