English
Related papers

Related papers: Neural source-filter waveform models for statistic…

200 papers

This paper presents a reverberation module for source-filter-based neural vocoders that improves the performance of reverberant effect modeling. This module uses the output waveform of neural vocoders as an input and produces a reverberant…

Sound · Computer Science 2020-05-18 Yang Ai , Xin Wang , Junichi Yamagishi , Zhen-Hua Ling

We study the segmental recurrent neural network for end-to-end acoustic modelling. This model connects the segmental conditional random field (CRF) with a recurrent neural network (RNN) used for feature extraction. Compared to most previous…

Computation and Language · Computer Science 2016-06-21 Liang Lu , Lingpeng Kong , Chris Dyer , Noah A. Smith , Steve Renals

Nonnegative Matrix Factorization (NMF) is a powerful tool for decomposing mixtures of audio signals in the Time-Frequency (TF) domain. In the source separation framework, the phase recovery for each extracted component is necessary for…

Sound · Computer Science 2016-11-17 Paul Magron , Roland Badeau , Bertrand David

Conventionally, audio super-resolution models fixed the initial and the target sampling rates, which necessitate the model to be trained for each pair of sampling rates. We introduce NU-Wave 2, a diffusion model for neural audio upsampling…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-28 Seungu Han , Junhyeok Lee

Despite the recent popularity of deep generative state space models, few comparisons have been made between network architectures and the inference steps of the Bayesian filtering framework -- with most models simultaneously approximating…

Machine Learning · Statistics 2020-09-29 Bryan Lim , Stefan Zohren , Stephen Roberts

In our previous work, we proposed a neural vocoder called APNet, which directly predicts speech amplitude and phase spectra with a 5 ms frame shift in parallel from the input acoustic features, and then reconstructs the 16 kHz speech…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-21 Hui-Peng Du , Ye-Xin Lu , Yang Ai , Zhen-Hua Ling

This paper presents a waveform modeling and generation method using hierarchical recurrent neural networks (HRNN) for speech bandwidth extension (BWE). Different from conventional BWE methods which predict spectral parameters for…

Sound · Computer Science 2018-01-26 Zhen-Hua Ling , Yang Ai , Yu Gu , Li-Rong Dai

Recently, we proposed short-time Fourier transform (STFT)-based loss functions for training a neural speech waveform model. In this paper, we generalize the above framework and propose a training scheme for such models based on spectral…

Audio and Speech Processing · Electrical Eng. & Systems 2019-04-09 Shinji Takaki , Hirokazu Kameoka , Junichi Yamagishi

In this report we describe an ongoing line of research for solving single-channel source separation problems. Many monaural signal decomposition techniques proposed in the literature operate on a feature space consisting of a time-frequency…

Sound · Computer Science 2015-04-29 Pablo Sprechmann , Joan Bruna , Yann LeCun

Recent advances in generative visual models and neural radiance fields have greatly boosted 3D-aware image synthesis and stylization tasks. However, previous NeRF-based work is limited to single scene stylization, training a model to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-29 Zichen Tang , Hongyu Yang

In order to fully utilize spatial information for segmentation and address the challenge of handling areas with significant grayscale variations in remote sensing segmentation, we propose the SFFNet (Spatial and Frequency Domain Fusion…

Computer Vision and Pattern Recognition · Computer Science 2024-05-06 Yunsong Yang , Genji Yuan , Jinjiang Li

Acoustic beamformers have been widely used to enhance audio signals. Currently, the best methods are the deep neural network (DNN)-powered variants of the generalized eigenvalue and minimum-variance distortionless response beamformers and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-12 Yuichiro Koyama , Bhiksha Raj

We propose an independence-based joint dereverberation and separation method with a neural source model. We introduce a neural network in the framework of time-decorrelation iterative source steering, which is an extension of independent…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-04 Kohei Saijo , Robin Scheibler

In this paper, we propose a model to perform style transfer of speech to singing voice. Contrary to the previous signal processing-based methods, which require high-quality singing templates or phoneme synchronization, we explore a…

Sound · Computer Science 2022-08-29 Shrutina Agarwal , Sriram Ganapathy , Naoya Takahashi

Normalizing Flows (NFs) are likelihood-based models for continuous inputs. They have demonstrated promising results on both density estimation and generative modeling tasks, but have received relatively little attention in recent years. In…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Shuangfei Zhai , Ruixiang Zhang , Preetum Nakkiran , David Berthelot , Jiatao Gu , Huangjie Zheng , Tianrong Chen , Miguel Angel Bautista , Navdeep Jaitly , Josh Susskind

The talking head generation recently attracted considerable attention due to its widespread application prospects, especially for digital avatars and 3D animation design. Inspired by this practical demand, several works explored Neural…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Tianyong Wang , Xiangyu Liang , Wangguandong Zheng , Dan Niu , Haifeng Xia , Siyu Xia

This paper proposes a new unsupervised audio-visual speech enhancement (AVSE) approach that combines a diffusion-based audio-visual speech generative model with a non-negative matrix factorization (NMF) noise model. First, the diffusion…

Sound · Computer Science 2025-01-16 Jean-Eudes Ayilo , Mostafa Sadeghi , Romain Serizel , Xavier Alameda-Pineda

Most deep learning-based multi-channel speech enhancement methods focus on designing a set of beamforming coefficients to directly filter the low signal-to-noise ratio signals received by microphones, which hinders the performance of these…

Sound · Computer Science 2022-02-08 Wenzhe Liu , Andong Li , Chengshi Zheng , Xiaodong Li

This paper proposes voicing-aware conditional discriminators for Parallel WaveGAN-based waveform synthesis systems. In this framework, we adopt a projection-based conditioning method that can significantly improve the discriminator's…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-27 Ryuichi Yamamoto , Eunwoo Song , Min-Jae Hwang , Jae-Min Kim

While numerous forecasters have been proposed using different network architectures, the Transformer-based models have state-of-the-art performance in time series forecasting. However, forecasters based on Transformers are still suffering…

Machine Learning · Computer Science 2024-11-06 Kun Yi , Jingru Fei , Qi Zhang , Hui He , Shufeng Hao , Defu Lian , Wei Fan
‹ Prev 1 4 5 6 7 8 10 Next ›