English
Related papers

Related papers: Phase Vocoder Done Right

200 papers

Neural vocoders are now being used in a wide range of speech processing applications. In many of those applications, the vocoder can be the most complex component, so finding lower complexity algorithms can lead to significant practical…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-06 Jean-Marc Valin , Ahmed Mustafa , Jan Büthe

Deep generative models applied to audio have improved by a large margin the state-of-the-art in many speech and music related tasks. However, as raw waveform modelling remains an inherently difficult task, audio generative models are either…

Machine Learning · Computer Science 2021-12-16 Antoine Caillon , Philippe Esling

This note shows how to align a periodic signal with its the Fourier transform by means of frequency or time scaling. This may be useful in developing new algorithms, e.g. for pitch estimation. This note also convolves the signals and the…

Signal Processing · Electrical Eng. & Systems 2023-09-19 Matthew R. Flax , W. Harvey Holmes

Many audio signal processing methods are formulated in the time-frequency (T-F) domain which is obtained by the short-time Fourier transform (STFT). The properties of the STFT are fully characterized by window function, number of frequency…

Signal Processing · Electrical Eng. & Systems 2019-02-05 Tsubasa Kusano , Yoshiki Masuyama , Kohei Yatabe , Yasuhiro Oikawa

As a consequence of a general trend in the physics of oscillators and clocks towards optics, phase and frequency metrology is rapidly moving to optics too. Yet, optics is not replacing the traditional radio-frequency (RF) and microwave…

Instrumentation and Detectors · Physics 2017-08-23 Enrico Rubiola , Vincent Giordano , Kirill Volyanskiy , Laurent Larger

Time-frequency audio source separation is usually achieved by estimating the short-time Fourier transform (STFT) magnitude of each source, and then applying a phase recovery algorithm to retrieve time-domain signals. In particular, the…

Sound · Computer Science 2021-02-10 Paul Magron , Pierre-Hugo Vial , Thomas Oberlin , Cédric Févotte

End-to-end singing voice synthesis (SVS) model VISinger can achieve better performance than the typical two-stage model with fewer parameters. However, VISinger has several problems: text-to-phase problem, the end-to-end model learns the…

Sound · Computer Science 2022-11-08 Yongmao Zhang , Heyang Xue , Hanzhao Li , Lei Xie , Tingwei Guo , Ruixiong Zhang , Caixia Gong

While universal vocoders have achieved proficient waveform generation across diverse voices, their integration into text-to-speech (TTS) tasks often results in degraded synthetic quality. To address this challenge, we present a novel…

Sound · Computer Science 2024-09-05 Jeongmin Liu , Eunwoo Song

Vibrational resonance (VR) is a nonlinear phenomenon in which the system response to a weak signal can be resonantly enhanced by applying a high-frequency modulation signal with an appropriate amplitude. The majority of VR research has…

Optics · Physics 2024-03-01 Yinuo Wang , Shan Wu , Cuicui Li , Zhenglu Duan , Min Xie , Bixuan Fan

This paper presents a Bayesian approach to symbol and phase inference in a phase-unsynchronized digital receiver. It primarily extends [Quinn 2011] to the multi-symbol case, using the variational Bayes (VB) approximation to deal with the…

Applications · Statistics 2015-03-13 Arijit Das , Anthony Quinn

Phase Coded (PC) waveforms possess desirable Auto-Correlation Function (ACF) properties for use in radar and sonar systems. However, their spectra possess high spectral leakage due to the abrupt phase transitions between the chips in the…

Signal Processing · Electrical Eng. & Systems 2022-01-21 David A. Hague

The advancements of AI-synthesized human voices have introduced a growing threat of impersonation and disinformation. It is therefore of practical importance to developdetection methods for synthetic human voices. This work proposes a new…

Sound · Computer Science 2023-04-28 Chengzhe Sun , Shan Jia , Shuwei Hou , Ehab AlBadawy , Siwei Lyu

Data-driven decompositions of Particle Image Velocimetry (PIV) measurements are widely used for a variety of purposes, including the detection of coherent features (e.g., vortical structures), filtering operations (e.g., outlier removal or…

Fluid Dynamics · Physics 2020-08-26 Miguel A Mendez , David Hess , Bo B Watz , Jean-Marie Buchlin

Recent advances in zero-shot text-to-speech (TTS), driven by language models, diffusion models and masked generation, have achieved impressive naturalness in speech synthesis. Nevertheless, stability and fidelity remain key challenges,…

Sound · Computer Science 2025-10-24 Hualei Wang , Na Li , Chuke Wang , Shu Wu , Zhifeng Li , Dong Yu

Neural audio coding has emerged as a vivid research direction by promising good audio quality at very low bitrates unachievable by classical coding techniques. Here, end-to-end trainable autoencoder-like models represent the state of the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-20 Andreas Brendel , Nicola Pia , Kishan Gupta , Lyonel Behringer , Guillaume Fuchs , Markus Multrus

Phase retrieval(PR) problem is a kind of ill-condition inverse problem which can be found in various of applications. Utilizing the sparse priority, an algorithm called SWF(Sparse Wirtinger Flow) is proposed in this paper to deal with…

Information Theory · Computer Science 2017-04-12 Ziyang Yuan , Qi Wang , Hongxia Wang

Tone Transfer is a novel deep-learning technique for interfacing a sound source with a synthesizer, transforming the timbre of audio excerpts while keeping their musical form content. Due to its good audio quality results and continuous…

Sound · Computer Science 2023-10-10 Franco Caspe , Andrew McPherson , Mark Sandler

As a foundational technology for intelligent human-computer interaction, voice conversion (VC) seeks to transform speech from any source timbre into any target timbre. Traditional voice conversion methods based on Generative Adversarial…

Sound · Computer Science 2025-06-11 Wenhan Yao , Fen Xiao , Xiarun Chen , Jia Liu , YongQiang He , Weiping Wen

Entertainment-oriented singing voice synthesis (SVS) requires a vocoder to generate high-fidelity (e.g. 48kHz) audio. However, most text-to-speech (TTS) vocoders cannot reconstruct the waveform well in this scenario. In this paper, we…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-19 Chunhui Wang , Chang Zeng , Jun Chen , Xing He

Low frequency oscillator (LFO) driven audio effects such as phaser, flanger, and chorus, modify an input signal using time-varying filters and delays, resulting in characteristic sweeping or widening effects. It has been shown that these…

‹ Prev 1 8 9 10 Next ›