English
Related papers

Related papers: Phase Vocoder Done Right

200 papers

Positive time varying frequency representation for transient signals has been a hearty desire of signal analysts due to its theoretical and practical importance. During approximately the last two decades there has formulated a signal…

Complex Variables · Mathematics 2018-05-17 Tao Qian

This paper presents a novel neural vocoder named APNet which reconstructs speech waveforms from acoustic features by predicting amplitude and phase spectra directly. The APNet vocoder is composed of an amplitude spectrum predictor (ASP) and…

Sound · Computer Science 2023-05-16 Yang Ai , Zhen-Hua Ling

In recent text-to-speech synthesis and voice conversion systems, a mel-spectrogram is commonly applied as an intermediate representation, and the necessity for a mel-spectrogram vocoder is increasing. A mel-spectrogram vocoder must solve…

Sound · Computer Science 2022-03-07 Takuhiro Kaneko , Kou Tanaka , Hirokazu Kameoka , Shogo Seki

This paper explores the innovative application of the Fractional Fourier Transform (FrFT) in sound synthesis, highlighting its potential to redefine time-frequency analysis in audio processing. As an extension of the classical Fourier…

Sound · Computer Science 2025-06-12 Esteban Gutiérrez , Rodrigo Cádiz , Carlos Sing Long , Frederic Font , Xavier Serra

Phase-retrieval techniques aim to recover the original signal from just the modulus of its Fourier transform, which is usually much easier to measure than its phase, but the standard iterative techniques tend to fail if only part of the…

Image and Video Processing · Electrical Eng. & Systems 2023-07-06 Giovanni Pellegrini , Jacopo Bertolotti

With the scale of vision Transformer-based models continuing to grow, finetuning these large-scale pretrained models for new tasks has become increasingly parameter-intensive. Visual prompt tuning is introduced as a parameter-efficient…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Runjia Zeng , Cheng Han , Qifan Wang , Chunshu Wu , Tong Geng , Lifu Huang , Ying Nian Wu , Dongfang Liu

One of the obstacles in many-to-many voice conversion is the requirement of the parallel training data, which contain pairs of utterances with the same linguistic content spoken by different speakers. Since collecting such parallel data is…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-04 Keonnyeong Lee , In-Chul Yoo , Dongsuk Yook

In this work, we investigate the effectiveness of two techniques for improving variational autoencoder (VAE) based voice conversion (VC). First, we reconsider the relationship between vocoder features extracted using the high quality…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-09 Wen-Chin Huang , Yi-Chiao Wu , Chen-Chou Lo , Patrick Lumban Tobing , Tomoki Hayashi , Kazuhiro Kobayashi , Tomoki Toda , Yu Tsao , Hsin-Min Wang

The residual vector quantization (RVQ) technique plays a central role in recent advances in neural audio codecs. These models effectively synthesize high-fidelity audio from a limited number of codes due to the hierarchical structure among…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Hyeongju Kim , Junhyeok Lee , Jacob Morton , Juheon Lee , Jinhyeok Yang

Fourier ptychographic microscopy (FPM) is a recently developed imaging modality that uses angularly varying illumination to extend a system performance beyond the limit defined by its optical elements. The FPM technique applies a novel…

Optics · Physics 2014-05-02 Xiaoze Ou , Roarke Horstmeyer , Changhuei Yang , Guoan Zheng

Voice disorders significantly impact patient quality of life, yet non-invasive automated diagnosis remains under-explored due to both the scarcity of pathological voice data, and the variability in recording sources. This work introduces…

In this paper, we consider an OFDM radio link corrupted by oscillator phase noise in the receiver, namely the problem of estimating and compensating for the impairment. To lessen the computational burden and delay incurred onto the…

Information Theory · Computer Science 2017-04-05 Pramod Mathecken , Taneli Riihonen , Stefan Werner , Risto Wichman

Voice conversion is a task of synthesizing an utterance with target speaker's voice while maintaining linguistic information of the source utterance. While a speaker can produce varying utterances from a single script with different…

Sound · Computer Science 2025-04-17 Soobin Suh , Dabi Ahn , Heewoong Park , Jonghun Park

Fast Fourier Transform (FFT) relies on the HRV frequency-domain analysis techniques. It requires re-sampling of the inherently unevenly sampled heartbeat time-series (RR tachogram) to produce an evenly sampled time series of the heartbeat.…

Medical Physics · Physics 2022-08-04 Amin Gasmi

Fourier-transform spectroscopy (FTS) has been widely used as a standard analytical technique over the past half-century. FTS is a simple and robust autocorrelation-based technique that is compatible with both temporally coherent and…

Instrumentation and Detectors · Physics 2018-11-21 Kazuki Hashimoto , Takuro Ideguchi

Recent synthetic speech detection models typically adapt a pre-trained SSL model via finetuning, which is computationally demanding. Parameter-Efficient Fine-Tuning (PEFT) offers an alternative. However, existing methods lack the specific…

Sound · Computer Science 2025-10-30 Yassine El Kheir , Fabian Ritter-Guttierez , Arnab Das , Tim Polzehl , Sebastian Möller

The voting method, an ensemble approach for fundamental frequency estimation, is empirically known for its robustness but lacks thorough investigation. This paper provides a principled analysis and improvement of this technique. First, we…

Sound · Computer Science 2026-02-03 Junya Koguchi , Tomoki Koriyama

In this letter, we present a novel and extremely fast steganalysis method of Voice over IP (VoIP) streams, driven by the need for a quick and accurate detection of possible steganography in VoIP streams. We firstly analyzed the correlations…

Multimedia · Computer Science 2020-04-22 Hao Yang , ZhongLiang Yang , YongJian Bao , YongFeng Huang

Many problems on signal processing reduce to nonparametric function estimation. We propose a new methodology, piecewise convex fitting (PCF), and give a two-stage adaptive estimate. In the first stage, the number and location of the change…

Methodology · Statistics 2020-02-18 Kurt Riedel

A non-iterative method for the construction of the Short-Time Fourier Transform (STFT) phase from the magnitude is presented. The method is based on the direct relationship between the partial derivatives of the phase and the logarithm of…

Sound · Computer Science 2019-03-27 Zdeněk Průša , Peter Balazs , Peter L. Søndergaard
‹ Prev 1 3 4 5 6 7 10 Next ›