中文
相关论文

相关论文: Phase Vocoder Done Right

200 篇论文

Positive time varying frequency representation for transient signals has been a hearty desire of signal analysts due to its theoretical and practical importance. During approximately the last two decades there has formulated a signal…

复变函数 · 数学 2018-05-17 Tao Qian

This paper presents a novel neural vocoder named APNet which reconstructs speech waveforms from acoustic features by predicting amplitude and phase spectra directly. The APNet vocoder is composed of an amplitude spectrum predictor (ASP) and…

声音 · 计算机科学 2023-05-16 Yang Ai , Zhen-Hua Ling

In recent text-to-speech synthesis and voice conversion systems, a mel-spectrogram is commonly applied as an intermediate representation, and the necessity for a mel-spectrogram vocoder is increasing. A mel-spectrogram vocoder must solve…

声音 · 计算机科学 2022-03-07 Takuhiro Kaneko , Kou Tanaka , Hirokazu Kameoka , Shogo Seki

This paper explores the innovative application of the Fractional Fourier Transform (FrFT) in sound synthesis, highlighting its potential to redefine time-frequency analysis in audio processing. As an extension of the classical Fourier…

声音 · 计算机科学 2025-06-12 Esteban Gutiérrez , Rodrigo Cádiz , Carlos Sing Long , Frederic Font , Xavier Serra

Phase-retrieval techniques aim to recover the original signal from just the modulus of its Fourier transform, which is usually much easier to measure than its phase, but the standard iterative techniques tend to fail if only part of the…

图像与视频处理 · 电气工程与系统科学 2023-07-06 Giovanni Pellegrini , Jacopo Bertolotti

With the scale of vision Transformer-based models continuing to grow, finetuning these large-scale pretrained models for new tasks has become increasingly parameter-intensive. Visual prompt tuning is introduced as a parameter-efficient…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Runjia Zeng , Cheng Han , Qifan Wang , Chunshu Wu , Tong Geng , Lifu Huang , Ying Nian Wu , Dongfang Liu

One of the obstacles in many-to-many voice conversion is the requirement of the parallel training data, which contain pairs of utterances with the same linguistic content spoken by different speakers. Since collecting such parallel data is…

音频与语音处理 · 电气工程与系统科学 2020-02-04 Keonnyeong Lee , In-Chul Yoo , Dongsuk Yook

In this work, we investigate the effectiveness of two techniques for improving variational autoencoder (VAE) based voice conversion (VC). First, we reconsider the relationship between vocoder features extracted using the high quality…

音频与语音处理 · 电气工程与系统科学 2019-07-09 Wen-Chin Huang , Yi-Chiao Wu , Chen-Chou Lo , Patrick Lumban Tobing , Tomoki Hayashi , Kazuhiro Kobayashi , Tomoki Toda , Yu Tsao , Hsin-Min Wang

The residual vector quantization (RVQ) technique plays a central role in recent advances in neural audio codecs. These models effectively synthesize high-fidelity audio from a limited number of codes due to the hierarchical structure among…

音频与语音处理 · 电气工程与系统科学 2025-09-24 Hyeongju Kim , Junhyeok Lee , Jacob Morton , Juheon Lee , Jinhyeok Yang

Fourier ptychographic microscopy (FPM) is a recently developed imaging modality that uses angularly varying illumination to extend a system performance beyond the limit defined by its optical elements. The FPM technique applies a novel…

光学 · 物理学 2014-05-02 Xiaoze Ou , Roarke Horstmeyer , Changhuei Yang , Guoan Zheng

Voice disorders significantly impact patient quality of life, yet non-invasive automated diagnosis remains under-explored due to both the scarcity of pathological voice data, and the variability in recording sources. This work introduces…

In this paper, we consider an OFDM radio link corrupted by oscillator phase noise in the receiver, namely the problem of estimating and compensating for the impairment. To lessen the computational burden and delay incurred onto the…

信息论 · 计算机科学 2017-04-05 Pramod Mathecken , Taneli Riihonen , Stefan Werner , Risto Wichman

Voice conversion is a task of synthesizing an utterance with target speaker's voice while maintaining linguistic information of the source utterance. While a speaker can produce varying utterances from a single script with different…

声音 · 计算机科学 2025-04-17 Soobin Suh , Dabi Ahn , Heewoong Park , Jonghun Park

Fast Fourier Transform (FFT) relies on the HRV frequency-domain analysis techniques. It requires re-sampling of the inherently unevenly sampled heartbeat time-series (RR tachogram) to produce an evenly sampled time series of the heartbeat.…

医学物理 · 物理学 2022-08-04 Amin Gasmi

Fourier-transform spectroscopy (FTS) has been widely used as a standard analytical technique over the past half-century. FTS is a simple and robust autocorrelation-based technique that is compatible with both temporally coherent and…

仪器与探测器 · 物理学 2018-11-21 Kazuki Hashimoto , Takuro Ideguchi

Recent synthetic speech detection models typically adapt a pre-trained SSL model via finetuning, which is computationally demanding. Parameter-Efficient Fine-Tuning (PEFT) offers an alternative. However, existing methods lack the specific…

声音 · 计算机科学 2025-10-30 Yassine El Kheir , Fabian Ritter-Guttierez , Arnab Das , Tim Polzehl , Sebastian Möller

The voting method, an ensemble approach for fundamental frequency estimation, is empirically known for its robustness but lacks thorough investigation. This paper provides a principled analysis and improvement of this technique. First, we…

声音 · 计算机科学 2026-02-03 Junya Koguchi , Tomoki Koriyama

In this letter, we present a novel and extremely fast steganalysis method of Voice over IP (VoIP) streams, driven by the need for a quick and accurate detection of possible steganography in VoIP streams. We firstly analyzed the correlations…

多媒体 · 计算机科学 2020-04-22 Hao Yang , ZhongLiang Yang , YongJian Bao , YongFeng Huang

Many problems on signal processing reduce to nonparametric function estimation. We propose a new methodology, piecewise convex fitting (PCF), and give a two-stage adaptive estimate. In the first stage, the number and location of the change…

统计方法学 · 统计学 2020-02-18 Kurt Riedel

A non-iterative method for the construction of the Short-Time Fourier Transform (STFT) phase from the magnitude is presented. The method is based on the direct relationship between the partial derivatives of the phase and the logarithm of…

声音 · 计算机科学 2019-03-27 Zdeněk Průša , Peter Balazs , Peter L. Søndergaard