中文
相关论文

相关论文: Phase Vocoder Done Right

200 篇论文

We learn audio representations by solving a novel self-supervised learning task, which consists of predicting the phase of the short-time Fourier transform from its magnitude. A convolutional encoder is used to map the magnitude spectrum of…

音频与语音处理 · 电气工程与系统科学 2019-10-29 Félix de Chaumont Quitry , Marco Tagliasacchi , Dominik Roblek

The short-time Fourier transform (STFT) represents a window of audio samples as a set of complex coefficients. These are advantageously viewed as magnitudes and phases and the overall distribution of phases is very often assumed to be…

音频与语音处理 · 电气工程与系统科学 2024-09-16 Stephen D. Voran

This paper proposes a spectral-domain perceptual weighting technique for Parallel WaveGAN-based text-to-speech (TTS) systems. The recently proposed Parallel WaveGAN vocoder successfully generates waveform sequences using a fast…

音频与语音处理 · 电气工程与系统科学 2021-01-20 Eunwoo Song , Ryuichi Yamamoto , Min-Jae Hwang , Jin-Seob Kim , Ohsung Kwon , Jae-Min Kim

In this work, we propose a novel consistency-preserving loss function for recovering the phase information in the context of phase reconstruction (PR) and speech enhancement (SE). Different from conventional techniques that directly…

音频与语音处理 · 电气工程与系统科学 2024-09-25 Pin-Jui Ku , Chun-Wei Ho , Hao Yen , Sabato Marco Siniscalchi , Chin-Hui Lee

We consider the classical 1D phase retrieval problem. In order to overcome the difficulties associated with phase retrieval from measurements of the Fourier magnitude, we treat recovery from the magnitude of the short-time Fourier transform…

信息论 · 计算机科学 2015-06-23 Yonina C. Eldar , Pavel Sidorenko , Dustin G. Mixon , Shaby Barel , Oren Cohen

In most current approaches of speech processing, information is extracted from the magnitude spectrum. However recent perceptual studies have underlined the importance of the phase component. The goal of this paper is to investigate the…

声音 · 计算机科学 2020-01-03 Thomas Drugman , Thomas Dubuisson , Thierry Dutoit

This study investigates phase reconstruction for deep learning based monaural talker-independent speaker separation in the short-time Fourier transform (STFT) domain. The key observation is that, for a mixture of two sources, with their…

声音 · 计算机科学 2018-11-26 Zhong-Qiu Wang , Ke Tan , DeLiang Wang

We address the problem of recovering a signal (up to global phase) from its short-time Fourier transform (STFT) magnitude measurements. This problem arises in several applications, including optical imaging and speech processing. In this…

信息论 · 计算机科学 2015-10-06 Tamir Bendory , Yonina C. Eldar

To date, various speech technology systems have adopted the vocoder approach, a method for synthesizing speech waveform that shows a major role in the performance of statistical parametric speech synthesis. WaveNet one of the best models…

声音 · 计算机科学 2021-06-15 Mohammed Salah Al-Radhi , Tamás Gábor Csapó , Csaba Zainkó , Géza Németh

State-of-the-art statistical parametric speech synthesis (SPSS) generally uses a vocoder to represent speech signals and parameterize them into features for subsequent modeling. Magnitude spectrum has been a dominant feature over the years.…

声音 · 计算机科学 2015-10-08 Bo Fan , Siu Wa Lee , Xiaohai Tian , Lei Xie , Minghui Dong

This paper investigates the effect of oscillator phase noise in orthogonal time frequency space (OTFS) systems. The paper provides in-depth analysis of the interference due to phase noise in the delay-Doppler domain and derives expressions…

信号处理 · 电气工程与系统科学 2026-05-15 Stephen McWade , Arman Farhang

In this paper, we propose a new method for the accurate estimation and tracking of formants in speech signals using time-varying quasi-closed-phase (TVQCP) analysis. Conventional formant tracking methods typically adopt a two-stage…

音频与语音处理 · 电气工程与系统科学 2023-09-01 Dhananjaya Gowda , Sudarsana Reddy Kadiri , Brad Story , Paavo Alku

An estimation method is presented for polynomial phase signals, i.e., those adopting the form of a complex exponential whose phase is polynomial in its indices. Transcending the scope of existing techniques, the proposed estimator can…

信号处理 · 电气工程与系统科学 2024-11-12 Heedong Do , Namyoon Lee , Angel Lozano

In this paper, we present a vocoder-free framework for audio super-resolution that employs a flow matching generative model to capture the conditional distribution of complex-valued spectral coefficients. Unlike conventional two-stage…

音频与语音处理 · 电气工程与系统科学 2026-02-06 Woongjib Choi , Sangmin Lee , Hyungseob Lim , Hong-Goo Kang

Calibration is crucial for ensuring the performance of phased array since amplitude-phase imbalance between elements results in significant performance degradation. While amplitude-only calibration methods offer advantages when phase…

信号处理 · 电气工程与系统科学 2025-04-24 Shiyuan Li , Yuyue Zhou , Chi Zhang , Liang Kong , Kebin Liu , Yihan Xie , Chong He

The focus of this article is on shape and topology optimization of transient vibroacoustic problems. The main contribution is a transient problem formulation that enables optimization over wide ranges of frequencies with complex signals,…

最优化与控制 · 数学 2023-06-28 Cetin B. Dilgen , Niels Aage

The problem of phase-noise compensation for correlated phase noise in coded multichannel optical transmission is investigated. To that end, a simple multichannel phase-noise model is considered and the maximum a posteriori detector for this…

信息论 · 计算机科学 2019-10-15 Arni F. Alfredsson , Erik Agrell , Henk Wymeersch

The technique of transforming voices in order to hide the real identity of a speaker is called voice disguise, among which automatic voice disguise (AVD) by modifying the spectral and temporal characteristics of voices with miscellaneous…

音频与语音处理 · 电气工程与系统科学 2020-09-16 Linlin Zheng , Jiakang Li , Meng Sun , Xiongwei Zhang , Thomas Fang Zheng

Time-resolved optical filtering (TROF) measures the spectrogram or sonogram by a fast photodiode followed a tunable narrowband optical filter. For periodic signal and to match the sonogram, numerical TROF algorithm is used to find the…

光学 · 物理学 2013-01-15 KeangPo Ho , Hsi-Cheng Wang , Hau-Kai Chen , Cheng-Chen Wu

This paper defines the novel task of drum-to-vocal percussion (VP) sound conversion. VP imitates percussion instruments through human vocalization and is frequently employed in contemporary a cappella music. It exhibits acoustic properties…