中文
相关论文

相关论文: Mel-Spectrogram Inversion via Alternating Directio…

200 篇论文

Developing a versatile deep neural network to model music audio is crucial in MIR. This task is challenging due to the intricate spectral variations inherent in music signals, which convey melody, harmonics, and timbres of diverse…

声音 · 计算机科学 2024-09-10 Ju-Chiang Wang , Wei-Tsung Lu , Jitong Chen

This paper describes Mixer-TTS, a non-autoregressive model for mel-spectrogram generation. The model is based on the MLP-Mixer architecture adapted for speech synthesis. The basic Mixer-TTS contains pitch and duration predictors, with the…

音频与语音处理 · 电气工程与系统科学 2021-10-25 Oktai Tatanov , Stanislav Beliaev , Boris Ginsburg

Variational mode decomposition (VMD) and its extensions like Multivariate VMD (MVMD) decompose signals into ensembles of band-limited modes with narrow central frequencies. These methods utilize Fourier transformations to shift signals…

Hybrid spectral CT integrates energy integrating detectors (EID) and photon counting detectors (PCD) into a single system, combining the large field-of-view advantage of EID with the high energy and spatial resolution of PCD. This…

医学物理 · 物理学 2024-10-25 Huiying Pan , Jianing Sun , Xu Jiang , Xing Zhao

Phase noise correction is crucial to exploit full advantage of orthogonal frequency division multiplexing (OFDM) in modern high-data-rate communications. OFDM channel estimation with simultaneous phase noise compensation has therefore drawn…

信息论 · 计算机科学 2017-04-25 Zhongju Wang , Prabhu Babu , Daniel P. Palomar

Resampling is a common technique applied in digital signal processing. Based on the Fast Fourier Transformation (FFT), we apply an optimization called here the LMN method to achieve fast and robust re-sampling. In addition to performance…

高能物理 - 实验 · 物理学 2024-09-16 Lino Gerlach , Wenqiang Gu , Nitish Nayak , Xin Qian , Brett Viren

There are many selected mapping (SLM) schemes to reduce the peak-to-average power ratio (PAPR) of orthogonal frequency division multiplexing (OFDM) signals. Beginning with the conventional SLM scheme, there have been proposed many…

信息论 · 计算机科学 2012-09-03 Kee-Hoon Kim , Hyun-Seung Joo , Jong-Seon No , Dong-Joon Shin

In this paper, we propose a novel family of windowing technique to compute Mel Frequency Cepstral Coefficient (MFCC) for automatic speaker recognition from speech. The proposed method is based on fundamental property of discrete time…

计算机视觉与模式识别 · 计算机科学 2015-06-05 Md. Sahidullah , Goutam Saha

Fourier phase retrieval (FPR) is a challenging task widely used in various applications. It involves recovering an unknown signal from its Fourier phaseless measurements. FPR with few measurements is important for reducing time and hardware…

图像与视频处理 · 电气工程与系统科学 2023-07-19 Liyuan Ma , Hongxia Wang , Ningyi Leng , Ziyang Yuan

A novel formulation of the hyperspectral broadband phase retrieval is developed for the scenario where both object and modulation phase masks are spectrally varying. The proposed algorithm is based on a complex domain version of the…

图像与视频处理 · 电气工程与系统科学 2021-05-18 Vladimir Katkovnik , Igor Shevkunov , Karen Egiazarian

Inverse design of metasurfaces for the joint optimization of optical modulation and algorithmic decoding in computational optics presents significant challenges, especially in applications such as hyperspectral imaging. We introduce a…

图像与视频处理 · 电气工程与系统科学 2025-10-28 Rongzhou Chen , Haitao Nie , Shuo Zhu , Yaping Zhao , Chutian Wang , Edmund Y. Lam

We revisit the classical problem of Fourier-sparse signal reconstruction -- a variant of the \emph{Set Query} problem -- which asks to efficiently reconstruct (a subset of) a $d$-dimensional Fourier-sparse signal ($\|\hat{x}(t)\|_0 \leq…

数据结构与算法 · 计算机科学 2023-11-21 Yeqi Gao , Zhao Song , Baocheng Sun , Omri Weinstein , Ruizhe Zhang

This work introduces MELA-TTS, a novel joint transformer-diffusion framework for end-to-end text-to-speech synthesis. By autoregressively generating continuous mel-spectrogram frames from linguistic and speaker conditions, our architecture…

音频与语音处理 · 电气工程与系统科学 2026-01-27 Keyu An , Zhiyu Zhang , Changfeng Gao , Yabin Li , Zhendong Peng , Haoxu Wang , Zhihao Du , Han Zhao , Zhifu Gao , Xiangang Li

In this paper, multi-stream transmission in interference networks aided by multiple amplify-and-forward (AF) relays in the presence of direct links is considered. The objective is to minimize the sum power of transmitters and relays by…

信号处理 · 电气工程与系统科学 2018-12-31 Cenk M. Yetis , Ronald Y. Chang

The alternating direction method of multipliers (ADMM) is a widely used method for solving many convex minimization models arising in signal and image processing. In this paper, we propose an inertial ADMM for solving a two-block separable…

最优化与控制 · 数学 2021-04-02 Yang Yang , Yuchao Tang

In this paper, we propose a new stochastic alternating direction method of multipliers (ADMM) algorithm, which incrementally approximates the full gradient in the linearized ADMM formulation. Besides having a low per-iteration complexity as…

机器学习 · 计算机科学 2013-08-19 Leon Wenliang Zhong , James T. Kwok

In this paper we propose an iterative method using alternating direction method of multipliers (ADMM) strategy to solve linear inverse problems in Hilbert spaces with general convex penalty term. When the data is given exactly, we give a…

数值分析 · 数学 2016-01-13 Yuling Jiao , Qinian Jin , Xiliang Lu , Weijie Wang

Separating vocal elements from musical tracks is a longstanding challenge in audio signal processing. This study tackles the distinct separation of vocal components from musical spectrograms. We employ the Short Time Fourier Transform…

声音 · 计算机科学 2024-05-31 Adam Sorrenti

This paper addresses the problem of speech separation and enhancement from multichannel convolutive and noisy mixtures, \emph{assuming known mixing filters}. We propose to perform the speech separation and enhancement task in the short-time…

声音 · 计算机科学 2019-01-31 Xiaofei Li , Laurent Girin , Sharon Gannot , Radu Horaud

In time series classification and regression, signals are typically mapped into some intermediate representation used for constructing models. Since the underlying task is often insensitive to time shifts, these representations are required…

声音 · 计算机科学 2019-07-16 Joakim Andén , Vincent Lostanlen , Stéphane Mallat