中文
相关论文

相关论文: An Investigation on Combining Geometry and Consist…

200 篇论文

Phase retrieval is a problem encountered not only in speech and audio processing, but in many other fields such as optics. Iterative algorithms based on non-convex set projections are effective and frequently used for retrieving the phase…

音频与语音处理 · 电气工程与系统科学 2022-11-10 Tal Peer , Simon Welker , Timo Gerkmann

This paper presents a novel phase reconstruction method (only from a given amplitude spectrogram) by combining a signal-processing-based approach and a deep neural network (DNN). To retrieve a time-domain signal from its amplitude…

声音 · 计算机科学 2019-03-12 Yoshiki Masuyama , Kohei Yatabe , Yuma Koizumi , Yasuhiro Oikawa , Noboru Harada

The phase retrieval problem is found in various areas of applications of engineering and applied physics. It is also a very active field of research in mathematics, signal processing and machine learning. In this paper, we present an…

最优化与控制 · 数学 2023-04-26 Rossen Nenov , Dang-Khoa Nguyen , Peter Balazs

The performance of text-to-speech (TTS) systems heavily depends on spectrogram to waveform generation, also known as the speech reconstruction phase. The time required for the same is known as synthesis delay. In this paper, an approach to…

Recent advances in diffusion models have positioned them as powerful generative frameworks for speech synthesis, demonstrating substantial improvements in audio quality and stability. Nevertheless, their effectiveness in vocoders…

声音 · 计算机科学 2025-12-01 Teysir Baoueb , Xiaoyu Bie , Mathieu Fontaine , Gaël Richard

Diffusion models are receiving a growing interest for a variety of signal generation tasks such as speech or music synthesis. WaveGrad, for example, is a successful diffusion model that conditionally uses the mel spectrogram to guide a…

声音 · 计算机科学 2024-02-27 Haocheng Liu , Teysir Baoueb , Mathieu Fontaine , Jonathan Le Roux , Gael Richard

We propose an optimization-based method for reconstructing a time-domain signal from a low-dimensional spectral representation such as a mel-spectrogram. Phase reconstruction has been studied to reconstruct a time-domain signal from the…

声音 · 计算机科学 2023-07-25 Yoshiki Masuyama , Natsuki Ueno , Nobutaka Ono

The recovery of a signal from the magnitudes of its transformation, like the Fourier transform, is known as the phase retrieval problem and is of big relevance in various fields of engineering and applied physics. In this paper, we present…

最优化与控制 · 数学 2023-06-23 Rossen Nenov , Dang-Khoa Nguyen , Peter Balazs , Radu Ioan Bot

Several recent contributions in the field of iterative STFT phase retrieval have demonstrated that the performance of the classical Griffin-Lim method can be considerably improved upon. By using the same projection operators as Griffin-Lim,…

音频与语音处理 · 电气工程与系统科学 2023-09-14 Tal Peer , Simon Welker , Johannes Kolhoff , Timo Gerkmann

A speech enhancement method based on probabilistic geometric approach to spectral subtraction (PGA) performed on short time magnitude spectrum is presented in this paper. A confidence parameter of noise estimation is introduced in the gain…

音频与语音处理 · 电气工程与系统科学 2018-02-15 Md Tauhidul Islam , Celia Shahnaz , Wei-Ping Zhu , M. Omair Ahmad

In this paper, a speech enhancement method based on noise compensation performed on short time magnitude as well phase spectra is presented. Unlike the conventional geometric approach (GA) to spectral subtraction (SS), here the noise…

音频与语音处理 · 电气工程与系统科学 2018-03-09 Md Tauhidul Islam , Udoy Saha , K. T. Shahid , Ahmed Bin Hussain , Celia Shahnaz

In this work, we propose a novel consistency-preserving loss function for recovering the phase information in the context of phase reconstruction (PR) and speech enhancement (SE). Different from conventional techniques that directly…

音频与语音处理 · 电气工程与系统科学 2024-09-25 Pin-Jui Ku , Chun-Wei Ho , Hao Yen , Sabato Marco Siniscalchi , Chin-Hui Lee

Alignment with human preference is a desired property of large language models (LLMs). Currently, the main alignment approach is based on reinforcement learning from human feedback (RLHF). Despite the effectiveness of RLHF, it is intricate…

计算与语言 · 计算机科学 2024-04-16 Geyang Guo , Ranchi Zhao , Tianyi Tang , Wayne Xin Zhao , Ji-Rong Wen

The problem of recovering a signal from the magnitude of its short-time Fourier transform (STFT) is a longstanding one in audio signal processing. Existing approaches rely on heuristics that often perform poorly because of the nonconvexity…

应用统计 · 统计学 2012-09-11 Dennis L. Sun , Julius O. Smith

For the constrained LiGME model, a nonconvexly regularized least squares estimation model, we present an iterative algorithm of guaranteed convergence to its globally optimal solution. The proposed algorithm can deal with two different…

最优化与控制 · 数学 2024-04-05 Wataru Yata , Isao Yamada

In speech enhancement (SE), phase estimation is important for perceptual quality, so many methods take clean speech's complex short-time Fourier transform (STFT) spectrum or the complex ideal ratio mask (cIRM) as the learning target. To…

音频与语音处理 · 电气工程与系统科学 2023-10-12 Yuewei Zhang , Huanbin Zou , Jie Zhu

Time-Scale Modification (TSM) of speech aims to alter the playback rate of audio without changing its pitch. While classical methods like Waveform Similarity-based Overlap-Add (WSOLA) provide strong baselines, they often introduce artifacts…

音频与语音处理 · 电气工程与系统科学 2025-10-06 Dyah A. M. G. Wisnu , Ryandhimas E. Zezario , Stefano Rini , Fo-Rui Li , Yan-Tsung Peng , Hsin-Min Wang , Yu Tsao

Many applications of generalised linear models (GLMs) can be improved by applying constraints that impose assumptions on the associations or improve consistency of the estimators. Yet, there are still barriers to the implementation and…

统计方法学 · 统计学 2026-02-19 Pierre Masselot , Devon Nenon , Jacopo Vanoli , Zaid Chalabi , Antonio Gasparrini

Automatic Speech Recognition (ASR) systems remain prone to errors that affect downstream applications. In this paper, we propose LIR-ASR, a heuristic optimized iterative correction framework using LLMs, inspired by human auditory…

音频与语音处理 · 电气工程与系统科学 2025-09-23 Yutong Liu , Ziyue Zhang , Cheng Huang , Yongbin Yu , Xiangxiang Wang , Yuqing Cai , Nyima Tashi

This paper presents a novel speech phase prediction model which predicts wrapped phase spectra directly from amplitude spectra by neural networks. The proposed model is a cascade of a residual convolutional network and a parallel estimation…

声音 · 计算机科学 2023-02-17 Yang Ai , Zhen-Hua Ling
‹ 上一页 1 2 3 10 下一页 ›