中文
相关论文

相关论文: Experimental investigation on STFT phase represent…

200 篇论文

Music source separation with deep neural networks typically relies only on amplitude features. In this paper we show that additional phase features can improve the separation performance. Using the theoretical relationship between STFT…

Maximum Voiced Frequency (MVF) is used in various speech models as the spectral boundary separating periodic and aperiodic components during the production of voiced sounds. Recent studies have shown that its proper estimation and modeling…

音频与语音处理 · 电气工程与系统科学 2020-06-02 Thomas Drugman , Yannis Stylianou

Dysarthric speech recognition (DSR) presents a formidable challenge due to inherent inter-speaker variability, leading to severe performance degradation when applying DSR models to new dysarthric speakers. Traditional speaker adaptation…

声音 · 计算机科学 2024-09-25 Shiyao Wang , Shiwan Zhao , Jiaming Zhou , Aobo Kong , Yong Qin

In this paper, a speech enhancement method based on noise compensation performed on short time magnitude as well phase spectra is presented. Unlike the conventional geometric approach (GA) to spectral subtraction (SS), here the noise…

音频与语音处理 · 电气工程与系统科学 2018-03-09 Md Tauhidul Islam , Udoy Saha , K. T. Shahid , Ahmed Bin Hussain , Celia Shahnaz

Recent advancements in deep learning have significantly impacted the field of speech signal processing, particularly in the analysis and manipulation of complex spectrograms. This survey provides a comprehensive overview of the…

音频与语音处理 · 电气工程与系统科学 2025-10-06 Yuying Xie , Zheng-Hua Tan

Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech by improving the intelligibility and naturalness. This is a challenging task especially for patients with severe dysarthria and speaking in…

声音 · 计算机科学 2024-02-01 Xueyuan Chen , Yuejiao Wang , Xixin Wu , Disong Wang , Zhiyong Wu , Xunying Liu , Helen Meng

The scattering transform is a non-linear signal representation method based on cascaded wavelet transform magnitudes. In this paper we introduce phase scattering, a novel approach where we use phase derivatives in a scattering procedure. We…

声音 · 计算机科学 2024-07-09 Daniel Haider , Peter Balazs , Nicki Holighaus

Dysarthria, a motor speech disorder, affects intelligibility and requires targeted interventions for effective communication. In this work, we investigate automated mispronunciation feedback by collecting a dysarthric speech dataset from…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Seohyun Park , Chitralekha Gupta , Michelle Kah Yian Kwan , Xinhui Fung , Alexander Wenjun Yip , Suranga Nanayakkara

We present STFTCodec, a novel spectral-based neural audio codec that efficiently compresses audio using Short-Time Fourier Transform (STFT). Unlike waveform-based approaches that require large model capacity and substantial memory…

声音 · 计算机科学 2025-03-24 Tao Feng , Zhiyuan Zhao , Yifan Xie , Yuqi Ye , Xiangyang Luo , Xun Guan , Yu Li

In this paper, a novel multitaper modified group delay function-based representation for speech signals is proposed. With a set of phoneme-based experiments, it is shown that the proposed method performs better that an existing multitaper…

声音 · 计算机科学 2018-07-04 K. C. Narendra , R. Kumaraswamy , S. Gurugopinath

Most speech enhancement algorithms make use of the short-time Fourier transform (STFT), which is a simple and flexible time-frequency decomposition that estimates the short-time spectrum of a signal. However, the duration of short STFT…

声音 · 计算机科学 2015-09-03 Scott Wisdom , Thomas Powers , Les Atlas , James Pitton

This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural speech waveform model that predicts raw continuous speech waveform samples directly. Not only amplitude…

音频与语音处理 · 电气工程与系统科学 2018-10-31 Shinji Takaki , Toru Nakashika , Xin Wang , Junichi Yamagishi

Speech synthesis from intracranial EEG (iEEG) signals offers a promising avenue for restoring communication in individuals with severe speech impairments. However, achieving intelligible and natural speech remains challenging due to…

声音 · 计算机科学 2025-08-06 Mohammed Salah Al-Radhi , Géza Németh , Branislav Gerazov

Current anti-spoofing and audio deepfake detection systems use either magnitude spectrogram-based features (such as CQT or Melspectrograms) or raw audio processed through convolution or sinc-layers. Both methods have drawbacks: magnitude…

声音 · 计算机科学 2023-08-24 Nicolas M. Müller , Philip Sperl , Konstantin Böttinger

In this work, we propose a novel consistency-preserving loss function for recovering the phase information in the context of phase reconstruction (PR) and speech enhancement (SE). Different from conventional techniques that directly…

音频与语音处理 · 电气工程与系统科学 2024-09-25 Pin-Jui Ku , Chun-Wei Ho , Hao Yen , Sabato Marco Siniscalchi , Chin-Hui Lee

The decoupling-style concept begins to ignite in the speech enhancement area, which decouples the original complex spectrum estimation task into multiple easier sub-tasks i.e., magnitude-only recovery and the residual complex spectrum…

声音 · 计算机科学 2022-08-02 Guochen Yu , Andong Li , Hui Wang , Yutian Wang , Yuxuan Ke , Chengshi Zheng

Modern neural speech enhancement models usually include various forms of phase information in their training loss terms, either explicitly or implicitly. However, these loss terms are typically designed to reduce the distortion of phase…

声音 · 计算机科学 2022-02-25 Doyeon Kim , Hyewon Han , Hyeon-Kyeong Shin , Soo-Whan Chung , Hong-Goo Kang

Phase information has a significant impact on speech perceptual quality and intelligibility. However, existing speech enhancement methods encounter limitations in explicit phase estimation due to the non-structural nature and wrapping…

音频与语音处理 · 电气工程与系统科学 2024-04-02 Ye-Xin Lu , Yang Ai , Zhen-Hua Ling

Developing and selecting hearing aids is a time consuming process which is simplified by using objective models. Previously, the framework for auditory discrimination experiments (FADE) accurately simulated benefits of hearing aid…

音频与语音处理 · 电气工程与系统科学 2021-02-12 David Hülsmeier , Marc René Schädler , Birger Kollmeier

Purpose: Speech intelligibility is a critical outcome in the assessment and management of dysarthria, yet most research and clinical practices have focused on English, limiting their applicability across languages. This commentary…

计算与语言 · 计算机科学 2025-05-09 Eunjung Yeo , Julie Liss , Visar Berisha , David Mortensen