中文
相关论文

相关论文: Experimental investigation on STFT phase represent…

200 篇论文

Most audio processing pipelines involve transformations that act on fixed-dimensional input representations of audio. For example, when using the Short Time Fourier Transform (STFT) the DFT size specifies a fixed dimension for the input…

音频与语音处理 · 电气工程与系统科学 2022-03-28 Krishna Subramani , Paris Smaragdis

Conventional time-delay neural networks (TDNNs) struggle to handle long-range context, their ability to represent speaker information is therefore limited in long utterances. Existing solutions either depend on increasing model complexity…

声音 · 计算机科学 2023-08-02 Yangfu Li , Jiapan Gan , Xiaodan Lin

Acoustic-to-articulatory inversion (AAI) involves mapping from the acoustic to the articulatory space. Signal-processing features like the MFCCs, have been widely used for the AAI task. For subjects with dysarthric speech, AAI is…

音频与语音处理 · 电气工程与系统科学 2024-02-13 Sarthak Kumar Maharana , Krishna Kamal Adidam , Shoumik Nandi , Ajitesh Srivastava

We study the problem of multimodal physical scene understanding, where an embodied agent needs to find fallen objects by inferring object properties, direction, and distance of an impact sound source. Previous works adopt feed-forward…

机器人学 · 计算机科学 2024-07-17 Jie Yin , Andrew Luo , Yilun Du , Anoop Cherian , Tim K. Marks , Jonathan Le Roux , Chuang Gan

Speaker Verification (SV) systems trained on adults speech often underperform on children's SV due to the acoustic mismatch, and limited children speech data makes fine-tuning not very effective. In this paper, we propose an innovative…

音频与语音处理 · 电气工程与系统科学 2025-08-12 Vishwas M. Shetty , Jiusi Zheng , Abeer Alwan

The problem of recovering a signal from its Fourier magnitude is of paramount importance in various fields of engineering and applied physics. Due to the absence of Fourier phase information, some form of additional information is required…

信息论 · 计算机科学 2016-05-25 Kishore Jaganathan , Yonina C. Eldar , Babak Hassibi

Existing deepfake detection research has primarily focused on scenarios where the manipulated subject is actively speaking, i.e., generating fabricated content by altering the speaker's appearance or voice. However, in realistic interaction…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Miao Liu , Fangda Wei , Jing Wang , Xinyuan Qian

Recent high-performance transformer-based speech enhancement models demonstrate that time domain methods could achieve similar performance as time-frequency domain methods. However, time-domain speech enhancement systems typically receive…

声音 · 计算机科学 2023-10-31 Junhui Li , Pu Wang , Jialu Li , Xinzhe Wang , Youshan Zhang

Short-time Fourier transform (STFT) is used as the front end of many popular successful monaural speech separation methods, such as deep clustering (DPCL), permutation invariant training (PIT) and their various variants. Since the frequency…

声音 · 计算机科学 2019-02-05 Ziqiang Shi , Huibin Lin , Liu Liu , Rujie Liu , Jiqing Han

Dysarthric speech severity assessment typically requires trained clinicians or supervised models built from labelled pathological speech, limiting scalability across languages and clinical settings. We present a training-free method that…

计算与语言 · 计算机科学 2026-04-14 Bernard Muller , Antonio Armando Ortiz Barrañón , LaVonne Roberts

Automated dysarthria detection and severity assessment from speech have attracted significant research attention due to their potential clinical impact. Despite rapid progress in acoustic modeling and deep learning, models still fall short…

音频与语音处理 · 电气工程与系统科学 2025-10-28 Krishna Gurugubelli

Speech deepfake detection (SDD) systems perform well on standard benchmarks datasets but often fail to generalize to expressive and emotional spoofing attacks. Many methods rely on spoof-heavy training data, learning dataset-specific…

音频与语音处理 · 电气工程与系统科学 2026-04-16 Aurosweta Mahapatra , Ismail Rasim Ulgen , Kong Aik Lee , Nicholas Andrews , Berrak Sisman

In this work, we propose a new mathematical vocoder algorithm(modified spectral inversion) that generates a waveform from acoustic features without phase estimation. The main benefit of using our proposed method is that it excludes the…

音频与语音处理 · 电气工程与系统科学 2021-06-17 Hyun Gon Ryu , Jeong-Hoon Kim , Simon See

Speech enhancement techniques based on deep learning have brought significant improvement on speech quality and intelligibility. Nevertheless, a large gain in speech quality measured by objective metrics, such as perceptual evaluation of…

音频与语音处理 · 电气工程与系统科学 2020-07-06 Bo Wu , Meng Yu , Lianwu Chen , Yong Xu , Chao Weng , Dan Su , Dong Yu

In speech enhancement (SE), phase estimation is important for perceptual quality, so many methods take clean speech's complex short-time Fourier transform (STFT) spectrum or the complex ideal ratio mask (cIRM) as the learning target. To…

音频与语音处理 · 电气工程与系统科学 2023-10-12 Yuewei Zhang , Huanbin Zou , Jie Zhu

This paper proposes a novel Stage-wise and Prior-aware Neural Speech Phase Prediction (SP-NSPP) model, which predicts the phase spectrum from input amplitude spectrum by two-stage neural networks. In the initial prior-construction stage, we…

声音 · 计算机科学 2024-10-08 Fei Liu , Yang Ai , Hui-Peng Du , Ye-Xin Lu , Rui-Chen Zheng , Zhen-Hua Ling

A divide and conquer strategy for enhancement of noisy speeches in adverse environments involving lower levels of SNR is presented in this paper, where the total system of speech enhancement is divided into two separate steps. The first…

音频与语音处理 · 电气工程与系统科学 2018-02-09 Md Tauhidul Islam , Celia Shahnaz , Wei-Ping Zhu , M. Omair Ahmad

Deep learning based speech enhancement and source separation systems have recently reached unprecedented levels of quality, to the point that performance is reaching a new ceiling. Most systems rely on estimating the magnitude of a target…

声音 · 计算机科学 2019-06-26 Jonathan Le Roux , Gordon Wichern , Shinji Watanabe , Andy Sarroff , John R. Hershey

Clarifying the neural basis of speech intelligibility is critical for computational neuroscience and digital speech processing. Recent neuroimaging studies have shown that intelligibility modulates cortical activity beyond simple acoustics,…

神经元与认知 · 定量生物学 2025-11-05 Ching-Chih Sung , Shuntaro Suzuki , Francis Pingfan Chien , Komei Sugiura , Yu Tsao

Dysarthric speech reconstruction (DSR) aims to convert dysarthric speech into comprehensible speech while maintaining the speaker's identity. Despite significant advancements, existing methods often struggle with low speech intelligibility…

声音 · 计算机科学 2025-06-03 Xueyuan Chen , Dongchao Yang , Wenxuan Wu , Minglin Wu , Jing Xu , Xixin Wu , Zhiyong Wu , Helen Meng
‹ 上一页 1 8 9 10 下一页 ›