中文
相关论文

相关论文: Time-domain Speech Enhancement Assisted by Multi-r…

200 篇论文

Recently, phase processing is attracting increasinginterest in speech enhancement community. Some researchersintegrate phase estimations module into speech enhancementmodels by using complex-valued short-time Fourier transform(STFT)…

声音 · 计算机科学 2019-01-03 Xingjian Du , Mengyao Zhu , Xuan Shi , Xinpeng Zhang , Wen Zhang , Jingdong Chen

Speech enhancement (SE) enables robust speech recognition, real-time communication, hearing aids, and other applications where speech quality is crucial. However, deploying such systems on resource-constrained devices involves choosing a…

音频与语音处理 · 电气工程与系统科学 2025-12-02 Riccardo Miccini , Minje Kim , Clément Laroche , Luca Pezzarossa , Paris Smaragdis

Prompting and adapter tuning have emerged as efficient alternatives to fine-tuning (FT) methods. However, existing studies on speech prompting focused on classification tasks and failed on more complex sequence generation tasks. Besides,…

音频与语音处理 · 电气工程与系统科学 2023-11-16 Kai-Wei Chang , Ming-Hsin Chen , Yun-Ping Lin , Jing Neng Hsu , Paul Kuo-Ming Huang , Chien-yu Huang , Shang-Wen Li , Hung-yi Lee

Single-channel speech separation in time domain and frequency domain has been widely studied for voice-driven applications over the past few years. Most of previous works assume known number of speakers in advance, however, which is not…

音频与语音处理 · 电气工程与系统科学 2020-04-02 Yiming Xiao , Haijian Zhang

Speech separation has recently made significant progress thanks to the fine-grained vision used in time-domain methods. However, several studies have shown that adopting Short-Time Fourier Transform (STFT) for feature extraction could be…

声音 · 计算机科学 2024-03-05 Kuan-Hsun Ho , Jeih-weih Hung , Berlin Chen

The multi-codebook speech codec enables the application of large language models (LLM) in TTS but bottlenecks efficiency and robustness due to multi-sequence prediction. To avoid this obstacle, we propose Single-Codec, a single-codebook…

音频与语音处理 · 电气工程与系统科学 2024-06-12 Hanzhao Li , Liumeng Xue , Haohan Guo , Xinfa Zhu , Yuanjun Lv , Lei Xie , Yunlin Chen , Hao Yin , Zhifei Li

Speech enhancement (SE) aims to improve the clarity, intelligibility, and quality of speech signals for various speech enabled applications. However, air-conducted (AC) speech is highly susceptible to ambient noise, particularly in low…

音频与语音处理 · 电气工程与系统科学 2025-01-14 Fuyuan Feng , Longting Xu , Rohan Kumar Das

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of…

声音 · 计算机科学 2019-04-03 Hyeong-Seok Choi , Jang-Hyun Kim , Jaesung Huh , Adrian Kim , Jung-Woo Ha , Kyogu Lee

A typical neural speech enhancement (SE) approach mainly handles speech and noise mixtures, which is not optimal for singing voice enhancement scenarios. Music source separation (MSS) models treat vocals and various accompaniment components…

声音 · 计算机科学 2023-10-09 Weiming Xu , Zhouxuan Chen , Zhili Tan , Shubo Lv , Runduo Han , Wenjiang Zhou , Weifeng Zhao , Lei Xie

From hearing aids to augmented and virtual reality devices, binaural speech enhancement algorithms have been established as state-of-the-art techniques to improve speech intelligibility and listening comfort. In this paper, we present an…

音频与语音处理 · 电气工程与系统科学 2025-07-29 Vikas Tokala , Eric Grinstein , Mike Brookes , Simon Doclo , Jesper Jensen , Patrick A. Naylor

In audio processing applications, phase retrieval (PR) is often performed from the magnitude of short-time Fourier transform (STFT) coefficients. Although PR performance has been observed to depend on the considered STFT parameters and…

信号处理 · 电气工程与系统科学 2021-06-10 Andrés Marafioti , Nicki Holighaus , Piotr Majdak

Deep learning-based models have greatly advanced the performance of speech enhancement (SE) systems. However, two problems remain unsolved, which are closely related to model generalizability to noisy conditions: (1) mismatched noisy…

音频与语音处理 · 电气工程与系统科学 2020-12-29 Cheng Yu , Ryandhimas E. Zezario , Syu-Siang Wang , Jonathan Sherman , Yi-Yen Hsieh , Xugang Lu , Hsin-Min Wang , Yu Tsao

In this work, we explore the constant-Q transform (CQT) for speech emotion recognition (SER). The CQT-based time-frequency analysis provides variable spectro-temporal resolution with higher frequency resolution at lower frequencies. Since…

音频与语音处理 · 电气工程与系统科学 2021-02-09 Premjeet Singh , Goutam Saha , Md Sahidullah

Tacotron-based text-to-speech (TTS) systems directly synthesize speech from text input. Such frameworks typically consist of a feature prediction network that maps character sequences to frequency-domain acoustic features, followed by a…

音频与语音处理 · 电气工程与系统科学 2020-04-08 Rui Liu , Berrak Sisman , Feilong Bao , Guanglai Gao , Haizhou Li

Speech translation (ST) automatically converts utterances in a source language into text in another language. Splitting continuous speech into shorter segments, known as speech segmentation, plays an important role in ST. Recent…

音频与语音处理 · 电气工程与系统科学 2023-12-19 Ryo Fukuda , Katsuhito Sudoh , Satoshi Nakamura

Self-supervised learning has demonstrated impressive performance in speech tasks, yet there remains ample opportunity for advancement in the realm of speech enhancement research. In addressing speech tasks, confining the attention mechanism…

音频与语音处理 · 电气工程与系统科学 2024-08-14 Tao Zheng , Liejun Wang , Yinfeng Yu

Speech translation (ST) aims to learn transformations from speech in the source language to the text in the target language. Previous works show that multitask learning improves the ST performance, in which the recognition decoder generates…

计算与语言 · 计算机科学 2020-07-07 Shun-Po Chuang , Tzu-Wei Sung , Alexander H. Liu , Hung-yi Lee

The data analysis of space-based gravitational wave detectors like Taiji faces significant challenges from non-stationary noise, which compromises the efficacy of traditional frequency-domain analysis. This work proposes a unified framework…

广义相对论与量子宇宙学 · 物理学 2025-06-23 Minghui Du , Ziren Luo , Peng Xu

In this work, we address EmoFake Detection (EFD). We hypothesize that multilingual speech foundation models (SFMs) will be particularly effective for EFD due to their pre-training across diverse languages, enabling a nuanced understanding…

音频与语音处理 · 电气工程与系统科学 2025-07-18 Orchid Chetia Phukan , Mohd Mujtaba Akhtar , Girish , Arun Balaji Buduru

Neural speech codecs have demonstrated their ability to compress high-quality speech and audio by converting them into discrete token representations. Most existing methods utilize Residual Vector Quantization (RVQ) to encode speech into…

声音 · 计算机科学 2024-10-22 Peiji Yang , Fengping Wang , Yicheng Zhong , Huawei Wei , Zhisheng Wang