中文
相关论文

相关论文: Moving Speaker Separation via Parallel Spectral-Sp…

200 篇论文

Speech translation (ST) automatically converts utterances in a source language into text in another language. Splitting continuous speech into shorter segments, known as speech segmentation, plays an important role in ST. Recent…

音频与语音处理 · 电气工程与系统科学 2023-12-19 Ryo Fukuda , Katsuhito Sudoh , Satoshi Nakamura

The spatial semantic segmentation task focuses on separating and classifying sound objects from multichannel signals. To achieve two different goals, conventional methods fine-tune a large classification model cascaded with the separation…

音频与语音处理 · 电气工程与系统科学 2025-09-22 Younghoo Kwon , Jung-Woo Choi

Dual-path processing along the temporal and spectral dimensions has shown to be effective in various speech processing applications. While the sound source localization (SSL) models utilizing dual-path processing such as the FN-SSL and…

音频与语音处理 · 电气工程与系统科学 2025-09-24 Yuseon Choi , Hyeonseung Kim , Jewoo Jun , Jong Won Shin

Signal separation in the passive underwater acoustic domain has heavily relied on deep learning techniques to isolate ship radiated noise. However, the separation networks commonly used in this domain stem from speech separation…

声音 · 计算机科学 2025-04-14 Yucheng Liu , Longyu Jiang

Deep learning has the potential to enhance speech signals and increase their intelligibility for users of hearing aids. Deep models suited for real-world application should feature a low computational complexity and low processing delay of…

音频与语音处理 · 电气工程与系统科学 2024-10-31 Nils L. Westhausen , Hendrik Kayser , Theresa Jansen , Bernd T. Meyer

Localizing sounds and detecting events in different room environments is a difficult task, mainly due to the wide range of reflections and reverberations. When training neural network models with sounds recorded in only a few room…

音频与语音处理 · 电气工程与系统科学 2023-06-06 Yusun Shul , Byeong-Yun Ko , Jung-Woo Choi

We tackle the multi-party speech recovery problem through modeling the acoustic of the reverberant chambers. Our approach exploits structured sparsity models to perform room modeling and speech recovery. We propose a scheme for…

机器学习 · 计算机科学 2012-10-26 Afsaneh Asaei , Mohammad Golbabaee , Hervé Bourlard , Volkan Cevher

The end-to-end approach for single-channel speech separation has been studied recently and shown promising results. This paper extended the previous approach and proposed a new end-to-end model for multi-channel speech separation. The…

声音 · 计算机科学 2019-05-29 Rongzhi Gu , Jian Wu , Shi-Xiong Zhang , Lianwu Chen , Yong Xu , Meng Yu , Dan Su , Yuexian Zou , Dong Yu

In this paper we address the problems of modeling the acoustic space generated by a full-spectrum sound source and of using the learned model for the localization and separation of multiple sources that simultaneously emit sparse-spectrum…

声音 · 计算机科学 2015-02-06 Antoine Deleforge , Florence Forbes , Radu Horaud

To estimate the direction of arrival (DOA) of multiple speakers with methods that use prototype transfer functions, frequency-dependent spatial spectra (SPS) are usually constructed. To make the DOA estimation robust, SPS from different…

音频与语音处理 · 电气工程与系统科学 2026-02-11 Daniel Fejgin , Elior Hadad , Sharon Gannot , Zbyněk Koldovský , Simon Doclo

Disentanglement of a speaker's timbre and style is very important for style transfer in multi-speaker multi-style text-to-speech (TTS) scenarios. With the disentanglement of timbres and styles, TTS systems could synthesize expressive speech…

声音 · 计算机科学 2022-11-23 Wei Song , Yanghao Yue , Ya-jie Zhang , Zhengchen Zhang , Youzheng Wu , Xiaodong He

While modern best practices advocate for scalable architectures that support long-range interactions, object-centric models are yet to fully embrace these architectures. In particular, existing object-centric models for handling sequential…

机器学习 · 计算机科学 2024-02-28 Gautam Singh , Yue Wang , Jiawei Yang , Boris Ivanovic , Sungjin Ahn , Marco Pavone , Tong Che

Spoken dialogue is an intuitive form of human-computer interaction, yet current speech language models often remain constrained to turn-based exchanges, lacking real-time adaptability such as user barge-in. We propose a novel duplex speech…

We present a multi-channel database of overlapping speech for training, evaluation, and detailed analysis of source separation and extraction algorithms: SMS-WSJ -- Spatialized Multi-Speaker Wall Street Journal. It consists of artificially…

声音 · 计算机科学 2019-10-31 Lukas Drude , Jens Heitkaemper , Christoph Boeddeker , Reinhold Haeb-Umbach

Teleconferencing is becoming essential during the COVID-19 pandemic. However, in real-world applications, speech quality can deteriorate due to, for example, background interference, noise, or reverberation. To solve this problem, target…

音频与语音处理 · 电气工程与系统科学 2022-05-02 Yicheng Hsu , Yonghan Lee , Mingsian R. Bai

Recently, multi-channel speech enhancement has drawn much interest due to the use of spatial information to distinguish target speech from interfering signal. To make full use of spatial information and neural network based masking…

音频与语音处理 · 电气工程与系统科学 2022-10-18 Shubo Lv , Yihui Fu , Yukai Jv , Lei Xie , Weixin Zhu , Wei Rao , Yannan Wang

In recent years there have been many deep learning approaches towards the multi-speaker source separation problem. Most use Long Short-Term Memory - Recurrent Neural Networks (LSTM-RNN) or Convolutional Neural Networks (CNN) to model the…

机器学习 · 计算机科学 2019-12-20 Jeroen Zegers , Hugo Van hamme

The performance of automatic speech recognition (ASR) systems severely degrades when multi-talker speech overlap occurs. In meeting environments, speech separation is typically performed to improve the robustness of ASR systems. Recently,…

音频与语音处理 · 电气工程与系统科学 2023-01-18 Hassan Taherian , DeLiang Wang

Acoustic-to-articulatory inversion (AAI) aims to estimate the parameters of articulators from speech audio. There are two common challenges in AAI, which are the limited data and the unsatisfactory performance in speaker independent…

声音 · 计算机科学 2023-02-28 Jianrong Wang , Jinyu Liu , Li Liu , Xuewei Li , Mei Yu , Jie Gao , Qiang Fang

In this paper we propose a method for separation of moving sound sources. The method is based on first tracking the sources and then estimation of source spectrograms using multichannel non-negative matrix factorization (NMF) and extracting…

声音 · 计算机科学 2017-10-30 Joonas Nikunen , Aleksandr Diment , Tuomas Virtanen