中文
相关论文

相关论文: Real-time Stereo Speech Enhancement with Spatial-C…

200 篇论文

Imagine being in a crowded space where people speak a different language and having hearables that transform the auditory space into your native language, while preserving the spatial cues for all speakers. We introduce spatial speech…

计算与语言 · 计算机科学 2025-04-29 Tuochao Chen , Qirui Wang , Runlin He , Shyam Gollakota

Recently, deep representation learning has shown strong performance in multiple audio tasks. However, its use for learning spatial representations from multichannel audio is underexplored. We investigate the use of a pretraining stage based…

In this paper, we present RT-GCC-NMF: a real-time (RT), two-channel blind speech enhancement algorithm that combines the non-negative matrix factorization (NMF) dictionary learning algorithm with the generalized cross-correlation (GCC)…

音频与语音处理 · 电气工程与系统科学 2019-04-08 Sean U. N. Wood , Jean Rouat

This paper describes the practical response- and performance-aware development of online speech enhancement for an augmented reality (AR) headset that helps a user understand conversations made in real noisy echoic environments (e.g.,…

音频与语音处理 · 电气工程与系统科学 2022-07-18 Kouhei Sekiguchi , Aditya Arie Nugraha , Yicheng Du , Yoshiaki Bando , Mathieu Fontaine , Kazuyoshi Yoshii

This paper addresses the problem of single-channel speech separation, where the number of speakers is unknown, and each speaker may speak multiple utterances. We propose a speech separation model that simultaneously performs separation,…

音频与语音处理 · 电气工程与系统科学 2025-05-23 Yuzhu Wang , Archontis Politis , Konstantinos Drossos , Tuomas Virtanen

Recently, multi-channel speech enhancement has drawn much interest due to the use of spatial information to distinguish target speech from interfering signal. To make full use of spatial information and neural network based masking…

音频与语音处理 · 电气工程与系统科学 2022-10-18 Shubo Lv , Yihui Fu , Yukai Jv , Lei Xie , Weixin Zhu , Wei Rao , Yannan Wang

We introduce and analyze a novel approach to the problem of speaker identification in multi-party recorded meetings. Given a speech segment and a set of available candidate profiles, we propose a novel data-driven way to model the distance…

音频与语音处理 · 电气工程与系统科学 2021-02-23 Nikolaos Flemotomos , Dimitrios Dimitriadis

Existing multi-channel continuous speech separation (CSS) models are heavily dependent on supervised data - either simulated data which causes data mismatch between the training and real-data testing, or the real transcribed overlapping…

音频与语音处理 · 电气工程与系统科学 2022-04-08 Xiaofei Wang , Dongmei Wang , Naoyuki Kanda , Sefik Emre Eskimez , Takuya Yoshioka

Modern neural network-based speech processing systems usually need to have reverberation resistance, so the training of such systems requires a large amount of reverberation data. In the process of system training, it is now more inclined…

音频与语音处理 · 电气工程与系统科学 2025-08-19 Dong Yang

For real-time speech enhancement (SE) including noise suppression, dereverberation and acoustic echo cancellation, the time-variance of the audio signals becomes a severe challenge. The causality and memory usage limit that only the…

音频与语音处理 · 电气工程与系统科学 2023-02-22 Chengyu Zheng , Yuan Zhou , Xiulian Peng , Yuan Zhang , Yan Lu

Replay speech attacks pose a significant threat to voice-controlled systems, especially in smart environments where voice assistants are widely deployed. While multi-channel audio offers spatial cues that can enhance replay detection…

音频与语音处理 · 电气工程与系统科学 2025-09-19 Michael Neri , Tuomas Virtanen

Personalized speech enhancement (PSE) is a real-time SE approach utilizing a speaker embedding of a target person to remove background noise, reverberation, and interfering voices. To deploy a PSE model for full duplex communications, the…

音频与语音处理 · 电气工程与系统科学 2023-05-29 Sefik Emre Eskimez , Takuya Yoshioka , Alex Ju , Min Tang , Tanel Parnamaa , Huaming Wang

In multimedia applications such as films and video games, spatial audio techniques are widely employed to enhance user experiences by simulating 3D sound: transforming mono audio into binaural formats. However, this process is often complex…

多媒体 · 计算机科学 2025-02-14 Xiaojing Liu , Ogulcan Gurelli , Yan Wang , Joshua Reiss

Recent works have shown that Deep Recurrent Neural Networks using the LSTM architecture can achieve strong single-channel speech enhancement by estimating time-frequency masks. However, these models do not naturally generalize to…

声音 · 计算机科学 2020-12-04 Felix Grezes , Zhaoheng Ni , Viet Anh Trinh , Michael Mandel

Personalized speech enhancement has been a field of active research for suppression of speechlike interferers such as competing speakers or TV dialogues. Compared with single channel approaches, multichannel PSE systems can be more…

音频与语音处理 · 电气工程与系统科学 2022-11-17 Yicheng Hsu , Yonghan Lee , Mingsian R. Bai

It is commonly believed that multipath hurts various audio processing algorithms. At odds with this belief, we show that multipath in fact helps sound source separation, even with very simple propagation models. Unlike most existing…

声音 · 计算机科学 2019-05-08 Robin Scheibler , Diego Di Carlo , Antoine Deleforge , Ivan Dokmanić

Multi-channel multi-talker speech recognition presents formidable challenges in the realm of speech processing, marked by issues such as background noise, reverberation, and overlapping speech. Overcoming these complexities requires…

音频与语音处理 · 电气工程与系统科学 2023-10-09 Yiwen Shao

Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speakers, background noise and room reverberation remains a…

声音 · 计算机科学 2022-04-11 Guinan Li , Jianwei Yu , Jiajun Deng , Xunying Liu , Helen Meng

This paper describes a system that gives a mobile robot the ability to perform automatic speech recognition with simultaneous speakers. A microphone array is used along with a real-time implementation of Geometric Source Separation and a…

Building a single universal speech enhancement (SE) system that can handle arbitrary input is a demanded but underexplored research topic. Towards this ultimate goal, one direction is to build a single model that handles diverse audio…

音频与语音处理 · 电气工程与系统科学 2024-02-19 Wangyou Zhang , Jee-weon Jung , Shinji Watanabe , Yanmin Qian
‹ 上一页 1 8 9 10 下一页 ›