English
Related papers

Related papers: A Computationally Efficient and Practically Feasib…

200 papers

Speech separation involves extracting an individual speaker's voice from a multi-speaker audio signal. The increasing complexity of real-world environments, where multiple speakers might converse simultaneously, underscores the importance…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Renana Opochinsky , Mordehay Moradi , Sharon Gannot

Separating a song into vocal and accompaniment components is an active research topic, and recent years witnessed an increased performance from supervised training using deep learning techniques. We propose to apply the visual information…

Sound · Computer Science 2021-07-02 Bochen Li , Yuxuan Wang , Zhiyao Duan

We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with simulated room impulse responses (RIRs) using omni-directional…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-18 Yang Yang , George Sung , Shao-Fu Shih , Hakan Erdogan , Chehung Lee , Matthias Grundmann

We describe a joint acoustic echo cancellation (AEC) and blind source extraction (BSE) approach for multi-microphone acoustic frontends. The proposed algorithm blindly estimates AEC and beamforming filters by maximizing the statistical…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-11 Thomas Haubner , Zbyněk Koldovský , Walter Kellermann

This paper addresses the challenge of joint communication and sensing (JCAS) in next-generation wireless networks, with an emphasis on in-band full-duplex (IBFD) multiple-input multiple-output (MIMO) systems. Traditionally,…

Emerging Technologies · Computer Science 2025-08-29 Siyao Li , Conrad Prisby , Thomas Yang

Speech separation approaches for single-channel, dry speech mixtures have significantly improved. However, real-world spatial and reverberant acoustic environments remain challenging, limiting the effectiveness of these approaches for…

Spoofing attacks posed by generating artificial speech can severely degrade the performance of a speaker verification system. Recently, many anti-spoofing countermeasures have been proposed for detecting varying types of attacks from…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-08 Yuanjun Zhao , Roberto Togneri , Victor Sreeram

Augmented listening devices such as hearing aids often perform poorly in noisy and reverberant environments with many competing sound sources. Large distributed microphone arrays can improve performance, but data from remote microphones…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-12 Ryan M. Corey , Matthew D. Skarha , Andrew C. Singer

Blind-audio-source-separation (BASS) techniques, particularly those with low latency, play an important role in a wide range of real-time systems, e.g., hearing aids, in-car hand-free voice communication, real-time human-machine…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-17 Kaien Mo , Xianrui Wang , Yichen Yang , Shoji Makino , Jingdong Chen

In this paper we propose a method for separation of moving sound sources. The method is based on first tracking the sources and then estimation of source spectrograms using multichannel non-negative matrix factorization (NMF) and extracting…

Sound · Computer Science 2017-10-30 Joonas Nikunen , Aleksandr Diment , Tuomas Virtanen

The increasing popularity of spatial audio in applications such as teleconferencing, entertainment, and virtual reality has led to the recent developments of binaural reproduction methods. However, only a few of these methods are…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-17 Ami Berger , Vladimir Tourbabin , Jacob Donley , Zamir Ben-Hur , Boaz Rafaely

The expansion of telecommunications incurs increasingly severe crosstalk and interference, and a physical layer cognitive method, called blind source separation (BSS), can effectively address these issues. BSS requires minimal prior…

Signal Processing · Electrical Eng. & Systems 2023-03-22 Weipeng Zhang , Alexander Tait , Chaoran Huang , Thomas Ferreira de Lima , Simon Bilodeau , Eric Blow , Aashu Jha , Bhavin J. Shastri , Paul Prucnal

In this paper, we present a statistical beamforming algorithm as a pre-processing step for robust automatic speech recognition (ASR). By modeling the target speech as a non-stationary Laplacian distribution, a mask-based statistical…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-08 Ui-Hyeop Shin , Hyung-Min Park

Considering a mixed signal composed of various audio sources and recorded with a single microphone, we consider on this paper the blind audio source separation problem which consists in isolating and extracting each of the sources. To…

Signal Processing · Electrical Eng. & Systems 2020-07-15 Valentin Leplat , Nicolas Gillis , Man Shun Ang

The extraction of a desired speech signal from a noisy environment has become a challenging issue. In the recent years, the scientific community has particularly focused on multichannel techniques which are dealt with in this review. In…

Sound · Computer Science 2013-01-01 Adel Hidri , Souad Meddeb , Hamid Amiri

This paper presents a neural method for distant speech recognition (DSR) that jointly separates and diarizes speech mixtures without supervision by isolated signals. A standard separation method for multi-talker DSR is a statistical…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-13 Yoshiaki Bando , Tomohiko Nakamura , Shinji Watanabe

Multichannel audio blind source separation (BSS) in the determined situation (the number of microphones is equal to that of the sources), or determined BSS, is performed by multichannel linear filtering in the time-frequency domain to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-15 Kohei Yatabe

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Hyeonggon Ryu , Seongyu Kim , Joon Son Chung , Arda Senocak

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is suffering from the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Qiushi Zhu , Jie Zhang , Yu Gu , Yuchen Hu , Lirong Dai

Under noisy conditions, automatic speech recognition (ASR) can greatly benefit from the addition of visual signals coming from a video of the speaker's face. However, when multiple candidate speakers are visible this traditionally requires…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-12 Otavio Braga , Olivier Siohan