English
Related papers

Related papers: LuSeeL: Language-queried Binaural Universal Sound …

200 papers

Human auditory perception is shaped by moving sound sources in 3D space, yet prior work in generative sound modelling has largely been restricted to mono signals or static spatial audio. In this work, we introduce a framework for generating…

Sound · Computer Science 2025-09-29 Yunyi Liu , Shaofan Yang , Kai Li , Xu Li

In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Helin Wang , Jiarui Hai , Yen-Ju Lu , Karan Thakkar , Mounya Elhilali , Najim Dehak

Wireless distributed systems as used in sensor networks, Internet-of-Things and cyber-physical systems, impose high requirements on resource efficiency. Advanced preprocessing and classification of data at the network edge can help to…

Computer Vision and Pattern Recognition · Computer Science 2018-08-17 Matthias Meyer , Lukas Cavigelli , Lothar Thiele

Sound event localization and detection (SELD) systems estimate both the direction-of-arrival (DOA) and class of sound sources over time. In the DCASE 2022 SELD Challenge (Task 3), models are designed to operate in a 4-channel setting. While…

In this paper, we propose a convolutional recurrent neural network for joint sound event localization and detection (SELD) of multiple overlapping sound events in three-dimensional (3D) space. The proposed network takes a sequence of…

Sound · Computer Science 2018-12-18 Sharath Adavanne , Archontis Politis , Joonas Nikunen , Tuomas Virtanen

We propose DeepASA, a multi-purpose model for auditory scene analysis that performs multi-input multi-output (MIMO) source separation, dereverberation, sound event detection (SED), audio classification, and direction-of-arrival estimation…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-16 Dongheon Lee , Younghoo Kwon , Jung-Woo Choi

This report presents the dataset and the evaluation setup of the Sound Event Localization & Detection (SELD) task for the DCASE 2020 Challenge. The SELD task refers to the problem of trying to simultaneously classify a known set of sound…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-30 Archontis Politis , Sharath Adavanne , Tuomas Virtanen

To date a number of studies have shown that receptive field shapes of early sensory neurons can be reproduced by optimizing coding efficiency of natural stimulus ensembles. A still unresolved question is whether the efficient coding…

Neurons and Cognition · Quantitative Biology 2014-03-18 Wiktor Mlynarski

In recent years, user-generated audio content has proliferated across various media platforms, creating a growing need for efficient retrieval methods that allow users to search for audio clips using natural language queries. This task,…

Sound · Computer Science 2024-12-31 Haoran Sun , Zimu Wang , Qiuyi Chen , Jianjun Chen , Jia Wang , Haiyang Zhang

In this paper, we investigate a novel approach for Target Speech Extraction (TSE), which relies solely on textual context to extract the target speech. We refer to this task as Contextual Speech Extraction (CSE). Unlike traditional TSE…

Sound · Computer Science 2025-03-13 Minsu Kim , Rodrigo Mira , Honglie Chen , Stavros Petridis , Maja Pantic

Separating an audio scene into isolated sources is a fundamental problem in computer audition, analogous to image segmentation in visual scene analysis. Source separation systems based on deep learning are currently the most successful…

Sound · Computer Science 2018-11-07 Prem Seetharaman , Gordon Wichern , Jonathan Le Roux , Bryan Pardo

We propose a novel mixture of experts framework for field-of-view enhancement in binaural signal matching. Our approach enables dynamic spatial audio rendering that adapts to continuous talker motion, allowing users to emphasize or suppress…

Most existing sound event detection~(SED) algorithms operate under a closed-set assumption, restricting their detection capabilities to predefined classes. While recent efforts have explored language-driven zero-shot SED by exploiting…

Sound · Computer Science 2025-10-28 Pengfei Cai , Yan Song , Qing Gu , Nan Jiang , Haoyu Song , Ian McLoughlin

We propose listen to extract (LExt), a highly-effective while extremely-simple algorithm for monaural target speaker extraction (TSE). Given an enrollment utterance of a target speaker, LExt aims at extracting the target speaker from the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-06 Pengjie Shen , Kangrui Chen , Shulin He , Pengru Chen , Shuqi Yuan , He Kong , Xueliang Zhang , Zhong-Qiu Wang

Binaural audio provides a listener with 3D sound sensation, allowing a rich perceptual experience of the scene. However, binaural recordings are scarcely available and require nontrivial expertise and equipment to obtain. We propose to…

Computer Vision and Pattern Recognition · Computer Science 2019-04-10 Ruohan Gao , Kristen Grauman

In this paper we describe a speaker diarization system that enables localization and identification of all speakers present in a conversation or meeting. We propose a novel systematic approach to tackle several long-standing challenges in…

Sound · Computer Science 2021-07-21 Siqi Zheng , Weilong Huang , Xianliang Wang , Hongbin Suo , Jinwei Feng , Zhijie Yan

Estimating the positions of multiple speakers can be helpful for tasks like automatic speech recognition or speaker diarization. Both applications benefit from a known speaker position when, for instance, applying beamforming or assigning…

The current paradigm for creating and deploying immersive audio content is based on audio objects, which are composed of an audio track and position metadata. While rendering an object-based production into a multichannel mix is…

Sound · Computer Science 2021-12-22 Daniel Arteaga , Jordi Pons

Binaural speech separation in real-world scenarios often involves moving speakers. Most current speech separation methods use utterance-level permutation invariant training (u-PIT) for training. In inference time, however, the order of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-15 Cong Han , Nima Mesgarani

Strong representations of target speakers can help extract important information about speakers and detect corresponding temporal regions in multi-speaker conversations. In this study, we propose a neural architecture that simultaneously…

Sound · Computer Science 2023-06-07 Chin-Yi Cheng , Hung-Shin Lee , Yu Tsao , Hsin-Min Wang
‹ Prev 1 8 9 10 Next ›