English
Related papers

Related papers: An Experimental Study on Joint Modeling for Sound …

200 papers

Sound event detection (SED) methods are tasked with labeling segments of audio recordings by the presence of active sound sources. SED is typically posed as a supervised machine learning problem, requiring strong annotations for the…

Sound · Computer Science 2018-08-13 Brian McFee , Justin Salamon , Juan Pablo Bello

In this paper we present our work on Task 1 Acoustic Scene Classi- fication and Task 3 Sound Event Detection in Real Life Recordings. Among our experiments we have low-level and high-level features, classifier optimization and other…

This paper focuses on few-shot Sound Event Detection (SED), which aims to automatically recognize and classify sound events with limited samples. However, prevailing methods methods in few-shot SED predominantly rely on segment-level…

Sound · Computer Science 2024-03-20 Liang Zou , Genwei Yan , Ruoyu Wang , Jun Du , Meng Lei , Tian Gao , Xin Fang

In conventional sound event detection (SED) models, two types of events, namely, those that are present and those that do not occur in an acoustic scene, are regarded as the same type of events. The conventional SED methods cannot…

Sound · Computer Science 2021-02-11 Noriyuki Tonami , Keisuke Imoto , Yuki Okamoto , Takahiro Fukumori , Yoichi Yamashita

Polyphonic sound event localization and detection (SELD) has many practical applications in acoustic sensing and monitoring. However, the development of real-time SELD has been limited by the demanding computational requirement of most…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-07 Thi Ngoc Tho Nguyen , Douglas L. Jones , Karn N. Watcharasupat , Huy Phan , Woon-Seng Gan

A popular approach for 3D source localization using multiple microphones is the steered-response power method, where the source position is directly estimated by maximizing a function of three continuous position variables. Instead of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-19 Klaus Brümann , Simon Doclo

The Dynamic Saliency Prediction (DSP) task simulates the human selective attention mechanism to perceive the dynamic scene, which is significant and imperative in many vision tasks. Most of existing methods only consider visual cues, while…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Hailong Ning , Bin Zhao , Zhanxuan Hu , Lang He , Ercheng Pei

Sound event detection (SED) often suffers from the data deficiency problem. The recent baseline system in the DCASE2023 challenge task 4 leverages the large pretrained self-supervised learning (SelfSL) models to mitigate such restriction,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-01 Nian Shao , Xian Li , Xiaofei Li

Sound Source Localization (SSL) involves estimating the Direction of Arrival (DOA) of sound sources. Since the DOA estimation output space is continuous, regression might be more suitable for DOA, offering higher precision. However, in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-30 Linfeng Feng , Xiao-Lei Zhang , Xuelong Li

There are two sub-tasks implied in the weakly-supervised SED: audio tagging and event boundary detection. Current methods which combine multi-task learning with SED requires annotations both for these two sub-tasks. Since there are only…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-25 Yuxin Huang , Xiangdong Wang , Liwei Lin , Hong Liu , Yueliang Qian

In this paper, we propose a temporal-frequential attention model for sound event detection (SED). Our network learns how to listen with two attention models: a temporal attention model and a frequential attention model. Proposed system…

Sound · Computer Science 2025-05-06 Yu-Han Shen , Ke-Xin He , Wei-Qiang Zhang

In this paper, we propose an effective sound event detection (SED) method based on the audio spectrogram transformer (AST) model, pretrained on the large-scale AudioSet for audio tagging (AT) task, termed AST-SED. Pretrained AST models have…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-08 Kang Li , Yan Song , Li-Rong Dai , Ian McLoughlin , Xin Fang , Lin Liu

Existing methods utilizing spatial information for sound source separation require prior knowledge of the direction of arrival (DOA) of the source or utilize estimated but imprecise localization results, which impairs the separation…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-08 Donghang Wu , Xihong Wu , Tianshu Qu

Localizing a moving sound source in the real world involves determining its direction-of-arrival (DOA) and distance relative to a microphone. Advancements in DOA estimation have been facilitated by data-driven methods optimized with large…

Sound · Computer Science 2023-09-19 Saksham Singh Kushwaha , Iran R. Roman , Magdalena Fuentes , Juan Pablo Bello

This paper describes that semi-supervised learning called peer collaborative learning (PCL) can be applied to the polyphonic sound event detection (PSED) task, which is one of the tasks in the Detection and Classification of Acoustic Scenes…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-31 Hayato Endo , Hiromitsu Nishizaki

While many deep learning methods on other domains have been applied to sound event detection (SED), differences between original domains of the methods and SED have not been appropriately considered so far. As SED uses audio data with two…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-24 Hyeonuk Nam , Seong-Hu Kim , Deokki Min , Byeong-Yun Ko , Seung-Deok Choi , Yong-Hwa Park

Pre-training methods have greatly improved the performance of sound event localization and detection (SELD). However, existing Transformer-based models still face high computational cost. To solve this problem, we present a stereo SELD…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-15 Wenmiao Gao , Han Yin

In recent years, deep learning systems have shown a concerning trend toward increased complexity and higher energy consumption. As researchers in this domain and organizers of one of the Detection and Classification of Acoustic Scenes and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-18 Francesca Ronchini , Romain Serizel

Sound event detection is the task of recognizing sounds and determining their extent (onset/offset times) within an audio clip. Existing systems commonly predict sound presence confidence in short time frames. Then, thresholding produces…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-07 Janek Ebbers , Francois G. Germain , Gordon Wichern , Jonathan Le Roux

Sound separation (SS) and target sound extraction (TSE) are fundamental techniques for addressing complex acoustic scenarios. While existing SS methods struggle with determining the unknown number of sound sources, TSE approaches require…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-25 Hongyu Wang , Chenda Li , Xin Zhou , Shuai Wang , Yanmin Qian