English
Related papers

Related papers: Towards duration robust weakly supervised sound ev…

200 papers

Sound Event Early Detection (SEED) is an essential task in recognizing the acoustic environments and soundscapes. However, most of the existing methods focus on the offline sound event detection, which suffers from the over-confidence issue…

Sound · Computer Science 2022-02-15 Xujiang Zhao , Xuchao Zhang , Wei Cheng , Wenchao Yu , Yuncong Chen , Haifeng Chen , Feng Chen

Joint sound event localization and detection (SELD) is an emerging audio signal processing task adding spatial dimensions to acoustic scene analysis and sound event detection. A popular approach to modeling SELD jointly is using…

Sound · Computer Science 2021-09-28 Parthasaarathy Sudarsanam , Archontis Politis , Konstantinos Drossos

Acoustic Scene Classification (ASC) and Sound Event Detection (SED) are two separate tasks in the field of computational sound scene analysis. In this work, we present a new dataset with both sound scene and sound event labels and use this…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-02 Helen L. Bear , Ines Nolasco , Emmanouil Benetos

Sound event detection is the task of recognizing sounds and determining their extent (onset/offset times) within an audio clip. Existing systems commonly predict sound presence confidence in short time frames. Then, thresholding produces…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-07 Janek Ebbers , Francois G. Germain , Gordon Wichern , Jonathan Le Roux

Audio tagging aims to predict one or several labels in an audio clip. Many previous works use weakly labelled data (WLD) for audio tagging, where only presence or absence of sound events is known, but the order of sound events is unknown.…

Sound · Computer Science 2018-08-07 Yuanbo Hou , Qiuqiang Kong , Shengchen Li

Audio Event Detection is an important task for content analysis of multimedia data. Most of the current works on detection of audio events is driven through supervised learning approaches. We propose a weakly supervised learning framework…

Sound · Computer Science 2016-06-14 Anurag Kumar , Bhiksha Raj

The Audio-Visual Video Parsing task aims to identify and temporally localize the events that occur in either or both the audio and visual streams of audible videos. It often performs in a weakly-supervised manner, where only video event…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Jinxing Zhou , Dan Guo , Yiran Zhong , Meng Wang

In this paper, we propose addressing the lack of strongly labeled data by using pseudo strongly labeled data approximated using Convolutive Nonnegative Matrix Factorization. Using this set of data, we then train a novel architecture called…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-03 Teck Kai Chan , Cheng Siong Chin

This study presents an audio-visual information fusion approach to sound event localization and detection (SELD) in low-resource scenarios. We aim at utilizing audio and video modality information through cross-modal learning and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-24 Ya Jiang , Qing Wang , Jun Du , Maocheng Hu , Pengfei Hu , Zeyan Liu , Shi Cheng , Zhaoxu Nian , Yuxuan Dong , Mingqi Cai , Xin Fang , Chin-Hui Lee

Sound event detection (SED) methods are tasked with labeling segments of audio recordings by the presence of active sound sources. SED is typically posed as a supervised machine learning problem, requiring strong annotations for the…

Sound · Computer Science 2018-08-13 Brian McFee , Justin Salamon , Juan Pablo Bello

In this paper, we study the use of soft labels to train a system for sound event detection (SED). Soft labels can result from annotations which account for human uncertainty about categories, or emerge as a natural representation of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-01 Irene Martín-Morató , Manu Harju , Paul Ahokas , Annamaria Mesaros

Sound event detection (SED) often suffers from the data deficiency problem. The recent baseline system in the DCASE2023 challenge task 4 leverages the large pretrained self-supervised learning (SelfSL) models to mitigate such restriction,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-01 Nian Shao , Xian Li , Xiaofei Li

Sound event localization frameworks based on deep neural networks have shown increased robustness with respect to reverberation and noise in comparison to classical parametric approaches. In particular, recurrent architectures that…

We target the problem of developing new low-complexity networks for the sound event detection task. Our goal is to meticulously analyze the performance-complexity trade-off, aiming to be competitive with the large state-of-the-art models,…

Sound · Computer Science 2025-06-13 Tobias Morocutti , Florian Schmid , Jonathan Greif , Francesco Foscarin , Gerhard Widmer

Sound event detection is a core module for acoustic environmental analysis. Semi-supervised learning technique allows to largely scale up the dataset without increasing the annotation budget, and recently attracts lots of research…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-02 Xiaofei Li

Self-supervised learning (SSL) models offer powerful representations for sound event detection (SED), yet their synergistic potential remains underexplored. This study systematically evaluates state-of-the-art SSL models to guide optimal…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-27 Hanfang Cui , Longfei Song , Li Li , Dongxing Xu , Yanhua Long

Localizing sounds and detecting events in different room environments is a difficult task, mainly due to the wide range of reflections and reverberations. When training neural network models with sounds recorded in only a few room…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-06 Yusun Shul , Byeong-Yun Ko , Jung-Woo Choi

Audio event classification is an important task for several applications such as surveillance, audio, video and multimedia retrieval etc. There are approximately 3M people with hearing loss who can't perceive events happening around them.…

Sound · Computer Science 2020-01-14 Harishchandra Dubey , Dimitra Emmanouilidou , Ivan J. Tashev

Sound event localization and detection (SELD) is a task for the classification of sound events and the identification of direction of arrival (DoA) utilizing multichannel acoustic signals. For effective classification and localization, a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-18 Yusun Shul , Dayun Choi , Jung-Woo Choi

Existing systems for sound event localization and detection (SELD) typically operate by estimating a source location for all classes at every time instant. In this paper, we propose an alternative class-conditioned SELD model for situations…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-09 Olga Slizovskaia , Gordon Wichern , Zhong-Qiu Wang , Jonathan Le Roux