English
Related papers

Related papers: SP-SEDT: Self-supervised Pre-training for Sound Ev…

200 papers

Due to the limitation of strong-labeled sound event detection data set, using synthetic data to improve the sound event detection system performance has been a new research focus. In this paper, we try to exploit the usage of synthetic data…

Sound · Computer Science 2020-11-03 Yuxin Huang , Liwei Lin , Xiangdong Wang , Hong Liu , Yueliang Qian , Min Liu , Kazushige Ouchi

This paper presents our work of training acoustic event detection (AED) models using unlabeled dataset. Recent acoustic event detectors are based on large-scale neural networks, which are typically trained with huge amounts of labeled data.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-01 Bowen Shi , Ming Sun , Chieh-Chi Kao , Viktor Rozgic , Spyros Matsoukas , Chao Wang

In recent years, there has been a growing demand for improved autonomy for in-orbit operations such as rendezvous, docking, and proximity maneuvers, leading to increased interest in employing Deep Learning-based Spacecraft Pose Estimation…

Computer Vision and Pattern Recognition · Computer Science 2023-11-10 Arunkumar Rathinam , Haytam Qadadri , Djamila Aouada

In this paper, we propose a novel four-stage data augmentation approach to ResNet-Conformer based acoustic modeling for sound event localization and detection (SELD). First, we explore two spatial augmentation techniques, namely audio…

Sound · Computer Science 2023-03-08 Qing Wang , Jun Du , Hua-Xin Wu , Jia Pan , Feng Ma , Chin-Hui Lee

This technical report outlines our approach to Task 3A of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024, focusing on Sound Event Localization and Detection (SELD). SELD provides valuable insights by estimating…

Sound · Computer Science 2025-07-25 Quoc Thinh Vo , David Han

Sound event localization and detection (SELD) has seen substantial advancements through learning-based methods. These systems, typically trained from scratch on specific datasets, have shown considerable generalization capabilities.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-03 Jinbo Hu , Yin Cao , Ming Wu , Fang Kang , Feiran Yang , Wenwu Wang , Mark D. Plumbley , Jun Yang

This paper proposes an effective modelling of sound event spectra with a hidden data-size-imbalance, for improved Acoustic Event Detection (AED). The proposed method models each event as an aggregated representation of a few latent factors,…

Audio and Speech Processing · Electrical Eng. & Systems 2019-04-08 Chaitanya Narisetty , Tatsuya Komatsu , Reishi Kondo

Sound event localization and detection (SELD) involves predicting active sound event classes over time while estimating their positions. The localization subtask in SELD is usually treated as a direction of arrival estimation problem,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-21 Davide Berghi , Philip J. B. Jackson

In this paper we investigate the importance of the extent of memory in sequential self attention for sound recognition. We propose to use a memory controlled sequential self attention mechanism on top of a convolutional recurrent neural…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Arjun Pankajakshan , Helen L. Bear , Vinod Subramanian , Emmanouil Benetos

Environment shifts and conflicts present significant challenges for learning-based sound event localization and detection (SELD) methods. SELD systems, when trained in particular acoustic settings, often show restricted generalization…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-08 Jinbo Hu , Yin Cao , Ming Wu , Qiuqiang Kong , Feiran Yang , Mark D. Plumbley , Jun Yang

While recent Transformer-based approaches have shown impressive performances on event-based object detection tasks, their high computational costs still diminish the low power consumption advantage of event cameras. Image-based works…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Yansong Peng , Hebei Li , Yueyi Zhang , Xiaoyan Sun , Feng Wu

This technical report details our systems submitted for Task 3 of the DCASE 2024 Challenge: Audio and Audiovisual Sound Event Localization and Detection (SELD) with Source Distance Estimation (SDE). We address only the audio-only SELD with…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-15 Jun Wei Yeow , Ee-Leng Tan , Jisheng Bai , Santi Peksi , Woon-Seng Gan

This paper introduces a zero-shot sound event classification (ZS-SEC) method to identify sound events that have never occurred in training data. In our previous work, we proposed a ZS-SEC method using sound attribute vectors (SAVs), where a…

Sound · Computer Science 2023-03-21 Yi-Han Lin , Xunquan Chen , Ryoichi Takashima , Tetsuya Takiguchi

This work introduces Sample-Efficient Speech Diffusion (SESD), an algorithm for effective speech synthesis in modest data regimes through latent diffusion. It is based on a novel diffusion architecture, that we call U-Audio Transformer…

Sound · Computer Science 2024-09-06 Justin Lovelace , Soham Ray , Kwangyoun Kim , Kilian Q. Weinberger , Felix Wu

Sound event detection (SED) and acoustic scene classification (ASC) are major tasks in environmental sound analysis. Considering that sound events and scenes are closely related to each other, some works have addressed joint analyses of…

Existing systems for sound event localization and detection (SELD) typically operate by estimating a source location for all classes at every time instant. In this paper, we propose an alternative class-conditioned SELD model for situations…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-09 Olga Slizovskaia , Gordon Wichern , Zhong-Qiu Wang , Jonathan Le Roux

Sound event localization and detection (SELD) is a combined task of identifying the sound event and its direction. Deep neural networks (DNNs) are utilized to associate them with the sound signals observed by a microphone array. Although…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-18 Kento Nagatomo , Masahiro Yasuda , Kohei Yatabe , Shoichiro Saito , Yasuhiro Oikawa

This paper is concerned with self-supervised learning for small models. The problem is motivated by our empirical studies that while the widely used contrastive self-supervised learning method has shown great progress on large model…

Computer Vision and Pattern Recognition · Computer Science 2021-04-19 Zhiyuan Fang , Jianfeng Wang , Lijuan Wang , Lei Zhang , Yezhou Yang , Zicheng Liu

This paper proposes a benchmark of submissions to Detection and Classification Acoustic Scene and Events 2021 Challenge (DCASE) Task 4 representing a sampling of the state-of-the-art in Sound Event Detection task. The submissions are…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-02 Francesca Ronchini , Romain Serizel

This report proposes a frequency dynamic convolution (FDY) with a large kernel attention (LKA)-convolutional recurrent neural network (CRNN) with a pre-trained bidirectional encoder representation from audio transformers (BEATs)…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-13 Ji Won Kim , Sang Won Son , Yoonah Song , Hong Kook Kim , Il Hoon Song , Jeong Eun Lim