English
Related papers

Related papers: Improving Stereo 3D Sound Event Localization and D…

200 papers

A good joint training framework is very helpful to improve the performances of weakly supervised audio tagging (AT) and acoustic event detection (AED) simultaneously. In this study, we propose three methods to improve the best…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-15 Yunhao Liang , Yanhua Long , Yijie Li , Jiaen Liang , Yuping Wang

In this work, we propose a novel event based stereo method which addresses the problem of motion blur for a moving event camera. Our method uses the velocity of the camera and a range of disparities to synchronize the positions of the…

Computer Vision and Pattern Recognition · Computer Science 2018-10-22 Alex Zihao Zhu , Yibo Chen , Kostas Daniilidis

Acoustic Scene Classification (ASC) and Sound Event Detection (SED) are two separate tasks in the field of computational sound scene analysis. In this work, we present a new dataset with both sound scene and sound event labels and use this…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-02 Helen L. Bear , Ines Nolasco , Emmanouil Benetos

In this paper, we study the problem of 3D object detection from stereo images, in which the key challenge is how to effectively utilize stereo information. Different from previous methods using pixel-level depth maps, we propose employing…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Zengyi Qin , Jinglu Wang , Yan Lu

An important problem in machine auditory perception is to recognize and detect sound events. In this paper, we propose a sequential self-teaching approach to learning sounds. Our main proposition is that it is harder to learn sounds in…

Sound · Computer Science 2020-07-02 Anurag Kumar , Vamsi Krishna Ithapu

Identification and localization of sounds are both integral parts of computational auditory scene analysis. Although each can be solved separately, the goal of forming coherent auditory objects and achieving a comprehensive spatial scene…

Sound · Computer Science 2019-12-24 Ivo Trowitzsch , Christopher Schymura , Dorothea Kolossa , Klaus Obermayer

Recent advances in generating synthetic captions based on audio and related metadata allow using the information contained in natural language as input for other audio tasks. In this paper, we propose a novel method to guide a sound event…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-29 Manu Harju , Annamaria Mesaros

Humans are able to localize objects in the environment using both visual and auditory cues, integrating information from multiple modalities into a common reference frame. We introduce a system that can leverage unlabeled audio-visual data…

Computer Vision and Pattern Recognition · Computer Science 2019-10-28 Chuang Gan , Hang Zhao , Peihao Chen , David Cox , Antonio Torralba

This paper introduces the task description for the Detection and Classification of Acoustic Scenes and Events (DCASE) 2025 Challenge Task 2, titled "First-shot unsupervised anomalous sound detection (ASD) for machine condition monitoring".…

This paper introduces Binaural Sound Event Localization and Detection (BiSELD), a task that aims to jointly detect and localize multiple sound events using binaural audio, inspired by the spatial hearing mechanism of humans. To support this…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-29 Gyeong-Tae Lee , Hyeonuk Nam , Yong-Hwa Park

This report introduces our novel method named STHG for the Audio-Visual Diarization task of the Ego4D Challenge 2023. Our key innovation is that we model all the speakers in a video using a single, unified heterogeneous graph learning…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Kyle Min

Human listeners exhibit the remarkable ability to segregate a desired sound from complex acoustic scenes through selective auditory attention, motivating the study of Targeted Sound Detection (TSD). The task requires detecting and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-19 Shubham Gupta , Adarsh Arigala , B. R. Dilleswari , Sri Rama Murty Kodukula

Target sound extraction (TSE) aims to extract the sound part of a target sound event class from a mixture audio with multiple sound events. The previous works mainly focus on the problems of weakly-labelled data, jointly learning and new…

Sound · Computer Science 2022-04-05 Helin Wang , Dongchao Yang , Chao Weng , Jianwei Yu , Yuexian Zou

Anomalous Sound Detection (ASD) has gained significant interest through the application of various Artificial Intelligence (AI) technologies in industrial settings. Though possessing great potential, ASD systems can hardly be readily…

Sound · Computer Science 2025-05-08 Xinhu Zheng , Anbai Jiang , Bing Han , Yanmin Qian , Pingyi Fan , Jia Liu , Wei-Qiang Zhang

The L3DAS22 Challenge is aimed at encouraging the development of machine learning strategies for 3D speech enhancement and 3D sound localization and detection in office-like environments. This challenge improves and extends the tasks of the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-16 Eric Guizzo , Christian Marinoni , Marco Pennese , Xinlei Ren , Xiguang Zheng , Chen Zhang , Bruno Masiero , Aurelio Uncini , Danilo Comminiello

Convolutional recurrent neural networks (CRNNs) have achieved state-of-the-art performance for sound event detection (SED). In this paper, we propose to use a dilated CRNN, namely a CRNN with a dilated convolutional kernel, as the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-21 Yanxiong Li , Mingle Liu , Konstantinos Drossos , Tuomas Virtanen

Consistency regularization (CR), which enforces agreement between model predictions on augmented views, has found recent benefits in automatic speech recognition [1]. In this paper, we propose the use of consistency regularization for audio…

Sound · Computer Science 2025-09-15 Shanmuka Sadhu , Weiran Wang

In this paper, we present a gated convolutional recurrent neural network based approach to solve task 4, large-scale weakly labelled semi-supervised sound event detection in domestic environments, of the DCASE 2018 challenge. Gated linear…

Sound · Computer Science 2018-10-17 Robert Harb , Franz Pernkopf

Current mainstream audio generation methods primarily rely on simple text prompts, often failing to capture the nuanced details necessary for multi-style audio generation. To address this limitation, the Sound Event Enhanced Prompt Adapter…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Chenxu Xiong , Ruibo Fu , Shuchen Shi , Zhengqi Wen , Jianhua Tao , Tao Wang , Chenxing Li , Chunyu Qiang , Yuankun Xie , Xin Qi , Guanjun Li , Zizheng Yang

Deep learning-based Sound Event Localization and Detection (SELD) systems degrade significantly on real-world, long-tailed datasets. Standard regression losses bias learning toward frequent classes, causing rare events to be systematically…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-22 Jun-Wei Yeow , Ee-Leng Tan , Santi Peksi , Woon-Seng Gan
‹ Prev 1 8 9 10 Next ›