English
Related papers

Related papers: Multimodal Urban Sound Tagging with Spatiotemporal…

200 papers

Acoustic scene recordings are often collected from a diverse range of cities. Most existing acoustic scene classification (ASC) approaches focus on identifying common acoustic scene patterns across cities to enhance generalization. However,…

Sound · Computer Science 2025-06-16 Yiqiang Cai , Yizhou Tan , Shengchen Li , Xi Shao , Mark D. Plumbley

The analysis, processing, and extraction of meaningful information from sounds all around us is the subject of the broader area of audio analytics. Audio captioning is a recent addition to the domain of audio analytics, a cross-modal…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-04 Sandeep Kothinti , Dimitra Emmanouilidou

Multimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Xiantao Hu , Ying Tai , Xu Zhao , Chen Zhao , Zhenyu Zhang , Jun Li , Bineng Zhong , Jian Yang

This work presents a text-to-audio-retrieval system based on pre-trained text and spectrogram transformers. Our method projects recordings and textual descriptions into a shared audio-caption space in which related examples from different…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-09 Paul Primus , Khaled Koutini , Gerhard Widmer

Audio perception is a key to solving a variety of problems ranging from acoustic scene analysis, music meta-data extraction, recommendation, synthesis and analysis. It can potentially also augment computers in doing tasks that humans do…

Sound · Computer Science 2020-02-12 Prateek Verma , Kenneth Salisbury

In recent years, multi-modal fusion has attracted a lot of research interest, both in academia, and in industry. Multimodal fusion entails the combination of information from a set of different types of sensors. Exploiting complementary…

Machine Learning · Computer Science 2020-08-27 Siddharth Roheda , Hamid Krim , Benjamin S. Riggan

This paper proposes to use low-level spatial features extracted from multichannel audio for sound event detection. We extend the convolutional recurrent neural network to handle more than one type of these multichannel features by learning…

Sound · Computer Science 2017-06-09 Sharath Adavanne , Pasi Pertilä , Tuomas Virtanen

This paper presents a robust multi-channel speaker extraction algorithm designed to handle inaccuracies in reference information. While existing approaches often rely solely on either spatial or spectral cues to identify the target speaker,…

Sound · Computer Science 2025-12-24 Aviad Eisenberg , Sharon Gannot , Shlomo E. Chazan

This technical report presents submission systems for Task 4 of the DCASE 2025 Challenge. This model incorporates additional audio features (spectral roll-off and chroma features) into the embedding feature extracted from the mel-spectral…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-27 Jongyeon Park , Joonhee Lee , Do-Hyeon Lim , Hong Kook Kim , Hyeongcheol Geum , Jeong Eun Lim

High quality labeled datasets have allowed deep learning to achieve impressive results on many sound analysis tasks. Yet, it is labor-intensive to accurately annotate large amount of audio data, and the dataset may contain noisy labels in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-17 Boqing Zhu , Kele Xu , Qiuqiang Kong , Huaimin Wang , Yuxing Peng

A range of applications of multi-modal music information retrieval is centred around the problem of connecting large collections of sheet music (images) to corresponding audio recordings, that is, identifying pairs of audio and score…

Sound · Computer Science 2023-09-22 Luis Carvalho , Gerhard Widmer

Target sound extraction (TSE) aims to extract the sound part of a target sound event class from a mixture audio with multiple sound events. The previous works mainly focus on the problems of weakly-labelled data, jointly learning and new…

Sound · Computer Science 2022-04-05 Helin Wang , Dongchao Yang , Chao Weng , Jianwei Yu , Yuexian Zou

Previous studies in automated audio captioning have faced difficulties in accurately capturing the complete temporal details of acoustic scenes and events within long audio sequences. This paper presents AudioLog, a large language models…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-05 Jisheng Bai , Han Yin , Mou Wang , Dongyuan Shi , Woon-Seng Gan , Jianfeng Chen , Susanto Rahardja

This paper introduces a curated dataset of urban scenes for audio-visual scene analysis which consists of carefully selected and recorded material. The data was recorded in multiple European cities, using the same equipment, in multiple…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-12 Shanshan Wang , Annamaria Mesaros , Toni Heittola , Tuomas Virtanen

Multimodal learning systems often face substantial uncertainty due to noisy data, low-quality labels, and heterogeneous modality characteristics. These issues become especially critical in human-computer interaction settings, where data…

Artificial Intelligence · Computer Science 2025-11-21 Hyo-Jeong Jang

A noise map facilitates the monitoring of environmental noise pollution in urban areas. However, state-of-the-art techniques for rendering noise maps in urban areas are expensive and rarely updated, as they rely on population and traffic…

Other Computer Science · Computer Science 2013-10-17 Rajib Rana , Chun Tung Chou , Nirupama Bulusu , Salil Kanhere , Wen Hu

Target speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extraction has been proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-03 Hiroshi Sato , Tsubasa Ochiai , Keisuke Kinoshita , Marc Delcroix , Tomohiro Nakatani , Shoko Araki

This paper presents a novel approach for enhancing the multiple sets of acoustic patterns automatically discovered from a given corpus. In a previous work it was proposed that different HMM configurations (number of states per model, number…

Computation and Language · Computer Science 2015-09-09 Cheng-Tao Chung , Wei-Ning Hsu , Cheng-Yi Lee , Lin-Shan Lee

Human beings can perceive a target sound type from a multi-source mixture signal by the selective auditory attention, however, such functionality was hardly ever explored in machine hearing. This paper addresses the target sound detection…

Sound · Computer Science 2022-07-08 Dongchao Yang , Helin Wang , Yuexian Zou , Fan Cui , Yujun Wang

Acoustic scene classification (ASC) and sound event detection (SED) are major topics in environmental sound analysis. Considering that acoustic scenes and sound events are closely related to each other, the joint analysis of acoustic scenes…

Sound · Computer Science 2022-06-22 Kayo Nada , Keisuke Imoto , Takao Tsuchiya