English
Related papers

Related papers: Multimodal Urban Sound Tagging with Spatiotemporal…

200 papers

Audio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audio-visual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensive fields of dense…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Juhyeong Seon , Woobin Im , Sebin Lee , Jumin Lee , Sung-Eui Yoon

Current multichannel speech enhancement algorithms typically assume a stationary sound source, a common mismatch with reality that limits their performance in real-world scenarios. This paper focuses on attention-driven spatial filtering…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-19 Yuzhu Wang , Archontis Politis , Tuomas Virtanen

Target audio source separation with natural language queries presents a promising paradigm for extracting arbitrary audio events through arbitrary text descriptions. Existing methods mainly face two challenges, the difficulty in jointly…

Sound · Computer Science 2025-12-03 Xinlei Yin , Xiulian Peng , Xue Jiang , Zhiwei Xiong , Yan Lu

Audio-based pedestrian detection is a challenging task and has, thus far, only been explored in noise-limited environments. We present a new dataset, results, and a detailed analysis of the state-of-the-art in audio-based pedestrian…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Yonghyun Kim , Chaeyeon Han , Akash Sarode , Noah Posner , Subhrajit Guhathakurta , Alexander Lerch

Urban spatio-temporal prediction is crucial for informed decision-making, such as traffic management, resource optimization, and emergence response. Despite remarkable breakthroughs in pretrained natural language models that enable one…

Machine Learning · Computer Science 2024-07-02 Yuan Yuan , Jingtao Ding , Jie Feng , Depeng Jin , Yong Li

Underwater noise pollution from human activities, particularly shipping, has been recognised as a serious threat to marine life. The sound generated by vessels can have various adverse effects on fish and aquatic ecosystems in general. In…

Stance detection plays a pivotal role in enabling an extensive range of downstream applications, from discourse parsing to tracing the spread of fake news and the denial of scientific facts. While most stance classification models rely on…

Computation and Language · Computer Science 2024-12-13 Guy Barel , Oren Tsur , Dan Vilenchik

Environmental sound classification is a field of growing importance for urban monitoring and cultural soundscape analysis, especially within the acoustically rich environments of South Asia. These regions present a unique challenge as…

Surgical instrument segmentation is instrumental to minimally invasive surgeries and related applications. Most previous methods formulate this task as single-frame-based instance segmentation while ignoring the natural temporal and stereo…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Qiyuan Wang , Shang Zhao , Zikang Xu , S Kevin Zhou

This paper presents MAST, a new model for Multimodal Abstractive Text Summarization that utilizes information from all three modalities -- text, audio and video -- in a multimodal video. Prior work on multimodal abstractive text…

Computation and Language · Computer Science 2020-10-19 Aman Khullar , Udit Arora

Recent advances in active noise control have enabled the development of hearables with spatial selectivity, which actively suppress undesired noise while preserving desired sound from specific directions. In this work, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-16 Tong Xiao , Simon Doclo

Most recent work in visual sound source localization relies on semantic audio-visual representations learned in a self-supervised manner, and by design excludes temporal information present in videos. While it proves to be effective for…

Sound · Computer Science 2023-04-18 Rajsuryan Singh , Pablo Zinemanas , Xavier Serra , Juan Pablo Bello , Magdalena Fuentes

Attributes of sound inherent to objects can provide valuable cues to learn rich representations for object detection and tracking. Furthermore, the co-occurrence of audiovisual events in videos can be exploited to localize objects over the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-05 Francisco Rivera Valverde , Juana Valeria Hurtado , Abhinav Valada

Most existing sound event detection~(SED) algorithms operate under a closed-set assumption, restricting their detection capabilities to predefined classes. While recent efforts have explored language-driven zero-shot SED by exploiting…

Sound · Computer Science 2025-10-28 Pengfei Cai , Yan Song , Qing Gu , Nan Jiang , Haoyu Song , Ian McLoughlin

Traffic forecasting, crucial for urban planning, requires accurate predictions of spatial-temporal traffic patterns across urban areas. Existing research mainly focuses on designing complex models that capture spatial-temporal dependencies…

Machine Learning · Computer Science 2024-07-30 Jiarui Sun , Yujie Fan , Chin-Chia Michael Yeh , Wei Zhang , Girish Chowdhary

Sound event detection (SED) entails identifying the type of sound and estimating its temporal boundaries from acoustic signals. These events are uniquely characterized by their spatio-temporal features, which are determined by the way they…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-19 Tanmay Khandelwal , Rohan Kumar Das

Sound event localization and detection (SELD) is a task for the classification of sound events and the identification of direction of arrival (DoA) utilizing multichannel acoustic signals. For effective classification and localization, a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-18 Yusun Shul , Dayun Choi , Jung-Woo Choi

We introduce and explore a new multimodal input representation for vision-language models: acoustic field video. Unlike conventional video (RGB with stereo/mono audio), our video stream provides a spatially grounded visualization of sound…

Human-Computer Interaction · Computer Science 2026-01-27 Daehwa Kim , Chris Harrison

Modern audio source separation techniques rely on optimizing sequence model architectures such as, 1D-CNNs, on mixture recordings to generalize well to unseen mixtures. Specifically, recent focus is on time-domain based architectures such…

Machine Learning · Computer Science 2019-04-09 Vivek Sivaraman Narayanaswamy , Sameeksha Katoch , Jayaraman J. Thiagarajan , Huan Song , Andreas Spanias

Air pollution remains one of the most formidable environmental threats to human health globally, particularly in urban areas, contributing to nearly 7 million premature deaths annually. Megacities, defined as cities with populations…

Machine Learning · Computer Science 2024-07-17 Harun Khan , Joseph Tso , Nathan Nguyen , Nivaan Kaushal , Ansh Malhotra , Nayel Rehman