中文
相关论文

相关论文: The LOCATA Challenge: Acoustic Source Localization…

200 篇论文

Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a two-stage learning framework to perform self-supervised…

计算机视觉与模式识别 · 计算机科学 2020-10-13 Di Hu , Rui Qian , Minyue Jiang , Xiao Tan , Shilei Wen , Errui Ding , Weiyao Lin , Dejing Dou

Sound event localization and detection (SELD) is an emerging research topic that aims to unify the tasks of sound event detection and direction-of-arrival estimation. As a result, SELD inherits the challenges of both tasks, such as noise,…

音频与语音处理 · 电气工程与系统科学 2021-10-28 Thi Ngoc Tho Nguyen , Karn N. Watcharasupat , Zhen Jian Lee , Ngoc Khanh Nguyen , Douglas L. Jones , Woon Seng Gan

In this paper, we study the underwater acoustic localization in the presence of environmental mismatch. Especially, we exploit a pre-trained neural network for the acoustic wave propagation in a gradient-based optimization framework to…

声音 · 计算机科学 2025-04-01 Dariush Kari , Yongjie Zhuang , Andrew C. Singer

The goal of the multi-sound source localization task is to localize sound sources from the mixture individually. While recent multi-sound source localization methods have shown improved performance, they face challenges due to their…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Dongjin Kim , Sung Jin Um , Sangmin Lee , Jung Uk Kim

Formally verifying audio classification systems is essential to ensure accurate signal classification across real-world applications like surveillance, automotive voice commands, and multimedia content management, preventing potential…

声音 · 计算机科学 2023-11-22 Neelanjana Pal , Taylor T Johnson

This work focuses on reliable detection of bird sound emissions as recorded in the open field. Acoustic detection of avian sounds can be used for the automatized monitoring of multiple bird taxa and querying in long-term recordings for…

声音 · 计算机科学 2016-09-28 Ilyas Potamitis

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Hyeonggon Ryu , Seongyu Kim , Joon Son Chung , Arda Senocak

This survey paper provides a comprehensive overview of the recent advancements and challenges in applying large language models to the field of audio signal processing. Audio processing, with its diverse signal representations and a wide…

Urban noise maps and noise visualizations traditionally provide macroscopic representations of noise levels across cities. However, those representations fail at accurately gauging the sound perception associated with these sound…

计算机与社会 · 计算机科学 2024-07-25 Modan Tailleur , Pierre Aumond , Vincent Tourre , Mathieu Lagrange

This paper explores the data cleaning challenges that arise in using WiFi connectivity data to locate users to semantic indoor locations such as buildings, regions, rooms. WiFi connectivity data consists of sporadic connections between…

We study two cases of acoustic source localization in a reverberant room, from a number of point-wise narrowband measurements. In the first case, the room is perfectly known. We show that using a sparse recovery algorithm with a dictionary…

信息论 · 计算机科学 2013-07-19 Gilles Chardon , Laurent Daudet

Existing audio question answering benchmarks largely emphasize sound event classification or caption-grounded queries, often enabling models to succeed through shortcut strategies, short-duration cues, lexical priors, dataset-specific…

计算与语言 · 计算机科学 2026-04-24 Tasnim Kabir , Dmytro Kurdydyk , Aadi Palnitkar , Liam Dorn , Ahmed Haj Ahmed , Jordan Lee Boyd-Graber

In recent years, Event Sound Source Localization has been widely applied in various fields. Recent works typically relying on the contrastive learning framework show impressive performance. However, all work is based on large relatively…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Yue Li , Baiqiao Yin , Jinfu Liu , Jiajun Wen , Jiaying Lin , Mengyuan Liu

While direction of arrival (DOA) of sound events is generally estimated from multichannel audio data recorded in a microphone array, sound events usually derive from visually perceptible source objects, e.g., sounds of footsteps come from…

In multi-lingual societies, where multiple languages are spoken in a small geographic vicinity, informal conversations often involve mix of languages. Existing speech technologies may be inefficient in extracting information from such…

音频与语音处理 · 电气工程与系统科学 2024-01-04 Shikha Baghel , Shreyas Ramoji , Somil Jain , Pratik Roy Chowdhuri , Prachi Singh , Deepu Vijayasenan , Sriram Ganapathy

This paper presents a robust multi-channel speaker extraction algorithm designed to handle inaccuracies in reference information. While existing approaches often rely solely on either spatial or spectral cues to identify the target speaker,…

声音 · 计算机科学 2025-12-24 Aviad Eisenberg , Sharon Gannot , Shlomo E. Chazan

This challenge aims to evaluate the capabilities of audio encoders, especially in the context of multi-task learning and real-world applications. Participants are invited to submit pre-trained audio encoders that map raw waveforms to…

This technical report is an extended version of the paper 'Cooperative Multi-Target Localization With Noisy Sensors' accepted to the 2013 IEEE International Conference on Robotics and Automation (ICRA). This paper addresses the task of…

机器人学 · 计算机科学 2015-01-30 Philip Dames , Vijay Kumar

The use of wireless signals for purposes of localization enables a host of applications relating to the determination and verification of the positions of network participants, ranging from radar to satellite navigation. Consequently, it…

网络与互联网体系结构 · 计算机科学 2020-12-02 Matthias Schäfer , Martin Strohmeier , Mauro Leonardi , Vincent Lenders

Recent work on audio-visual navigation assumes a constantly-sounding target and restricts the role of audio to signaling the target's position. We introduce semantic audio-visual navigation, where objects in the environment make sounds…

计算机视觉与模式识别 · 计算机科学 2021-04-08 Changan Chen , Ziad Al-Halah , Kristen Grauman
‹ 上一页 1 8 9 10 下一页 ›