中文
相关论文

相关论文: Solution for Temporal Sound Localisation Task of E…

200 篇论文

The temporal answering grounding in the video (TAGV) is a new task naturally derived from temporal sentence grounding in the video (TSGV). Given an untrimmed video and a text question, this task aims at locating the matching span from the…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Bin Li , Yixuan Weng , Bin Sun , Shutao Li

The understanding of the surrounding environment plays a critical role in autonomous robotic systems, such as self-driving cars. Extensive research has been carried out concerning visual perception. Yet, to obtain a more complete perception…

音频与语音处理 · 电气工程与系统科学 2021-01-13 Karim Guirguis , Christoph Schorn , Andre Guntoro , Sherif Abdulatif , Bin Yang

Leveraging temporal synchronization and association within sight and sound is an essential step towards robust localization of sounding objects. To this end, we propose a space-time memory network for sounding object localization in videos.…

计算机视觉与模式识别 · 计算机科学 2021-11-11 Sizhe Li , Yapeng Tian , Chenliang Xu

Audio-Visual Source Localization (AVSL) aims to localize the source of sound within a video. In this paper, we identify a significant issue in existing benchmarks: the sounding objects are often easily recognized based solely on visual…

多媒体 · 计算机科学 2024-09-12 Liangyu Chen , Zihao Yue , Boshen Xu , Qin Jin

Sound source localization in visual scenes aims to localize objects emitting the sound in a given image. Recent works showing impressive localization performance typically rely on the contrastive learning framework. However, the random…

计算机视觉与模式识别 · 计算机科学 2022-03-28 Zengjie Song , Yuxi Wang , Junsong Fan , Tieniu Tan , Zhaoxiang Zhang

In traditional sound event localization and detection (SELD) tasks, the focus is typically on sound event detection (SED) and direction-of-arrival (DOA) estimation, but they fall short of providing full spatial information about the sound…

声音 · 计算机科学 2025-01-22 Yuxuan Dong , Qing Wang , Hengyi Hong , Ya Jiang , Shi Cheng

Temporal sentence grounding in videos (TSGV), \aka natural language video localization (NLVL) or video moment retrieval (VMR), aims to retrieve a temporal moment that semantically corresponds to a language query from an untrimmed video.…

计算机视觉与模式识别 · 计算机科学 2023-04-26 Hao Zhang , Aixin Sun , Wei Jing , Joey Tianyi Zhou

Sound event detection (SED) is the task of identifying sound events along with their onset and offset times. A recent, convolutional neural networks based SED method, proposed the usage of depthwise separable (DWS) and time-dilated…

声音 · 计算机科学 2020-07-13 Konstantinos Drossos , Stylianos I. Mimilakis , Tuomas Virtanen

Environment shifts and conflicts present significant challenges for learning-based sound event localization and detection (SELD) methods. SELD systems, when trained in particular acoustic settings, often show restricted generalization…

音频与语音处理 · 电气工程与系统科学 2024-10-08 Jinbo Hu , Yin Cao , Ming Wu , Qiuqiang Kong , Feiran Yang , Mark D. Plumbley , Jun Yang

Weakly Labelled learning has garnered lot of attention in recent years due to its potential to scale Sound Event Detection (SED) and is formulated as Multiple Instance Learning (MIL) problem. This paper proposes a Multi-Task Learning (MTL)…

音频与语音处理 · 电气工程与系统科学 2020-11-02 Soham Deshmukh , Bhiksha Raj , Rita Singh

Sound event detection (SED) methods typically rely on either strongly labelled data or weakly labelled data. As an alternative, sequentially labelled data (SLD) was proposed. In SLD, the events and the order of events in audio clips are…

声音 · 计算机科学 2019-04-30 Yuanbo Hou , Qiuqiang Kong , Shengchen Li , Mark D. Plumbley

Natural language video localization (NLVL) is a crucial task in video understanding that aims to localize the target moment in videos specified by a given language description. Recently, a point-supervised paradigm has been presented to…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Zhuo Tao , Liang Li , Qi Chen , Yunbin Tu , Zheng-Jun Zha , Ming-Hsuan Yang , Yuankai Qi , Qingming Huang

Continuous sign language recognition (cSLR) is a public significant task that transcribes a sign language video into an ordered gloss sequence. It is important to capture the fine-grained gloss-level details, since there is no explicit…

计算机视觉与模式识别 · 计算机科学 2021-07-28 Pan Xie , Zhi Cui , Yao Du , Mengyi Zhao , Jianwei Cui , Bin Wang , Xiaohui Hu

Audio-based video object segmentation aims to locate and segment objects in videos conditioned on audio cues, requiring precise understanding of both appearance and motion. Recent audio-driven video segmentation methods extend MLLMs by…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Zhiyu Wang , Xudong Kang , Shutao Li

Recent advancements in machine learning have fueled research on multimodal tasks, such as for instance text-to-video and text-to-audio retrieval. These tasks require models to understand the semantic content of video and audio data,…

信息检索 · 计算机科学 2024-09-04 Andreea-Maria Oncescu , João F. Henriques , A. Sophia Koepke

Sound event localization and detection with distance estimation (3D SELD) in video involves identifying active sound events at each time frame while estimating their spatial coordinates. This multimodal task requires joint reasoning across…

音频与语音处理 · 电气工程与系统科学 2026-01-27 Davide Berghi , Philip J. B. Jackson

Audio-Visual Source Localization (AVSL) is the task of identifying specific sounding objects in the scene given audio cues. In our work, we focus on semi-supervised AVSL with pseudo-labeling. To address the issues with vanilla hard…

计算机视觉与模式识别 · 计算机科学 2024-03-06 Yuxin Guo , Shijie Ma , Yuhao Zhao , Hu Su , Wei Zou

This technical report outlines our approach to Task 3A of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024, focusing on Sound Event Localization and Detection (SELD). SELD provides valuable insights by estimating…

声音 · 计算机科学 2025-07-25 Quoc Thinh Vo , David Han

Sound event detection (SED) and acoustic scene classification (ASC) are major tasks in environmental sound analysis. Considering that sound events and scenes are closely related to each other, some works have addressed joint analyses of…

Identification and localization of sounds are both integral parts of computational auditory scene analysis. Although each can be solved separately, the goal of forming coherent auditory objects and achieving a comprehensive spatial scene…

声音 · 计算机科学 2019-12-24 Ivo Trowitzsch , Christopher Schymura , Dorothea Kolossa , Klaus Obermayer