中文
相关论文

相关论文: Audio-Visual Event Localization in Unconstrained V…

200 篇论文

Weakly-supervised audio-visual video parsing (WS-AVVP) aims to localize the temporal extents of audio, visual and audio-visual event instances as well as identify the corresponding event categories with only video-level category labels for…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Jie Fu , Junyu Gao , Changsheng Xu

Recent work on audio-visual navigation targets a single static sound in noise-free audio environments and struggles to generalize to unheard sounds. We introduce the novel dynamic audio-visual navigation benchmark in which an embodied AI…

计算机视觉与模式识别 · 计算机科学 2022-01-13 Abdelrahman Younes

We address the problem of retrieving a specific moment from an untrimmed video by a query sentence. This is a challenging problem because a target moment may take place in relations to other temporal moments in the untrimmed video. Existing…

计算机视觉与模式识别 · 计算机科学 2020-12-29 Songyang Zhang , Houwen Peng , Jianlong Fu , Jiebo Luo

This paper studies the problem of temporal moment localization in a long untrimmed video using natural language as the query. Given an untrimmed video and a sentence as the query, the goal is to determine the starting, and the ending, of…

计算机视觉与模式识别 · 计算机科学 2020-03-13 Cristian Rodriguez-Opazo , Edison Marrese-Taylor , Fatemeh Sadat Saleh , Hongdong Li , Stephen Gould

We propose a novel self-supervised approach for learning audio and visual representations from unlabeled videos, based on their correspondence. The approach uses an attention mechanism to learn the relative importance of convolutional…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Sudha Krishnamurthy

Video-based person re-identification (Re-ID) aims at matching video sequences of pedestrians across non-overlapping cameras. It is a practical yet challenging task of how to embed spatial and temporal information of a video into its feature…

计算机视觉与模式识别 · 计算机科学 2019-08-06 Chih-Ting Liu , Chih-Wei Wu , Yu-Chiang Frank Wang , Shao-Yi Chien

In this paper, we address the challenges in unsupervised video object segmentation (UVOS) by proposing an efficient algorithm, termed MTNet, which concurrently exploits motion and temporal cues. Unlike previous methods that focus solely on…

计算机视觉与模式识别 · 计算机科学 2025-01-15 Yunzhi Zhuge , Hongyu Gu , Lu Zhang , Jinqing Qi , Huchuan Lu

Supervised (pre-)training currently yields state-of-the-art performance for representation learning for visual recognition, yet it comes at the cost of (1) intensive manual annotations and (2) an inherent restriction in the scope of data…

计算机视觉与模式识别 · 计算机科学 2016-12-05 Ruohan Gao , Dinesh Jayaraman , Kristen Grauman

With the assumption that a video dataset is multimodality annotated in which auditory and visual modalities both are labeled or class-relevant, current multimodal methods apply modality fusion or cross-modality attention. However,…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Saghir Alfasly , Jian Lu , Chen Xu , Yuru Zou

Temporal action segmentation in untrimmed videos has gained increased attention recently. However, annotating action classes and frame-wise boundaries is extremely time consuming and cost intensive, especially on large-scale datasets. To…

计算机视觉与模式识别 · 计算机科学 2023-03-10 Wei Lin , Anna Kukleva , Horst Possegger , Hilde Kuehne , Horst Bischof

This paper presents our approach for the VA (Valence-Arousal) estimation task in the ABAW6 competition. We devised a comprehensive model by preprocessing video frames and audio segments to extract visual and audio features. Through the…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Jun Yu , Gongpeng Zhao , Yongqi Wang , Zhihong Wei , Yang Zheng , Zerui Zhang , Zhongpeng Cai , Guochen Xie , Jichao Zhu , Wangyuan Zhu

Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a two-stage learning framework to perform self-supervised…

计算机视觉与模式识别 · 计算机科学 2020-10-13 Di Hu , Rui Qian , Minyue Jiang , Xiao Tan , Shilei Wen , Errui Ding , Weiyao Lin , Dejing Dou

It has already been observed that audio-visual embedding is more robust than uni-modality embedding for person verification. Here, we proposed a novel audio-visual strategy that considers aggregators from a fusion perspective. First, we…

计算机视觉与模式识别 · 计算机科学 2022-10-27 Peiwen Sun , Shanshan Zhang , Zishan Liu , Yougen Yuan , Taotao Zhang , Honggang Zhang , Pengfei Hu

In the field of acoustic scene analysis, this paper presents a novel approach to find spatio-temporal latent representations from in-the-wild audio data. By using WE-LIVE, an in-house collected dataset that includes audio recordings in…

音频与语音处理 · 电气工程与系统科学 2024-12-11 Claudia Montero-Ramírez , Esther Rituerto-González , Carmen Peláez-Moreno

Audiovisual embodied navigation enables robots to locate audio sources by dynamically integrating visual observations from onboard sensors with the auditory signals emitted by the target. The core challenge lies in effectively leveraging…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Yinfeng Yu , Hailong Zhang , Meiling Zhu

Prior works on text-based video moment localization focus on temporally grounding the textual query in an untrimmed video. These works assume that the relevant video is already known and attempt to localize the moment on that relevant video…

计算机视觉与模式识别 · 计算机科学 2021-11-10 Sudipta Paul , Niluthpol Chowdhury Mithun , Amit K. Roy-Chowdhury

Audio-visual speech recognition (AVSR) system is thought to be one of the most promising solutions for robust speech recognition, especially in noisy environment. In this paper, we propose a novel multimodal attention based method for…

计算与语言 · 计算机科学 2019-04-24 Pan Zhou , Wenwen Yang , Wei Chen , Yanfeng Wang , Jia Jia

Recent dense audio-visual (AV) models achieve impressive retrieval and emergent localization, but almost all evidence comes from English-centric, caption-rich web video. It is unclear whether these objectives survive in low-resource,…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Sajay Raj

Audio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audio-visual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensive fields of dense…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Juhyeong Seon , Woobin Im , Sebin Lee , Jumin Lee , Sung-Eui Yoon

We address temporal localization of events in large-scale video data, in the context of the Youtube-8M Segments dataset. This emerging field within video recognition can enable applications to identify the precise time a specified event…

计算机视觉与模式识别 · 计算机科学 2019-10-28 Mikel Bober-Irizar , Miha Skalic , David Austin