English
Related papers

Related papers: Leveraging the Video-level Semantic Consistency of…

200 papers

Dense audio-visual event localization (DAVE) aims to identify event categories and locate the temporal boundaries in untrimmed videos. Most studies only employ event-related semantic constraints on the final outputs, lacking cross-modal…

Multimedia · Computer Science 2025-10-16 Huilai Li , Yonghao Dang , Ying Xing , Yiming Wang , Jianqin Yin

The audio-visual event localization task requires identifying concurrent visual and auditory events from unconstrained videos within a network model, locating them, and classifying their category. The efficient extraction and integration of…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Xiang He , Xiangxi Liu , Yang Li , Dongcheng Zhao , Guobin Shen , Qingqun Kong , Xin Yang , Yi Zeng

As a vital topic in media content interpretation, video anomaly detection (VAD) has made fruitful progress via deep neural network (DNN). However, existing methods usually follow a reconstruction or frame prediction routine. They suffer…

Computer Vision and Pattern Recognition · Computer Science 2020-08-28 Guang Yu , Siqi Wang , Zhiping Cai , En Zhu , Chuanfu Xu , Jianping Yin , Marius Kloft

The Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both. Capturing accurate event semantics for each audio/visual segment is vital. Prior works…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Pengcheng Zhao , Jinxing Zhou , Yang Zhao , Dan Guo , Yanxiang Chen

An audio-visual event (AVE) is denoted by the correspondence of the visual and auditory signals in a video segment. Precise localization of the AVEs is very challenging since it demands effective multi-modal feature correspondence to ground…

Computer Vision and Pattern Recognition · Computer Science 2022-10-12 Tanvir Mahmud , Diana Marculescu

Learning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations between vision and language data, such as information density, i.e., images can contain…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Yang Liu , Mengyuan Liu , Shudong Huang , Jiancheng Lv

As a critical clue of video super-resolution (VSR), inter-frame alignment significantly impacts overall performance. However, accurate pixel-level alignment is a challenging task due to the intricate motion interweaving in the video. In…

Computer Vision and Pattern Recognition · Computer Science 2024-01-22 Qi Tang , Yao Zhao , Meiqin Liu , Jian Jin , Chao Yao

In this paper, we introduce a novel problem of audio-visual event localization in unconstrained videos. We define an audio-visual event as an event that is both visible and audible in a video segment. We collect an Audio-Visual Event(AVE)…

Computer Vision and Pattern Recognition · Computer Science 2018-03-26 Yapeng Tian , Jing Shi , Bochen Li , Zhiyao Duan , Chenliang Xu

In this paper, we propose a framework centering around a novel architecture called the Event Decomposition Recomposition Network (EDRNet) to tackle the Audio-Visual Event (AVE) localization problem in the supervised and weakly supervised…

Computer Vision and Pattern Recognition · Computer Science 2021-12-23 Varshanth R. Rao , Md Ibrahim Khalil , Haoda Li , Peng Dai , Juwei Lu

Multimedia event detection is the task of detecting a specific event of interest in an user-generated video on websites. The most fundamental challenge facing this task lies in the enormously varying quality of the video as well as the…

Computer Vision and Pattern Recognition · Computer Science 2021-10-18 Minnan Luo , Xiaojun Chang , Chen Gong

Autonomous driving systems rely on robust 3D scene understanding. Recent advances in Semantic Scene Completion (SSC) for autonomous driving underscore the limitations of RGB-based approaches, which struggle under motion blur, poor lighting,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Shangwei Guo , Hao Shi , Song Wang , Xiaoting Yin , Kailun Yang , Kaiwei Wang

Event cameras harness advantages such as low latency, high temporal resolution, and high dynamic range (HDR), compared to standard cameras. Due to the distinct imaging paradigm shift, a dominant line of research focuses on event-to-video…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Kanghao Chen , Hangyu Li , JiaZhou Zhou , Zeyu Wang , Lin Wang

Segmenting video content into events provides semantic structures for indexing, retrieval, and summarization. Since motion cues are not available in continuous photo-streams, and annotations in lifelogging are scarce and costly, the frames…

Computer Vision and Pattern Recognition · Computer Science 2018-08-08 Ana Garcia del Molino , Joo-Hwee Lim , Ah-Hwee Tan

Video Moment Retrieval (VMR) aims to retrieve temporal segments in untrimmed videos corresponding to a given language query by constructing cross-modal alignment strategies. However, these existing strategies are often sub-optimal since…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Zhihang Liu , Jun Li , Hongtao Xie , Pandeng Li , Jiannan Ge , Sun-Ao Liu , Guoqing Jin

Visual Semantic Embedding (VSE) aims to extract the semantics of images and their descriptions, and embed them into the same latent space for cross-modal information retrieval. Most existing VSE networks are trained by adopting a hard…

Computer Vision and Pattern Recognition · Computer Science 2023-02-15 Yan Gong , Georgina Cosma

Visual neural decoding aims to extract and interpret original visual experiences directly from human brain activity. Recent studies have demonstrated the feasibility of decoding visual semantic categories from electroencephalography (EEG)…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Hongzhou Chen , Lianghua He , Yihang Liu , Longzhen Yang , Shaohua Shang , MengChu Zhou

Visual and audio signals often coexist in natural environments, forming audio-visual events (AVEs). Given a video, we aim to localize video segments containing an AVE and identify its category. In order to learn discriminative features for…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Jinxing Zhou , Liang Zheng , Yiran Zhong , Shijie Hao , Meng Wang

The exponential growth in wireless data traffic, driven by the proliferation of mobile devices and smart applications, poses significant challenges for modern communication systems. Ensuring the secure and reliable transmission of…

Signal Processing · Electrical Eng. & Systems 2024-11-05 Yuandi Li , Zhe Xiang , Fei Yu , Zhangshuang Guan , Hui Ji , Zhiguo Wan , Cheng Feng

Visible-infrared person re-identification (VIReID) retrieves pedestrian images with the same identity across different modalities. Existing methods learn visual content solely from images, lacking the capability to sense high-level…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Neng Dong , Shuanglin Yan , Liyan Zhang , Jinhui Tang

We study the problem of localizing audio-visual events that are both audible and visible in a video. Existing works focus on encoding and aligning audio and visual features at the segment level while neglecting informative correlation…

Computer Vision and Pattern Recognition · Computer Science 2021-08-31 Hao Wang , Zheng-Jun Zha , Liang Li , Xuejin Chen , Jiebo Luo
‹ Prev 1 2 3 10 Next ›