English
Related papers

Related papers: Frame Aggregation and Multi-Modal Fusion Framework…

200 papers

Recognition of low-quality face images remains a challenge due to invisible or deformation in partial facial regions. For low-quality images dominated by missing partial facial regions, local region similarity contributes more to face…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Wang Yu , Wei Wei

Existing deepfake detectors face several challenges in achieving robustness and generalization. One of the primary reasons is their limited ability to extract relevant information from forgery videos, especially in the presence of various…

Computer Vision and Pattern Recognition · Computer Science 2023-05-01 Zhiyuan Yan , Peng Sun , Yubo Lang , Shuo Du , Shanzhuo Zhang , Wei Wang , Lei Liu

Recently, deep neural network has shown promising performance in face image recognition. The inputs of most networks are face images, and there is hardly any work reported in literature on network with face videos as input. To sufficiently…

Computer Vision and Pattern Recognition · Computer Science 2016-03-23 Zhen Dong , Su Jia , Chi Zhang , Mingtao Pei

Video-based person re-identification (ReID) is challenging due to the presence of various interferences in video frames. Recent approaches handle this problem using temporal aggregation strategies. In this work, we propose a novel Context…

Computer Vision and Pattern Recognition · Computer Science 2022-07-07 Kan Wang , Changxing Ding , Jianxin Pang , Xiangmin Xu

In this paper we propose the two-stage approach of organizing information in video surveillance systems. At first, the faces are detected in each frame and a video stream is split into sequences of frames with face region of one person.…

Computer Vision and Pattern Recognition · Computer Science 2018-01-04 Anastasiia D. Sokolova , Angelina S. Kharchevnikova , Andrey V. Savchenko

Recent advancements in video understanding within visual large language models (VLLMs) have led to notable progress. However, the complexity of video data and contextual processing limitations still hinder long-video comprehension. A common…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Yanan Guo , Wenhui Dong , Jun Song , Shiding Zhu , Xuan Zhang , Hanqing Yang , Yingbo Wang , Yang Du , Xianing Chen , Bo Zheng

State-of-the-art video object detection methods maintain a memory structure, either a sliding window or a memory queue, to enhance the current frame using attention mechanisms. However, we argue that these memory structures are not…

Computer Vision and Pattern Recognition · Computer Science 2024-02-02 Guanxiong Sun , Yang Hua , Guosheng Hu , Neil Robertson

Human pose estimation plays an important role in many computer vision tasks and has been studied for many decades. However, due to complex appearance variations from poses, illuminations, occlusions and low resolutions, it still remains a…

Computer Vision and Pattern Recognition · Computer Science 2019-05-21 Zhihui Su , Ming Ye , Guohui Zhang , Lei Dai , Jianda Sheng

Video anomaly detection (VAD) is a challenging task that detects anomalous frames in continuous surveillance videos. Most previous work utilizes the spatio-temporal correlation of visual features to distinguish whether there are…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Guangyu Dai , Dong Chen , Siliang Tang , Yueting Zhuang

We consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in the image or video. In…

Computer Vision and Pattern Recognition · Computer Science 2021-02-10 Linwei Ye , Mrigank Rochan , Zhi Liu , Xiaoqin Zhang , Yang Wang

Cross-modality fusing complementary information of multispectral remote sensing image pairs can improve the perception ability of detection algorithms, making them more robust and reliable for a wider range of applications, such as…

Computer Vision and Pattern Recognition · Computer Science 2021-12-07 Qingyun Fang , Zhaokui Wang

Recent cutting-edge feature aggregation paradigms for video object detection rely on inferring feature correspondence. The feature correspondence estimation problem is fundamentally difficult due to poor image quality, motion blur, etc, and…

Computer Vision and Pattern Recognition · Computer Science 2019-07-12 Hao Luo , Lichao Huang , Han Shen , Yuan Li , Chang Huang , Xinggang Wang

How do humans recognize an object in a piece of video? Due to the deteriorated quality of single frame, it may be hard for people to identify an occluded object in this frame by just utilizing information within one image. We argue that…

Computer Vision and Pattern Recognition · Computer Science 2020-03-27 Yihong Chen , Yue Cao , Han Hu , Liwei Wang

Visual Place Recognition is a challenging task for robotics and autonomous systems, which must deal with the twin problems of appearance and viewpoint change in an always changing world. This paper introduces Patch-NetVLAD, which provides a…

Computer Vision and Pattern Recognition · Computer Science 2021-03-03 Stephen Hausler , Sourav Garg , Ming Xu , Michael Milford , Tobias Fischer

Multimodal object detection improves robustness in chal- lenging conditions by leveraging complementary cues from multiple sensor modalities. We introduce Filtered Multi- Modal Cross Attention Fusion (FMCAF), a preprocess- ing architecture…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Jad Berjawi , Yoann Dupas , Christophe C'erin

Image Forgery Localization (IFL) technology aims to detect and locate the forged areas in an image, which is very important in the field of digital forensics. However, existing IFL methods suffer from feature degradation during training…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Yakun Niu , Pei Chen , Lei Zhang , Lei Tan , Yingjian Chen

We summarize our TRECVID 2022 Ad-hoc Video Search (AVS) experiments. Our solution is built with two new techniques, namely Lightweight Attentional Feature Fusion (LAFF) for combining diverse visual / textual features and Bidirectional…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Xirong Li , Aozhu Chen , Ziyue Wang , Fan Hu , Kaibin Tian , Xinru Chen , Chengbo Dong

Video-based person re-identification matches video clips of people across non-overlapping cameras. Most existing methods tackle this problem by encoding each video frame in its entirety and computing an aggregate representation across all…

Computer Vision and Pattern Recognition · Computer Science 2018-03-28 Shuang Li , Slawomir Bak , Peter Carr , Xiaogang Wang

Models based on self-attention mechanisms have been successful in analyzing temporal data and have been widely used in the natural language domain. We propose a new model architecture for video face representation and recognition based on a…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Ihor Protsenko , Taras Lehinevych , Dmytro Voitekh , Ihor Kroosh , Nick Hasty , Anthony Johnson

Face detection has drawn much attention in recent decades since the seminal work by Viola and Jones. While many subsequences have improved the work with more powerful learning algorithms, the feature representation used for face detection…

Computer Vision and Pattern Recognition · Computer Science 2014-09-04 Bin Yang , Junjie Yan , Zhen Lei , Stan Z. Li