English
Related papers

Related papers: Recursive Fusion and Deformable Spatiotemporal Att…

200 papers

Image recapture seriously breaks the fairness of artificial intelligent (AI) systems, which deceives the system by recapturing others' images. Most of the existing recapture models can only address a single pattern of recapture (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-14 Shuyu Miao , Lin Zheng , Hong Jin

In recent years, Deep Learning has been successfully applied to multimodal learning problems, with the aim of learning useful joint representations in data fusion applications. When the available modalities consist of time series data such…

Computer Vision and Pattern Recognition · Computer Science 2017-04-12 Xitong Yang , Palghat Ramesh , Radha Chitta , Sriganesh Madhvanath , Edgar A. Bernal , Jiebo Luo

Early interlaced videos usually contain multiple and interlacing and complex compression artifacts, which significantly reduce the visual quality. Although the high-definition reconstruction technology for early videos has made great…

Image and Video Processing · Electrical Eng. & Systems 2022-10-19 Yang Zhao , Yanbo Ma , Yuan Chen , Wei Jia , Ronggang Wang , Xiaoping Liu

Video face restoration aims to enhance degraded face videos into high-quality results with realistic facial details, stable identity, and temporal coherence. Recent diffusion-based methods have brought strong generative priors to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Zheng Chen , Bowen Chai , Rongjun Gao , Mingtao Nie , Xi Li , Bingnan Duan , Jianping Fang , Xiaohong Liu , Linghe Kong , Yulun Zhang

In video-based emotion recognition (ER), it is important to effectively leverage the complementary relationship among audio (A) and visual (V) modalities, while retaining the intra-modal characteristics of individual modalities. In this…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 R Gnana Praveen , Eric Granger , Patrick Cardinal

Recent advancements in video generation have demonstrated the potential of using video diffusion models as world models, with autoregressive generation of infinitely long videos through masked conditioning. However, such models, usually…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Taiye Chen , Zihan Ding , Anjian Li , Christina Zhang , Zeqi Xiao , Yisen Wang , Chi Jin

Exploiting similar and sharper scene patches in spatio-temporal neighborhoods is critical for video deblurring. However, CNN-based methods show limitations in capturing long-range dependencies and modeling non-local self-similarity. In this…

Image and Video Processing · Electrical Eng. & Systems 2022-05-31 Jing Lin , Yuanhao Cai , Xiaowan Hu , Haoqian Wang , Youliang Yan , Xueyi Zou , Henghui Ding , Yulun Zhang , Radu Timofte , Luc Van Gool

Existing state-of-the-art methods for surgical phase recognition either rely on the extraction of spatial-temporal features at a short-range temporal resolution or adopt the sequential extraction of the spatial and temporal features across…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Shu Yang , Luyang Luo , Qiong Wang , Hao Chen

Infrared and visible video fusion is essential for achieving comprehensive perception in dynamic scenes. However, maintaining temporal consistency remains a formidable challenge. Conventional methods relying on optical flow often suffer…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Xingyuan Li , Haoyuan Xu , Shulin Li , Xiang Chen , Zhiying Jiang , Jinyuan Liu

High resolution images can be acquired using a non-regular sampling sensor which consists of an underlying low resolution sensor that is covered with a non-regular sampling mask. The reconstructed high resolution image is then obtained…

Image and Video Processing · Electrical Eng. & Systems 2022-04-08 Markus Jonscher , Karina Jaskolka , Jürgen Seiler , André Kaup

Large multimodal models (LMMs) excel in scene understanding but struggle with fine-grained spatiotemporal reasoning due to weak alignment between linguistic and visual representations. Existing methods map textual positions and durations…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Hanyu Zhou , Gim Hee Lee

Due to storage and bandwidth limitations, videos transmitted over the Internet often exhibit low quality, characterized by low-resolution and compression artifacts. Although video super-resolution (VSR) is an efficient video enhancing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Hongyu An , Xinfeng Zhang , Shijie Zhao , Li Zhang , Ruiqin Xiong

Diffusion-based zero-shot image restoration and enhancement models have achieved great success in various tasks of image restoration and enhancement. However, directly applying them to video restoration and enhancement results in severe…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Cong Cao , Huanjing Yue , Xin Liu , Jingyu Yang

Several speech processing systems have demonstrated considerable performance improvements when deep complex neural networks (DCNN) are coupled with self-attention (SA) networks. However, the majority of DCNN-based studies on speech…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-24 Vinay Kothapally , John H. L. Hansen

The recovery of 3D human mesh from monocular images has significantly been developed in recent years. However, existing models usually ignore spatial and temporal information, which might lead to mesh and image misalignment and temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Wei Yao , Hongwen Zhang , Yunlian Sun , Jinhui Tang

Most video super-resolution methods super-resolve a single reference frame with the help of neighboring frames in a temporal sliding window. They are less efficient compared to the recurrent-based methods. In this work, we propose a novel…

Computer Vision and Pattern Recognition · Computer Science 2020-08-04 Takashi Isobe , Xu Jia , Shuhang Gu , Songjiang Li , Shengjin Wang , Qi Tian

Video-based person re-identification is a crucial task of matching video sequences of a person across multiple camera views. Generally, features directly extracted from a single frame suffer from occlusion, blur, illumination and posture…

Computer Vision and Pattern Recognition · Computer Science 2018-12-27 Yiheng Liu , Zhenxun Yuan , Wengang Zhou , Houqiang Li

Turbulence-degraded image frames are distorted by both turbulent deformations and space-time-varying blurs. To suppress these effects, we propose a multi-frame reconstruction scheme to recover a latent image from the observed image…

Computer Vision and Pattern Recognition · Computer Science 2017-12-12 Chun Pong Lau , Yu Hin Lai , Lok Ming Lui

In autonomous driving, 3D object detection is essential for accurate perception and reliable decision-making. However, object motion and ego-motion often induce cross-frame spatiotemporal inconsistencies in BEV-based detectors, leading to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Wenxuan Li , Qin Zou , Shoubing Chen , Chi Chen , Yingyi Yang , Shoubing Chen , Qingxiang Meng

Learning based video compression attracts increasing attention in the past few years. The previous hybrid coding approaches rely on pixel space operations to reduce spatial and temporal redundancy, which may suffer from inaccurate motion…

Image and Video Processing · Electrical Eng. & Systems 2021-08-24 Zhihao Hu , Guo Lu , Dong Xu