中文
相关论文

相关论文: Learning Spatio-Temporal Feature Representations f…

200 篇论文

Appearance-based gaze estimation (AGE) has achieved remarkable performance in constrained settings, yet we reveal a significant generalization gap where existing AGE models often fail in practical, unconstrained scenarios, particularly…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Zhenhao Li , Zheng Liu , Seunghyun Lee , Amin Fadaeinejad , Yuanhao Yu

The target of space-time video super-resolution (STVSR) is to increase both the frame rate (also referred to as the temporal resolution) and the spatial resolution of a given video. Recent approaches solve STVSR using end-to-end deep neural…

计算机视觉与模式识别 · 计算机科学 2024-12-25 Zijie Yue , Miaojing Shi

This PhD. Thesis concerns the study and development of hierarchical representations for spatio-temporal visual attention modeling and understanding in video sequences. More specifically, we propose two computational models for visual…

计算机视觉与模式识别 · 计算机科学 2023-08-11 Miguel-Ángel Fernández-Torres

This paper addresses the challenging problem of estimating the general visual attention of people in images. Our proposed method is designed to work across multiple naturalistic social scenarios and provides a full picture of the subject's…

计算机视觉与模式识别 · 计算机科学 2018-07-30 Eunji Chong , Nataniel Ruiz , Yongxin Wang , Yun Zhang , Agata Rozga , James Rehg

Spatiotemporal graph convolutional networks (STGCNs) have emerged as a desirable model for skeleton-based human action recognition. Despite achieving state-of-the-art performance, there is a limited understanding of the representations…

图像与视频处理 · 电气工程与系统科学 2023-12-14 Pratyusha Das , Sarath Shekkizhar , Antonio Ortega

The objective of this work is human pose estimation in videos, where multiple frames are available. We investigate a ConvNet architecture that is able to benefit from temporal context by combining information across the multiple frames…

计算机视觉与模式识别 · 计算机科学 2015-11-10 Tomas Pfister , James Charles , Andrew Zisserman

Pre-trained on tremendous image-text pairs, vision-language models like CLIP have demonstrated promising zero-shot generalization across numerous image-based tasks. However, extending these capabilities to video tasks remains challenging…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Zichen Liu , Kunlun Xu , Bing Su , Xu Zou , Yuxin Peng , Jiahuan Zhou

Computer-aided pathology detection algorithms for video-based imaging modalities must accurately interpret complex spatiotemporal information by integrating findings across multiple frames. Current state-of-the-art methods operate by…

We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from most existing methods that only consider RGB images as…

计算机视觉与模式识别 · 计算机科学 2021-11-01 Yi-Wen Chen , Yi-Hsuan Tsai , Ming-Hsuan Yang

In this research, we present SLYKLatent, a novel approach for enhancing gaze estimation by addressing appearance instability challenges in datasets due to aleatoric uncertainties, covariant shifts, and test domain generalization. SLYKLatent…

计算机视觉与模式识别 · 计算机科学 2025-11-06 Samuel Adebayo , Joost C. Dessing , Seán McLoone

State-of-the-art transformer-based video instance segmentation (VIS) approaches typically utilize either single-scale spatio-temporal features or per-frame multi-scale features during the attention computations. We argue that such an…

Automatic generation of video captions is a fundamental challenge in computer vision. Recent techniques typically employ a combination of Convolutional Neural Networks (CNNs) and Recursive Neural Networks (RNNs) for video captioning. These…

计算机视觉与模式识别 · 计算机科学 2019-04-30 Nayyer Aafaq , Naveed Akhtar , Wei Liu , Syed Zulqarnain Gilani , Ajmal Mian

Gait recognition, which refers to the recognition or identification of a person based on their body shape and walking styles, derived from video data captured from a distance, is widely used in crime prevention, forensic identification, and…

计算机视觉与模式识别 · 计算机科学 2022-07-14 Hung-Min Hsu , Yizhou Wang , Cheng-Yen Yang , Jenq-Neng Hwang , Hoang Le Uyen Thuc , Kwang-Ju Kim

Facial expression spotting is a significant but challenging task in facial expression analysis. The accuracy of expression spotting is affected not only by irrelevant facial movements but also by the difficulty of perceiving subtle motions…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Yicheng Deng , Hideaki Hayashi , Hajime Nagahara

Enabling robots to understand human gaze target is a crucial step to allow capabilities in downstream tasks, for example, attention estimation and movement anticipation in real-world human-robot interactions. Prior works have addressed the…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Zhuangzhuang Dai , Vincent Gbouna Zakka , Luis J. Manso , Chen Li

In this paper, we investigate the challenge of spatio-temporal video prediction task, which involves generating future video frames based on historical spatio-temporal observation streams. Existing approaches typically utilize external…

计算机视觉与模式识别 · 计算机科学 2025-01-15 Hao Wu , Fan Xu , Chong Chen , Xian-Sheng Hua , Xiao Luo , Haixin Wang

We consider the problem of video-based person re-identification. The goal is to identify a person from videos captured under different cameras. In this paper, we propose an efficient spatial-temporal attention based model for person…

计算机视觉与模式识别 · 计算机科学 2018-10-29 Shivansh Rao , Tanzila Rahman , Mrigank Rochan , Yang Wang

In IoT based distributed network of cameras, real-time multi-camera video analytics is challenged by high bandwidth demands and redundant visual data, creating a fundamental tension where reducing data saves network overhead but can degrade…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Ragini Gupta , Lingzhi Zhao , Jiaxi Li , Volodymyr Vakhniuk , Claudiu Danilov , Josh Eckhardt , Keyshla Bernard , Klara Nahrstedt

Video-based person re-identification (Re-ID) aims at matching video sequences of pedestrians across non-overlapping cameras. It is a practical yet challenging task of how to embed spatial and temporal information of a video into its feature…

计算机视觉与模式识别 · 计算机科学 2019-08-06 Chih-Ting Liu , Chih-Wei Wu , Yu-Chiang Frank Wang , Shao-Yi Chien

Inter-personal anatomical differences limit the accuracy of person-independent gaze estimation networks. Yet there is a need to lower gaze errors further to enable applications requiring higher quality. Further gains can be achieved by…

计算机视觉与模式识别 · 计算机科学 2019-10-15 Seonwook Park , Shalini De Mello , Pavlo Molchanov , Umar Iqbal , Otmar Hilliges , Jan Kautz