English
Related papers

Related papers: Decoupling and Recoupling Spatiotemporal Represent…

200 papers

Planar grasp detection is one of the most fundamental tasks to robotic manipulation, and the recent progress of consumer-grade RGB-D sensors enables delivering more comprehensive features from both the texture and shape modalities. However,…

Robotics · Computer Science 2023-03-01 Ran Qin , Haoxiang Ma , Boyang Gao , Di Huang

Human motion prediction is challenging due to the complex spatiotemporal feature modeling. Among all methods, graph convolution networks (GCNs) are extensively utilized because of their superiority in explicit connection modeling. Within a…

Computer Vision and Pattern Recognition · Computer Science 2023-06-06 Jiajun Fu , Fuxing Yang , Yonghao Dang , Xiaoli Liu , Jianqin Yin

We all depend on mobility, and vehicular transportation affects the daily lives of most of us. Thus, the ability to forecast the state of traffic in a road network is an important functionality and a challenging task. Traffic data is often…

Machine Learning · Computer Science 2022-09-07 Zezhi Shao , Zhao Zhang , Wei Wei , Fei Wang , Yongjun Xu , Xin Cao , Christian S. Jensen

Despite the success in still image recognition, deep neural networks for spatiotemporal signal tasks (such as human action recognition in videos) still suffers from low efficacy and inefficiency over the past years. Recently, human experts…

Computer Vision and Pattern Recognition · Computer Science 2020-04-13 Yizhou Zhou , Xiaoyan Sun , Chong Luo , Zheng-Jun Zha , Wenjun Zeng

Spatiotemporal and motion features are two complementary and crucial information for video action recognition. Recent state-of-the-art methods adopt a 3D CNN stream to learn spatiotemporal features and another flow stream to learn motion…

Computer Vision and Pattern Recognition · Computer Science 2019-08-19 Boyuan Jiang , Mengmeng Wang , Weihao Gan , Wei Wu , Junjie Yan

Existing video Variational Autoencoders (VAEs) generally overlook the similarity between frame contents, leading to redundant latent modeling. In this paper, we propose decoupled VAE (DeCo-VAE) to achieve compact latent representation.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Xiangchen Yin , Jiahui Yuan , Zhangchi Hu , Wenzhang Sun , Jie Chen , Xiaozhen Qiao , Hao Li , Xiaoyan Sun

Hyperspectral imaging holds promises in surgical imaging by offering biological tissue differentiation capabilities with detailed information that is invisible to the naked eye. For intra-operative guidance, real-time spectral data capture…

Image and Video Processing · Electrical Eng. & Systems 2025-05-02 Peichao Li , Oscar MacCormac , Jonathan Shapey , Tom Vercauteren

Pedestrian detection is a critical task in robot perception. Multispectral modalities (visible light and thermal) can boost pedestrian detection performance by providing complementary visual information. Several gaps remain with…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Asiegbu Miracle Kanu-Asiegbu , Nitin Jotwani , Xiaoxiao Du

Spatially consistent long-horizon video generation aims to maintain temporal and spatial consistency along predefined camera trajectories. Existing methods mostly entangle memory modeling with video generation, leading to inconsistent…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Yanjun Guo , Zhengqiang Zhang , Pengfei Wang , Xinyue Liang , Zhiyuan Ma , Lei Zhang

Deformable 3D Gaussian Splatting (3D-GS) is limited by missing intermediate motion information due to the low temporal resolution of RGB cameras. To address this, we introduce the first approach combining event cameras, which capture…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Wenhao Xu , Wenming Weng , Yueyi Zhang , Ruikang Xu , Zhiwei Xiong

Inspired by the facts that retinal cells actually segregate the visual scene into different attributes (e.g., spatial details, temporal motion) for respective neuronal processing, we propose to first decompose the input video into…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Ming Lu , Tong Chen , Dandan Ding , Fengqing Zhu , Zhan Ma

Reconstructing dynamic 3D scenes from monocular video remains fundamentally challenging due to the need to jointly infer motion, structure, and appearance from limited observations. Existing dynamic scene reconstruction methods based on…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Jiahui Li , Shengeng Tang , Jingxuan He , Gang Huang , Zhangye Wang , Yantao Pan , Lechao Cheng

We present a method for temporally consistent motion segmentation from RGB-D videos assuming a piecewise rigid motion model. We formulate global energies over entire RGB-D sequences in terms of the segmentation of each frame into a number…

Computer Vision and Pattern Recognition · Computer Science 2016-08-17 Peter Bertholet , Alexandru-Eugen Ichim , Matthias Zwicker

Effectively modeling and utilizing spatiotemporal features from RGB and other modalities (\eg, depth, thermal, and event data, denoted as X) is the core of RGB-X tracker design. Existing methods often employ two parallel branches to…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 X. Feng , D. Zhang , S. Hu , X. Li , M. Wu , J. Zhang , X. Chen , K. Huang

We explore spatiotemporal data augmentation using video foundation models to diversify both camera viewpoints and scene dynamics. Unlike existing approaches based on simple geometric transforms or appearance perturbations, our method…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Jinfan Zhou , Lixin Luo , Sungmin Eum , Heesung Kwon , Jeong Joon Park

A popular and affordable option to provide room-scale human behaviour tracking is to rely on commodity RGB-D sensors %todo: such as the Kinect family of devices? as such devices offer body tracking capabilities at a reasonable price point.…

Human-Computer Interaction · Computer Science 2024-09-24 Adrien Coppens , Valérie Maquil

Expressive representation of pose sequences is crucial for accurate motion modeling in human motion prediction (HMP). While recent deep learning-based methods have shown promise in learning motion representations, these methods tend to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Jiexin Wang , Wenwen Qiang , Zhao Yang , Bing Su

Current video representations heavily rely on learning from manually annotated video datasets which are time-consuming and expensive to acquire. We observe videos are naturally accompanied by abundant text information such as YouTube titles…

Computer Vision and Pattern Recognition · Computer Science 2021-01-29 Tianhao Li , Limin Wang

Reconstructing interacting hands from monocular RGB data is a challenging task, as it involves many interfering factors, e.g. self- and mutual occlusion and similar textures. Previous works only leverage information from a single RGB image…

Computer Vision and Pattern Recognition · Computer Science 2024-01-08 Weichao Zhao , Hezhen Hu , Wengang Zhou , Li li , Houqiang Li

Functional magnetic resonance imaging produces high dimensional data, with a less then ideal number of labelled samples for brain decoding tasks (predicting brain states). In this study, we propose a new deep temporal convolutional neural…

Machine Learning · Computer Science 2015-01-13 Orhan Firat , Emre Aksan , Ilke Oztekin , Fatos T. Yarman Vural
‹ Prev 1 3 4 5 6 7 10 Next ›