English
Related papers

Related papers: Exploiting Spatial-Temporal Context for Interactin…

200 papers

People often interact with their surroundings by applying pressure with their hands. While hand pressure can be measured by placing pressure sensors between the hand and the environment, doing so can alter contact mechanics, interfere with…

Computer Vision and Pattern Recognition · Computer Science 2022-10-04 Patrick Grady , Chengcheng Tang , Samarth Brahmbhatt , Christopher D. Twigg , Chengde Wan , James Hays , Charles C. Kemp

Recent progress of video diffusion models have enabled extensive simulation of the physical world. While simulation with hand object interaction has been less explored. We propose DexSIM, a dexterous simulation framework for simulating…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Adam Lee

Spatio-temporal feature learning is of central importance for action recognition in videos. Existing deep neural network models either learn spatial and temporal features independently (C2D) or jointly with unconstrained parameters (C3D).…

Computer Vision and Pattern Recognition · Computer Science 2019-03-05 Chao Li , Qiaoyong Zhong , Di Xie , Shiliang Pu

Our work aims to reconstruct hand-object interactions from a single-view image, which is a fundamental but ill-posed task. Unlike methods that reconstruct from videos, multi-view images, or predefined 3D templates, single-view…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Yumeng Liu , Xiaoxiao Long , Zemin Yang , Yuan Liu , Marc Habermann , Christian Theobalt , Yuexin Ma , Wenping Wang

We present the first method for real-time full body capture that estimates shape and motion of body and hands together with a dynamic 3D face model from a single color image. Our approach uses a new neural network architecture that exploits…

Computer Vision and Pattern Recognition · Computer Science 2021-04-16 Yuxiao Zhou , Marc Habermann , Ikhsanul Habibie , Ayush Tewari , Christian Theobalt , Feng Xu

Egocentric action recognition is a challenging task due to erratic camera motion, frequent hand occlusion, and the difficulty of maintaining consistent visual representations over time. In this work, we propose a cross-modal architecture…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Juan Ignacio Bustos Gorostegui , Maria Elena Buemi

Video-based representations have gained prominence in planning and decision-making due to their ability to encode rich spatiotemporal dynamics and geometric relationships. These representations enable flexible and generalizable solutions…

Robotics · Computer Science 2026-02-11 Po-Chen Ko , Jiayuan Mao , Yu-Hsiang Fu , Hsien-Jeng Yeh , Chu-Rong Chen , Wei-Chiu Ma , Yilun Du , Shao-Hua Sun

Reconstructing spatially and temporally coherent videos from time-varying measurements is a fundamental challenge in many scientific domains. A major difficulty arises from the sparsity of measurements, which hinders accurate recovery of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Bingliang Zhang , Zihui Wu , Berthy T. Feng , Yang Song , Yisong Yue , Katherine L. Bouman

Humans effortlessly retrieve objects in cluttered, partially observable environments by combining visual reasoning, active viewpoint adjustment, and physical interaction-with only a single pair of eyes. In contrast, most existing robotic…

Robotics · Computer Science 2025-08-19 Hecheng Wang , Jiankun Ren , Jia Yu , Lizhe Qi , Yunquan Sun

We revisit the role of texture in monocular 3D hand reconstruction, not as an afterthought for photorealism, but as a dense, spatially grounded cue that can actively support pose and shape estimation. Our observation is simple: even in…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Giorgos Karvounas , Nikolaos Kyriazis , Iason Oikonomidis , Georgios Pavlakos , Antonis A. Argyros

Recent years have witnessed great success for hand reconstruction in real-time applications such as visual reality and augmented reality while interacting with two-hand reconstruction through efficient transformers is left unexplored. In…

Computer Vision and Pattern Recognition · Computer Science 2022-08-30 Xinhan Di , Pengqian Yu

This paper addresses the problem of text-to-video temporal grounding, which aims to identify the time interval in a video semantically relevant to a text query. We tackle this problem using a novel regression-based model that learns to…

Computer Vision and Pattern Recognition · Computer Science 2020-04-17 Jonghwan Mun , Minsu Cho , Bohyung Han

Monocular vertex-level human-scene contact prediction is a fundamental capability for interactive systems such as assistive monitoring, embodied AI, and rehabilitation analysis. In this work, we study this task jointly with single-image 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Xiaojian Lin , Yaomin Shen , Junyuan Ma , Yujie Sun , Chengqing Bu , Wenxin Zhang , Zongzheng Zhang , Hao Fei , Lei Jin , Hao Zhao

Egocentric interactive world models are essential for augmented reality and embodied AI, where visual generation must respond to user input with low latency, geometric consistency, and long-term stability. We study egocentric interaction…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Yuxi Wang , Wenqi Ouyang , Tianyi Wei , Yi Dong , Zhiqi Shen , Xingang Pan

This paper addresses the task of segmenting class-agnostic objects in semi-supervised setting. Although previous detection based methods achieve relatively good performance, these approaches extract the best proposal by a greedy strategy,…

Computer Vision and Pattern Recognition · Computer Science 2020-12-11 Daizong Liu , Shuangjie Xu , Xiao-Yang Liu , Zichuan Xu , Wei Wei , Pan Zhou

Existing multi-person human reconstruction approaches mainly focus on recovering accurate poses or avoiding penetration, but overlook the modeling of close interactions. In this work, we tackle the task of reconstructing closely interactive…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Buzhen Huang , Chen Li , Chongyang Xu , Liang Pan , Yangang Wang , Gim Hee Lee

We propose SelfRecon, a clothed human body reconstruction method that combines implicit and explicit representations to recover space-time coherent geometries from a monocular self-rotating human video. Explicit methods require a predefined…

Computer Vision and Pattern Recognition · Computer Science 2022-04-06 Boyi Jiang , Yang Hong , Hujun Bao , Juyong Zhang

Existing hand-object interactions (HOI) methods are largely limited to rigid objects, while 4D reconstruction methods of articulated objects generally require pre-scanning the object or even multi-view videos. It remains an unexplored but…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Zikai Wang , Zhilu Zhang , Yiqing Wang , Hui Li , Wangmeng Zuo

While supervised techniques in re-identification are extremely effective, the need for large amounts of annotations makes them impractical for large camera networks. One-shot re-identification, which uses a singular labeled tracklet for…

Computer Vision and Pattern Recognition · Computer Science 2020-07-23 Dripta S. Raychaudhuri , Amit K. Roy-Chowdhury

We propose Dyn-HaMR, to the best of our knowledge, the first approach to reconstruct 4D global hand motion from monocular videos recorded by dynamic cameras in the wild. Reconstructing accurate 3D hand meshes from monocular videos is a…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Zhengdi Yu , Stefanos Zafeiriou , Tolga Birdal