English
Related papers

Related papers: Exploiting Spatial-Temporal Context for Interactin…

200 papers

This paper presents a new method to describe spatio-temporal relations between objects and hands, to recognize both interactions and activities within video demonstrations of manual tasks. The approach exploits Scene Graphs to extract key…

Computer Vision and Pattern Recognition · Computer Science 2023-07-10 Elena Merlo , Marta Lagomarsino , Edoardo Lamon , Arash Ajoudani

We present GASPACHO, a method for generating photorealistic, controllable renderings of human-object interactions from multi-view RGB video. Unlike prior work that reconstructs only the human and treats objects as background, GASPACHO…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Aymen Mir , Arthur Moreau , Helisa Dhamo , Zhensong Zhang , Gerard Pons-Moll , Eduardo Pérez-Pellitero

3D hand pose estimation and shape recovery are challenging tasks in computer vision. We introduce a novel framework HandTailor, which combines a learning-based hand module and an optimization-based tailor module to achieve high-precision…

Computer Vision and Pattern Recognition · Computer Science 2021-10-25 Jun Lv , Wenqiang Xu , Lixin Yang , Sucheng Qian , Chongzhao Mao , Cewu Lu

A consistent spatial-temporal coordination across multiple agents is fundamental for collaborative perception, which seeks to improve perception abilities through information exchange among agents. To achieve this spatial-temporal…

Artificial Intelligence · Computer Science 2024-06-03 Zixing Lei , Zhenyang Ni , Ruize Han , Shuo Tang , Dingju Wang , Chen Feng , Siheng Chen , Yanfeng Wang

Prevailing High Dynamic Range (HDR) video reconstruction methods are fundamentally trapped in a fragile alignment-and-fusion paradigm. While explicit spatial alignment can successfully recover fine details in controlled environments, it…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Qianyu Zhang , Bolun Zheng , Lingyu Zhu , Aiai Huang , Zongpeng Li , Shiqi Wang

Human activity recognition in videos has been widely studied and has recently gained significant advances with deep learning approaches; however, it remains a challenging task. In this paper, we propose a novel framework that simultaneously…

Computer Vision and Pattern Recognition · Computer Science 2021-01-25 Dong-Gyu Lee , Seong-Whan Lee

Reconstructing dynamic hand-object interactions from monocular videos is critical for dexterous manipulation data collection and creating realistic digital twins for robotics and VR. However, current methods face two prohibitive barriers:…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Jin-Chuan Shi , Binhong Ye , Tao Liu , Xiaoyang Liu , Yangjinhui Xu , Junzhe He , Zeju Li , Hao Chen , Chunhua Shen

Estimating 3D interacting hand pose from a single RGB image is essential for understanding human actions. Unlike most previous works that directly predict the 3D poses of two interacting hands simultaneously, we propose to decompose the…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Hao Meng , Sheng Jin , Wentao Liu , Chen Qian , Mengxiang Lin , Wanli Ouyang , Ping Luo

Human-object interactions with articulated objects are common in everyday life. Despite much progress in single-view 3D reconstruction, it is still challenging to infer an articulated 3D object model from an RGB video showing a person…

Computer Vision and Pattern Recognition · Computer Science 2022-09-14 Sanjay Haresh , Xiaohao Sun , Hanxiao Jiang , Angel X. Chang , Manolis Savva

This paper focuses on a challenging setting of simultaneously modeling geometry and appearance of hand-object interaction scenes without any object priors. We follow the trend of dynamic 3D Gaussian Splatting based methods, and address…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Hao Tian , Chenyangguang Zhang , Rui Liu , Wen Shen , Xiaolin Qin

Tactile sensing is one of the modalities humans rely on heavily to perceive the world. Working with vision, this modality refines local geometry structure, measures deformation at the contact area, and indicates the hand-object contact…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Wenqiang Xu , Zhenjun Yu , Han Xue , Ruolin Ye , Siqiong Yao , Cewu Lu

We ask whether everyday open-world monocular videos can be turned into reusable 4D interaction primitives: articulated hand motion, object shape with 6D pose over time, and the when/where of contact. Such a capability would enable scalable…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Hao Xu , Yilin Liu , Yinqiao Wang , Chi-Wing Fu , Niloy J. Mitra

This paper presents an approach that reconstructs a hand-held object from a monocular video. In contrast to many recent methods that directly predict object geometry by a trained network, the proposed approach does not require any learned…

Computer Vision and Pattern Recognition · Computer Science 2022-12-01 Di Huang , Xiaopeng Ji , Xingyi He , Jiaming Sun , Tong He , Qing Shuai , Wanli Ouyang , Xiaowei Zhou

Robust video scene classification models should capture the spatial (pixel-wise) and temporal (frame-wise) characteristics of a video effectively. Transformer models with self-attention which are designed to get contextualized…

Computer Vision and Pattern Recognition · Computer Science 2021-10-28 Saurabh Sahu , Palash Goyal

We propose a self-supervised learning method to jointly reason about spatial and temporal context for video recognition. Recent self-supervised approaches have used spatial context [9, 34] as well as temporal coherency [32] but a…

Computer Vision and Pattern Recognition · Computer Science 2018-08-24 Unaiza Ahsan , Rishi Madhok , Irfan Essa

Recent advances have enabled 3d object reconstruction approaches using a single off-the-shelf RGB-D camera. Although these approaches are successful for a wide range of object classes, they rely on stable and distinctive geometric or…

Computer Vision and Pattern Recognition · Computer Science 2017-04-04 Dimitrios Tzionas , Juergen Gall

Capturing challenging human motions is critical for numerous applications, but it suffers from complex motion patterns and severe self-occlusion under the monocular setting. In this paper, we propose ChallenCap -- a template-based approach…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Yannan He , Anqi Pang , Xin Chen , Han Liang , Minye Wu , Yuexin Ma , Lan Xu

This work presents a first evaluation of using spatio-temporal receptive fields from a recently proposed time-causal spatio-temporal scale-space framework as primitives for video analysis. We propose a new family of video descriptors based…

Computer Vision and Pattern Recognition · Computer Science 2021-05-20 Ylva Jansson , Tony Lindeberg

Monocular 4D human-object interaction (HOI) reconstruction - recovering a moving human and a manipulated object from a single RGB video - remains challenging due to depth ambiguity and frequent occlusions. Existing methods often rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Haoyu Zhang , Wei Zhai , Yuhang Yang , Yang Cao , Zheng-Jun Zha

In domains such as healthcare, finance, and e-commerce, the temporal dynamics of relational data emerge from complex interactions-such as those between patients and providers, or users and products across diverse categories. To be broadly…

Machine Learning · Computer Science 2025-11-07 Divyansha Lachi , Mahmoud Mohammadi , Joe Meyer , Vinam Arora , Tom Palczewski , Eva L. Dyer
‹ Prev 1 4 5 6 7 8 10 Next ›