English
Related papers

Related papers: End-to-End Spatial-Temporal Transformer for Real-t…

200 papers

We present Human to Humanoid (H2O), a reinforcement learning (RL) based framework that enables real-time whole-body teleoperation of a full-sized humanoid robot with only an RGB camera. To create a large-scale retargeted motion dataset of…

Robotics · Computer Science 2024-03-08 Tairan He , Zhengyi Luo , Wenli Xiao , Chong Zhang , Kris Kitani , Changliu Liu , Guanya Shi

The current methods of video-based 3D human pose estimation have achieved significant progress.However, they still face pressing challenges, such as the underutilization of spatiotemporal bodystructure features in transformers and the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Yang Liu , Zhiyong Zhang

Generating realistic hand-object interactions (HOI) videos is a significant challenge due to the difficulty of modeling physical constraints (e.g., contact and occlusion between hands and manipulated objects). Current methods utilize HOI…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Haodong Yan , Hang Yu , Zhide Zhong , Weilin Yuan , Xin Gong , Zehang Luo , Chengxi Heyu , Junfeng Li , Wenxuan Song , Shunbo Zhou , Haoang Li

We present a new method, called MEsh TRansfOrmer (METRO), to reconstruct 3D human pose and mesh vertices from a single image. Our method uses a transformer encoder to jointly model vertex-vertex and vertex-joint interactions, and outputs 3D…

Computer Vision and Pattern Recognition · Computer Science 2021-06-16 Kevin Lin , Lijuan Wang , Zicheng Liu

We present a new pipeline for acquiring a textured mesh in the wild with a single smartphone which offers access to images, depth maps, and valid poses. Our method first introduces an RGBD-aided structure from motion, which can yield…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Jaehoon Choi , Dongki Jung , Taejae Lee , Sangwook Kim , Youngdong Jung , Dinesh Manocha , Donghwan Lee

Image retargeting aims to change the aspect-ratio of an image while maintaining its content and structure with less visual artifacts. Existing methods still generate many artifacts or fail to maintain original content or structure. To…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Yiran Xu , Siqi Xie , Zhuofang Li , Harris Shadmany , Yinxiao Li , Luciano Sbaiz , Miaosen Wang , Junjie Ke , Jose Lezama , Hang Qi , Han Zhang , Jesse Berent , Ming-Hsuan Yang , Irfan Essa , Jia-Bin Huang , Feng Yang

This paper introduces a novel Pre-trained Spatial Temporal Many-to-One (P-STMO) model for 2D-to-3D human pose estimation task. To reduce the difficulty of capturing spatial and temporal information, we divide this task into two stages:…

Computer Vision and Pattern Recognition · Computer Science 2022-08-01 Wenkang Shan , Zhenhua Liu , Xinfeng Zhang , Shanshe Wang , Siwei Ma , Wen Gao

Space situational awareness demands efficient monitoring of terrestrial sites and celestial bodies, necessitating advanced target recognition systems. Current target recognition systems exhibit limited operational speed due to challenges in…

Image and Video Processing · Electrical Eng. & Systems 2023-12-13 Julian Gamboa , Xi Shen , Tabassom Hamidfar , Selim M. Shahriar

Purpose: To develop new encoding and reconstruction techniques for fast multi-contrast quantitative imaging. Methods: The recently proposed Echo Planar Time-resolved Imaging (EPTI) technique can achieve fast distortion- and blurring-free…

Image and Video Processing · Electrical Eng. & Systems 2020-10-05 Zijing Dong , Fuyixue Wang , Timothy G. Reese , Berkin Bilgic , Kawin Setsompop

Spatio-temporal coherency is a major challenge in synthesizing high quality videos, particularly in synthesizing human videos that contain rich global and local deformations. To resolve this challenge, previous approaches have resorted to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Yaohui Wang , Xin Ma , Xinyuan Chen , Cunjian Chen , Antitza Dantcheva , Bo Dai , Yu Qiao

Recent advances in transformer-based text-to-motion generation have led to impressive progress in synthesizing high-quality human motion. Nevertheless, jointly achieving high fidelity, streaming capability, real-time responsiveness, and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Dongjie Fu , Tengjiao Sun , Pengcheng Fang , Xiaohao Cai , Hansung Kim

We present DreamHOI, a novel method for zero-shot synthesis of human-object interactions (HOIs), enabling a 3D human model to realistically interact with any given object based on a textual description. This task is complicated by the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Thomas Hanwen Zhu , Ruining Li , Tomas Jakab

We introduce D3D-HOI: a dataset of monocular videos with ground truth annotations of 3D object pose, shape and part motion during human-object interactions. Our dataset consists of several common articulated objects captured from diverse…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Xiang Xu , Hanbyul Joo , Greg Mori , Manolis Savva

Modeling 3D human-object interaction (HOI) is a problem of great interest for computer vision and a key enabler for virtual and mixed-reality applications. Existing methods work in a one-way direction: some recover plausible human…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Ilya A. Petrov , Riccardo Marin , Julian Chibane , Gerard Pons-Moll

We present an approach to reconstruct humans and track them over time. At the core of our approach, we propose a fully "transformerized" version of a network for human mesh recovery. This network, HMR 2.0, advances the state of the art and…

Computer Vision and Pattern Recognition · Computer Science 2023-09-01 Shubham Goel , Georgios Pavlakos , Jathushan Rajasegaran , Angjoo Kanazawa , Jitendra Malik

Neuromorphic cameras, also known as event cameras, are asynchronous brightness-change sensors that can capture extremely fast motion without suffering from motion blur, making them particularly promising for 3D reconstruction in extreme…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Chuanzhi Xu , Langyi Chen , Haodong Chen , Vera Chung , Qiang Qu

Inspired by the performance and scalability of autoregressive large language models (LLMs), transformer-based models have seen recent success in the visual domain. This study investigates a transformer adaptation for video prediction with a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Dean L Slack , G Thomas Hudson , Thomas Winterbottom , Noura Al Moubayed

Reasoning the human-object interactions (HOI) is essential for deeper scene understanding, while object affordances (or functionalities) are of great importance for human to discover unseen HOIs with novel objects. Inspired by this, we…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Zhi Hou , Baosheng Yu , Yu Qiao , Xiaojiang Peng , Dacheng Tao

Reconstructing 3D clothed humans from monocular images and videos is a fundamental problem with applications in virtual try-on, avatar creation, and mixed reality. Despite significant progress in human body recovery, accurately…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Yingxuan You , Ren Li , Corentin Dumery , Cong Cao , Hao Li , Pascal Fua

Video-based human pose estimation models aim to address scenarios that cannot be effectively solved by static image models such as motion blur, out-of-focus and occlusion. Most existing approaches consist of two stages: detecting human…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Zhihong Wei