English
Related papers

Related papers: CAMS: CAnonicalized Manipulation Spaces for Catego…

200 papers

Generating rich and controllable motion is a pivotal challenge in video synthesis. We propose Boximator, a new approach for fine-grained motion control. Boximator introduces two constraint types: hard box and soft box. Users select objects…

Computer Vision and Pattern Recognition · Computer Science 2024-02-05 Jiawei Wang , Yuchen Zhang , Jiaxin Zou , Yan Zeng , Guoqiang Wei , Liping Yuan , Hang Li

Recent generative models can synthesize high-quality images, but they often fail to generate humans interacting with objects using their hands. This arises mostly from the model's misunderstanding of such interactions and the hardships of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Patrick Kwon , Chen Chen , Hanbyul Joo

Recent progress on physics-based character animation has shown impressive breakthroughs on human motion synthesis, through imitating motion capture data via deep reinforcement learning. However, results have mostly been demonstrated on…

Computer Vision and Pattern Recognition · Computer Science 2020-12-17 Yu-Wei Chao , Jimei Yang , Weifeng Chen , Jia Deng

We present a system for learning generalizable hand-object tracking controllers purely from synthetic data, without requiring any human demonstrations. Our approach makes two key contributions: (1) HOP, a Hand-Object Planner, which can…

Robotics · Computer Science 2025-12-23 Yinhuai Wang , Runyi Yu , Hok Wai Tsui , Xiaoyi Lin , Hui Zhang , Qihan Zhao , Ke Fan , Miao Li , Jie Song , Jingbo Wang , Qifeng Chen , Ping Tan

Accurate capture of human-object interaction from ubiquitous sensors like RGB cameras is important for applications in human understanding, gaming, and robot learning. However, inferring 4D interactions from a single RGB view is highly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Xianghui Xie , Bowen Wen , Yan Chang , Hesam Rabeti , Jiefeng Li , Ye Yuan , Gerard Pons-Moll , Stan Birchfield

Classification activation map (CAM), utilizing the classification structure to generate pixel-wise localization maps, is a crucial mechanism for weakly supervised object localization (WSOL). However, CAM directly uses the classifier trained…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Lei Zhu , Qian Chen , Lujia Jin , Yunfei You , Yanye Lu

We present Implicit Two Hands (Im2Hands), the first neural implicit representation of two interacting hands. Unlike existing methods on two-hand reconstruction that rely on a parametric hand model and/or low-resolution meshes, Im2Hands can…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Jihyun Lee , Minhyuk Sung , Honggyu Choi , Tae-Kyun Kim

We consider the feedback design for stabilizing a rigid body system by making and breaking multiple contacts with the environment without prespecifying the timing or the number of occurrence of the contacts. We model such a system as a…

Systems and Control · Computer Science 2019-05-16 Weiqiao Han , Russ Tedrake

We consider the task of object grasping with a prosthetic hand capable of multiple grasp types. In this setting, communicating the intended grasp type often requires a high user cognitive load which can be reduced adopting shared autonomy…

Bimanual object manipulation involves multiple visuo-haptic sensory feedbacks arising from the interaction with the environment that are managed from the central nervous system and consequently translated in motor commands. Kinematic…

Robotics · Computer Science 2022-11-23 Elisa Galofaro , Erika D'Antonio , Nicola Lotti , Fabrizio Patane' , Maura Casadio , Lorenzo Masia

The goal of this paper is to estimate the 6D pose and dimensions of unseen object instances in an RGB-D image. Contrary to "instance-level" 6D pose estimation tasks, our problem assumes that no exact object CAD models are available during…

Computer Vision and Pattern Recognition · Computer Science 2019-06-25 He Wang , Srinath Sridhar , Jingwei Huang , Julien Valentin , Shuran Song , Leonidas J. Guibas

We introduce Egocentric Object Manipulation Graphs (Ego-OMG) - a novel representation for activity modeling and anticipation of near future actions integrating three components: 1) semantic temporal structure of activities, 2) short-term…

Computer Vision and Pattern Recognition · Computer Science 2020-06-08 Eadom Dessalene , Michael Maynord , Chinmaya Devaraj , Cornelia Fermuller , Yiannis Aloimonos

Semantic image synthesis (SIS) aims to generate realistic images that match given semantic masks. Despite recent advances allowing high-quality results and precise spatial control, they require a massive semantic segmentation dataset for…

Computer Vision and Pattern Recognition · Computer Science 2024-02-23 Jungwoo Chae , Hyunin Cho , Sooyeon Go , Kyungmook Choi , Youngjung Uh

We introduce a novel approach that takes a single semantic mask as input to synthesize multi-view consistent color images of natural scenes, trained with a collection of single images from the Internet. Prior works on 3D-aware image…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Shangzan Zhang , Sida Peng , Tianrun Chen , Linzhan Mou , Haotong Lin , Kaicheng Yu , Yiyi Liao , Xiaowei Zhou

Task-relevant grasping is critical for industrial assembly, where downstream manipulation tasks constrain the set of valid grasps. Learning how to perform this task, however, is challenging, since task-relevant grasp labels are hard to…

Robotics · Computer Science 2022-03-01 Bowen Wen , Wenzhao Lian , Kostas Bekris , Stefan Schaal

Eye-in-hand cameras have shown promise in enabling greater sample efficiency and generalization in vision-based robotic manipulation. However, for robotic imitation, it is still expensive to have a human teleoperator collect large amounts…

Robotics · Computer Science 2023-07-13 Moo Jin Kim , Jiajun Wu , Chelsea Finn

We aim to infer 3D shape and pose of object from a single image and propose a learning-based approach that can train from unstructured image collections, supervised by only segmentation outputs from off-the-shelf recognition systems (i.e.…

Computer Vision and Pattern Recognition · Computer Science 2021-02-12 Yufei Ye , Shubham Tulsiani , Abhinav Gupta

Purpose: We propose a formal framework for the modeling and segmentation of minimally-invasive surgical tasks using a unified set of motion primitives (MPs) to enable more objective labeling and the aggregation of different datasets.…

Robotics · Computer Science 2023-05-16 Kay Hutchinson , Ian Reyes , Zongyu Li , Homa Alemzadeh

It is now possible to estimate 3D human pose from monocular images with off-the-shelf 3D pose estimators. However, many practical applications require fine-grained absolute pose information for which multi-view cues and camera calibration…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 James Tang , Shashwat Suri , Daniel Ajisafe , Bastian Wandt , Helge Rhodin

We concentrate on a novel human-centric image synthesis task, that is, given only one reference facial photograph, it is expected to generate specific individual images with diverse head positions, poses, facial expressions, and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Chao Liang , Fan Ma , Linchao Zhu , Yingying Deng , Yi Yang