English
Related papers

Related papers: Bridging the Embodiment Gap: Disentangled Cross-Em…

200 papers

This paper focuses on transferring control policies between robot manipulators with different morphology. While reinforcement learning (RL) methods have shown successful results in robot manipulation tasks, transferring a trained policy…

Robotics · Computer Science 2024-06-05 Tianyu Wang , Dwait Bhatt , Xiaolong Wang , Nikolay Atanasov

A video autoencoder is proposed for learning disentan- gled representations of 3D structure and camera pose from videos in a self-supervised manner. Relying on temporal continuity in videos, our work assumes that the 3D scene structure in…

Computer Vision and Pattern Recognition · Computer Science 2021-10-07 Zihang Lai , Sifei Liu , Alexei A. Efros , Xiaolong Wang

There are many forms of feature information present in video data. Principle among them are object identity information which is largely static across multiple video frames, and object pose and style information which continuously…

Computer Vision and Pattern Recognition · Computer Science 2016-12-20 Will Grathwohl , Aaron Wilson

Learning robot manipulation from abundant human videos offers a scalable alternative to costly robot-specific data collection. However, domain gaps across visual, morphological, and physical aspects hinder direct imitation. To effectively…

Robotics · Computer Science 2025-09-16 Yangcen Liu , Woo Chul Shin , Yunhai Han , Zhenyang Chen , Harish Ravichandar , Danfei Xu

Humans are able to seamlessly visually imitate others, by inferring their intentions and using past experience to achieve the same end goal. In other words, we can parse complex semantic knowledge from raw video and efficiently translate…

Machine Learning · Computer Science 2020-11-12 Sudeep Dasari , Abhinav Gupta

As an increasing amount of image and video content will be analyzed by machines, there is demand for a new codec paradigm that is capable of compressing visual input primarily for the purpose of computer vision inference, while secondarily…

Image and Video Processing · Electrical Eng. & Systems 2023-01-12 Ezgi Ozyilkan , Mateen Ulhaq , Hyomin Choi , Fabien Racape

This paper presents a method for generating large-scale datasets to improve class-agnostic video segmentation across robots with different form factors. Specifically, we consider the question of whether video segmentation models trained on…

We study video-specific autoencoders that allow a human user to explore, edit, and efficiently transmit videos. Prior work has independently looked at these problems (and sub-problems) and proposed different formulations. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2022-01-11 Kevin Wang , Deva Ramanan , Aayush Bansal

Embodied control requires agents to leverage multi-modal pre-training to quickly learn how to act in new environments, where video demonstrations contain visual and motion details needed for low-level perception and control, and language…

Machine Learning · Computer Science 2023-04-20 Yao Mu , Shunyu Yao , Mingyu Ding , Ping Luo , Chuang Gan

Image-to-image translation aims to learn the mapping between two visual domains. There are two main challenges for many applications: 1) the lack of aligned training pairs and 2) multiple possible outputs from a single input image. In this…

Computer Vision and Pattern Recognition · Computer Science 2018-08-03 Hsin-Ying Lee , Hung-Yu Tseng , Jia-Bin Huang , Maneesh Kumar Singh , Ming-Hsuan Yang

For visual manipulation tasks, we aim to represent image content with semantically meaningful features. However, learning implicit representations from images often lacks interpretability, especially when attributes are intertwined. We…

Computer Vision and Pattern Recognition · Computer Science 2022-08-08 Xue Hu , Xinghui Li , Benjamin Busam , Yiren Zhou , Ales Leonardis , Shanxin Yuan

We propose a self-supervised approach for learning representations and robotic behaviors entirely from unlabeled videos recorded from multiple viewpoints, and study how this representation can be used in two robotic imitation settings:…

Computer Vision and Pattern Recognition · Computer Science 2018-03-21 Pierre Sermanet , Corey Lynch , Yevgen Chebotar , Jasmine Hsu , Eric Jang , Stefan Schaal , Sergey Levine

Observing a human demonstrator manipulate objects provides a rich, scalable and inexpensive source of data for learning robotic policies. However, transferring skills from human videos to a robotic manipulator poses several challenges, not…

Robotics · Computer Science 2023-03-08 Minttu Alakuijala , Gabriel Dulac-Arnold , Julien Mairal , Jean Ponce , Cordelia Schmid

Distilling knowledge from human demonstrations is a promising way for robots to learn and act. Existing methods, which often rely on coarsely-aligned video pairs, are typically constrained to learning global or task-level features. As a…

Robotics · Computer Science 2025-11-18 Sicheng Xie , Haidong Cao , Zejia Weng , Zhen Xing , Haoran Chen , Shiwei Shen , Jiaqi Leng , Zuxuan Wu , Yu-Gang Jiang

Recent advancements in human video synthesis have enabled the generation of high-quality videos through the application of stable diffusion models. However, existing methods predominantly concentrate on animating solely the human element…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Jinlin Liu , Kai Yu , Mengyang Feng , Xiefan Guo , Miaomiao Cui

Large-scale multi-task robotic manipulation systems often rely on text to specify the task. In this work, we explore whether a robot can learn by observing humans. To do so, the robot must understand a person's intent and perform the…

This paper introduces a novel deep-learning approach for human-to-robot motion retargeting, enabling robots to mimic human poses accurately. Contrary to prior deep-learning-based works, our method does not require paired human-to-robot…

Robotics · Computer Science 2024-04-09 Yashuai Yan , Esteve Valls Mascaro , Dongheui Lee

World models have made significant progress in modeling dynamic environments; however, most embodied world models are still restricted to 2D representations, lacking the comprehensive multi-view information essential for embodied spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Peiyan Tu , Hanxin Zhu , Jingwen Sun , Shaojie Ren , Cong Wang , Jiayi Luo , Xiaoqian Cheng , Zhibo Chen

By generating plausible and smooth transitions between two image frames, video inbetweening is an essential tool for video editing and long video synthesis. Traditional works lack the capability to generate complex large motions. While…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Maham Tanveer , Yang Zhou , Simon Niklaus , Ali Mahdavi Amiri , Hao Zhang , Krishna Kumar Singh , Nanxuan Zhao

We study the problem of learning disentangled representations for data across multiple domains and its applications in human retargeting. Our goal is to map an input image to an identity-invariant latent representation that captures…

Computer Vision and Pattern Recognition · Computer Science 2019-12-16 Chao Yang , Xiaofeng Liu , Qingming Tang , C. -C. Jay Kuo