English
Related papers

Related papers: V-JEPA 2: Self-Supervised Video Models Enable Unde…

200 papers

Invariance-based and generative methods have shown a conspicuous performance for 3D self-supervised representation learning (SSRL). However, the former relies on hand-crafted data augmentations that introduce bias not universally applicable…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Naiwen Hu , Haozhe Cheng , Yifan Xie , Shiqi Li , Jihua Zhu

Pretraining on large labeled datasets is a prerequisite to achieve good performance in many computer vision tasks like 2D object recognition, video classification etc. However, pretraining is not widely used for 3D recognition tasks where…

Computer Vision and Pattern Recognition · Computer Science 2021-01-08 Zaiwei Zhang , Rohit Girdhar , Armand Joulin , Ishan Misra

Inspired by the success of generative pretraining in natural language, we ask whether the same principles can yield strong self-supervised visual learners. Instead of training models to output features for downstream use, we train them to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Sihan Xu , Ziqiao Ma , Wenhao Chai , Xuweiyi Chen , Weiyang Jin , Joyce Chai , Saining Xie , Stella X. Yu

The emerging field of action prediction plays a vital role in various computer vision applications such as autonomous driving, activity analysis and human-computer interaction. Despite significant advancements, accurately predicting future…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Izzeddin Teeti , Rongali Sai Bhargav , Vivek Singh , Andrew Bradley , Biplab Banerjee , Fabio Cuzzolin

Autonomous robots frequently need to detect "interesting" scenes to decide on further exploration, or to decide which data to share for cooperation. These scenarios often require fast deployment with little or no training data. Prior work…

Robotics · Computer Science 2021-12-21 Chen Wang , Yuheng Qiu , Wenshan Wang , Yafei Hu , Seungchan Kim , Sebastian Scherer

Prediction, decision-making, and motion planning are essential for autonomous driving. In most contemporary works, they are considered as individual modules or combined into a multi-task learning paradigm with a shared backbone but separate…

Robotics · Computer Science 2023-10-17 Pengqin Wang , Meixin Zhu , Hongliang Lu , Hui Zhong , Xianda Chen , Shaojie Shen , Xuesong Wang , Yinhai Wang

Videos captured from multiple viewpoints can help in perceiving the 3D structure of the world and benefit computer vision tasks such as action recognition, tracking, etc. In this paper, we present a method for self-supervised learning from…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Ketul Shah , Robert Crandall , Jie Xu , Peng Zhou , Marian George , Mayank Bansal , Rama Chellappa

Recent successes in autoregressive (AR) generation models, such as the GPT series in natural language processing, have motivated efforts to replicate this success in visual tasks. Some works attempt to extend this approach to autonomous…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Xiaotao Hu , Wei Yin , Mingkai Jia , Junyuan Deng , Xiaoyang Guo , Qian Zhang , Xiaoxiao Long , Ping Tan

Many believe that the successes of deep learning on image understanding problems can be replicated in the realm of video understanding. However, due to the scale and temporal nature of video, the span of video understanding problems and the…

Computer Vision and Pattern Recognition · Computer Science 2021-10-05 Matthew Hutchinson , Vijay Gadepally

Enabling robots to autonomously perform hybrid motions in diverse environments can be beneficial for long-horizon tasks such as material handling, household chores, and work assistance. This requires extensive exploitation of intrinsic…

Robotics · Computer Science 2024-08-19 Jin Wang , Rui Dai , Weijie Wang , Luca Rossini , Francesco Ruscelli , Nikos Tsagarakis

Video generative models demonstrate great promise in robotics by serving as visual planners or as policy supervisors. When pretrained on internet-scale data, such video models intimately understand alignment with natural language, and can…

Machine Learning · Computer Science 2025-04-23 Calvin Luo , Zilai Zeng , Yilun Du , Chen Sun

Self-supervised learning aims to learn representations from the data itself without explicit manual supervision. Existing efforts ignore a crucial aspect of self-supervised learning - the ability to scale to large amount of data because…

Computer Vision and Pattern Recognition · Computer Science 2019-06-07 Priya Goyal , Dhruv Mahajan , Abhinav Gupta , Ishan Misra

Humans make extensive use of vision and touch as complementary senses, with vision providing global information about the scene and touch measuring local information during manipulation without suffering from occlusions. While prior work…

Robotics · Computer Science 2023-08-01 Justin Kerr , Huang Huang , Albert Wilcox , Ryan Hoque , Jeffrey Ichnowski , Roberto Calandra , Ken Goldberg

Human comprehension of a video stream is naturally broad: in a few instants, we are able to understand what is happening, the relevance and relationship of objects, and forecast what will follow in the near future, everything all at once.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Simone Alberto Peirone , Francesca Pistilli , Antonio Alliegro , Giuseppe Averta

Video generative models trained on expert demonstrations have been utilized as performant text-conditioned visual planners for solving robotic tasks. However, generalization to unseen tasks remains a challenge. Whereas improved…

Robotics · Computer Science 2026-03-12 Calvin Luo , Zilai Zeng , Mingxi Jia , Yilun Du , Chen Sun

In this paper we consider the problem of classifying fine-grained, multi-step activities (e.g., cooking different recipes, making disparate home improvements, creating various forms of arts and crafts) from long videos spanning up to…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Xudong Lin , Fabio Petroni , Gedas Bertasius , Marcus Rohrbach , Shih-Fu Chang , Lorenzo Torresani

Distilling knowledge from human demonstrations is a promising way for robots to learn and act. Existing methods, which often rely on coarsely-aligned video pairs, are typically constrained to learning global or task-level features. As a…

Robotics · Computer Science 2025-11-18 Sicheng Xie , Haidong Cao , Zejia Weng , Zhen Xing , Haoran Chen , Shiwei Shen , Jiaqi Leng , Zuxuan Wu , Yu-Gang Jiang

We present a novel method enabling robots to quickly learn to manipulate objects by leveraging a motion planner to generate "expert" training trajectories from a small amount of human-labeled data. In contrast to the traditional…

There is a large variation in the activities that humans perform in their everyday lives. We consider modeling these composite human activities which comprises multiple basic level actions in a completely unsupervised setting. Our model…

Computer Vision and Pattern Recognition · Computer Science 2016-03-14 Chenxia Wu , Jiemi Zhang , Ozan Sener , Bart Selman , Silvio Savarese , Ashutosh Saxena