English
Related papers

Related papers: iPOKE: Poking a Still Image for Controlled Stochas…

200 papers

We present a system for learning motion of independently moving objects from stereo videos. The only human annotation used in our system are 2D object bounding boxes which introduce the notion of objects to our system. Unlike prior learning…

Computer Vision and Pattern Recognition · Computer Science 2019-01-09 Zhe Cao , Abhishek Kar , Christian Haene , Jitendra Malik

Video stabilization is essential for improving visual quality of shaky videos. The current video stabilization methods usually take feature trajectories in the background to estimate one global transformation matrix or several…

Computer Vision and Pattern Recognition · Computer Science 2020-06-16 Minda Zhao , Qiang Ling

This paper strives for motion-focused video-language representations. Existing methods to learn video-language representations use spatial-focused data, where identifying the objects and scene is often enough to distinguish the relevant…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Hazel Doughty , Fida Mohammad Thoker , Cees G. M. Snoek

Human motion synthesis is an important problem with applications in graphics, gaming and simulation environments for robotics. Existing methods require accurate motion capture data for training, which is costly to obtain. Instead, we…

Computer Vision and Pattern Recognition · Computer Science 2022-08-15 Kevin Xie , Tingwu Wang , Umar Iqbal , Yunrong Guo , Sanja Fidler , Florian Shkurti

Human motion prediction is a stochastic process: Given an observed sequence of poses, multiple future motions are plausible. Existing approaches to modeling this stochasticity typically combine a random noise vector with information about…

Efficient video tokenization remains a key bottleneck in learning general purpose vision models that are capable of processing long video sequences. Prevailing approaches are restricted to encoding videos to a fixed number of tokens, where…

Machine Learning · Computer Science 2025-02-04 Wilson Yan , Volodymyr Mnih , Aleksandra Faust , Matei Zaharia , Pieter Abbeel , Hao Liu

Articulation modeling enables robots to learn joint parameters of articulated objects for effective manipulation which can then be used downstream for skill learning or planning. Existing approaches often rely on prior knowledge about the…

Robotics · Computer Science 2026-02-04 Anmol Gupta , Weiwei Gu , Omkar Patil , Jun Ki Lee , Nakul Gopalan

Masked Image Modeling (MIM) is a promising self-supervised learning approach that enables learning from unlabeled images. Despite its recent success, learning good representations through MIM remains challenging because it requires…

Computer Vision and Pattern Recognition · Computer Science 2024-02-28 Amir Bar , Florian Bordes , Assaf Shocher , Mahmoud Assran , Pascal Vincent , Nicolas Ballas , Trevor Darrell , Amir Globerson , Yann LeCun

Estimating 3D humans from images often produces implausible bodies that lean, float, or penetrate the floor. Such methods ignore the fact that bodies are typically supported by the scene. A physics engine can be used to enforce physical…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Shashank Tripathi , Lea Müller , Chun-Hao P. Huang , Omid Taheri , Michael J. Black , Dimitrios Tzionas

In this paper, we study the challenging problem of predicting the dynamics of objects in static images. Given a query object in an image, our goal is to provide a physical understanding of the object in terms of the forces acting upon it…

Computer Vision and Pattern Recognition · Computer Science 2015-11-13 Roozbeh Mottaghi , Hessam Bagherinezhad , Mohammad Rastegari , Ali Farhadi

When humans grasp objects in the real world, we often move our arms to hold the object in a different pose where we can use it. In contrast, typical lab settings only study the stability of the grasp immediately after lifting, without any…

Robotics · Computer Science 2022-09-13 Shubham Kanitkar , Helen Jiang , Wenzhen Yuan

Objects in videos are typically characterized by continuous smooth motion. We exploit continuous smooth motion in three ways. 1) Improved accuracy by using object motion as an additional source of supervision, which we obtain by…

Computer Vision and Pattern Recognition · Computer Science 2023-08-10 Xin Liu , Fatemeh Karimi Nejadasl , Jan C. van Gemert , Olaf Booij , Silvia L. Pintea

Despite the progress in semantic image synthesis, it remains a challenging problem to generate photo-realistic parts from input semantic map. Integrating part segmentation map can undoubtedly benefit image synthesis, but is bothersome and…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Yuxiang Wei , Zhilong Ji , Xiaohe Wu , Jinfeng Bai , Lei Zhang , Wangmeng Zuo

Long term human motion prediction is essential in safety-critical applications such as human-robot interaction and autonomous driving. In this paper we show that to achieve long term forecasting, predicting human pose at every time instant…

Computer Vision and Pattern Recognition · Computer Science 2022-09-05 Sena Kiciroglu , Wei Wang , Mathieu Salzmann , Pascal Fua

Accurate estimation of the environment structure simultaneously with the robot pose is a key capability of autonomous robotic vehicles. Classical simultaneous localization and mapping (SLAM) algorithms rely on the static world assumption to…

Robotics · Computer Science 2018-05-11 Mina Henein , Gerard Kennedy , Viorela Ila , Robert Mahony

State-of-the-art object pose estimation methods are prone to generating geometrically infeasible pose hypotheses. This problem is prevalent in dexterous manipulation, where estimated poses often intersect with the robotic hand or are not…

Robotics · Computer Science 2026-03-24 Anil Zeybek , Rhys Newbury , Snehal Dikhale , Nawid Jamali , Soshi Iba , Akansel Cosgun

In this paper, we propose a real-time adaptive prediction method to calculate smooth and accurate haptic feedback in complex scenarios. Smooth haptic feedback is an important task for haptic rendering with complex virtual objects. However,…

Human-Computer Interaction · Computer Science 2016-03-23 Xiyuan Hou , Olga Sourina

We propose novel motion representations for animating articulated objects consisting of distinct parts. In a completely unsupervised manner, our method identifies object parts, tracks them in a driving video, and infers their motions by…

Computer Vision and Pattern Recognition · Computer Science 2021-04-26 Aliaksandr Siarohin , Oliver J. Woodford , Jian Ren , Menglei Chai , Sergey Tulyakov

Reconstructing the motion of objects from videos is a key component for embodied AI and robot manipulation. While diverse approaches to object pose tracking have been studied, they rely heavily on strong external priors, such as depth data…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Jisu Shin , Junoh Lee , JunGyu Lee , Inhwan Bae , Dohyeon Lee , Hokyun Im , Youngwoon Lee , Hae-Gon Jeon

We propose to leverage the local information in image sequences to support global camera relocalization. In contrast to previous methods that regress global poses from single images, we exploit the spatial-temporal consistency in sequential…

Computer Vision and Pattern Recognition · Computer Science 2019-08-14 Fei Xue , Xin Wang , Zike Yan , Qiuyuan Wang , Junqiu Wang , Hongbin Zha