English
Related papers

Related papers: Inferring Dynamic Physical Properties from Video F…

200 papers

A long-standing question in physical reasoning is whether video-based models need to rely on factorized representations of physical variables in order to make physically accurate predictions, or whether they can implicitly represent such…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Sonia Joseph , Quentin Garrido , Randall Balestriero , Matthew Kowal , Thomas Fel , Shahab Bakhtiari , Blake Richards , Mike Rabbat

In order to reach human performance on complexvisual tasks, artificial systems need to incorporate a sig-nificant amount of understanding of the world in termsof macroscopic objects, movements, forces, etc. Inspiredby work on intuitive…

Artificial Intelligence · Computer Science 2020-02-12 Ronan Riochet , Mario Ynocente Castro , Mathieu Bernard , Adam Lerer , Rob Fergus , Véronique Izard , Emmanuel Dupoux

Learning the dynamics of robots from data can help achieve more accurate tracking controllers, or aid their navigation algorithms. However, when the actual dynamics of the robots change due to external conditions, on-line adaptation of…

Robotics · Computer Science 2019-03-14 Bilal Wehbe , Marc Hildebrandt , Frank Kirchner

3D human pose estimation from a monocular video has recently seen significant improvements. However, most state-of-the-art methods are kinematics-based, which are prone to physically implausible motions with pronounced artifacts. Current…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 Jiefeng Li , Siyuan Bian , Chao Xu , Gang Liu , Gang Yu , Cewu Lu

Data-driven learning approaches for physics simulation, sometimes referred to as world models, have emerged as promising alternatives to traditional physics simulators due to their differentiable nature. Prior work has demonstrated…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Chanho Kim , Suhas V. Sumukh , Li Fuxin

While stochastic video prediction models enable future prediction under uncertainty, they mostly fail to model the complex dynamics of real-world scenes. For example, they cannot provide reliable predictions for scenes with a moving camera…

Computer Vision and Pattern Recognition · Computer Science 2022-05-02 Adil Kaan Akan , Sadra Safadoust , Fatma Güney

The abundance of instructional videos and their narrations over the Internet offers an exciting avenue for understanding procedural activities. In this work, we propose to learn video representation that encodes both action steps and their…

Computer Vision and Pattern Recognition · Computer Science 2023-04-03 Yiwu Zhong , Licheng Yu , Yang Bai , Shangwen Li , Xueting Yan , Yin Li

We are interested in learning models of intuitive physics similar to the ones that animals use for navigation, manipulation and planning. In addition to learning general physical principles, however, we are also interested in learning ``on…

Computer Vision and Pattern Recognition · Computer Science 2019-05-28 Sébastien Ehrhardt , Aron Monszpart , Niloy J. Mitra , Andrea Vedaldi

Understanding physical phenomena is a key competence that enables humans and animals to act and interact under uncertain perception in previously unseen environments containing novel object and their configurations. Developmental psychology…

Computer Vision and Pattern Recognition · Computer Science 2016-04-04 Wenbin Li , Seyedmajid Azimi , Aleš Leonardis , Mario Fritz

Human motion capture from monocular videos has made significant progress in recent years. However, modern approaches often produce temporal artifacts, e.g. in form of jittery motion and struggle to achieve smooth and physically plausible…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Cuong Le , Viktor Johansson , Manon Kok , Bastian Wandt

Can we learn the physics of matter in motion directly from images and video--and trust it? Answering this question requires integrating experiments, physics-based simulation, and data across traditionally separate disciplines. Much of this…

Computational Engineering, Finance, and Science · Computer Science 2026-04-21 Hagen Holthusen , Kevin Linka , Ellen Kuhl

Video motion magnification techniques allow us to see small motions previously invisible to the naked eyes, such as those of vibrating airplane wings, or swaying buildings under the influence of the wind. Because the motion is small, the…

Computer Vision and Pattern Recognition · Computer Science 2019-02-18 Tae-Hyun Oh , Ronnachai Jaroensri , Changil Kim , Mohamed Elgharib , Frédo Durand , William T. Freeman , Wojciech Matusik

Humans are able to make rich predictions about the future dynamics of physical objects from a glance. On the other hand, most existing computer vision approaches require strong assumptions about the underlying system, ad-hoc modeling, or…

Neural and Evolutionary Computing · Computer Science 2018-09-19 Zhihua Wang , Stefano Rosa , Yishu Miao , Zihang Lai , Linhai Xie , Andrew Markham , Niki Trigoni

While text-to-video diffusion models have made significant strides, many still face challenges in generating videos with temporal consistency. Within diffusion frameworks, guidance techniques have proven effective in enhancing output…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Hyelin Nam , Jaemin Kim , Dohun Lee , Jong Chul Ye

Despite impressive visual fidelity, modern video generative models frequently produce sequences that violate intuitive physical laws, such as objects floating, teleporting, or morphing in ways that defy causality. While humans can easily…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Saman Motamed , Minghao Chen , Luc Van Gool , Iro Laina

A defining characteristic of intelligent systems is the ability to make action decisions based on the anticipated outcomes. Video prediction systems have been demonstrated as a solution for predicting how the future will unfold visually,…

Computer Vision and Pattern Recognition · Computer Science 2019-10-08 Manuel Serra Nunes , Atabak Dehban , Plinio Moreno , José Santos-Victor

Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elusive. Critically, existing pure-vision encoders and Multimodal Large Language Models (MLLMs)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Xavier Thomas , Youngsun Lim , Ananya Srinivasan , Audrey Zheng , Deepti Ghadiyaram

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Aritra Bhowmik , Denis Korzhenkov , Cees G. M. Snoek , Amirhossein Habibian , Mohsen Ghafoorian

Video action localization aims to find the timings of specific actions from a long video. Although existing learning-based approaches have been successful, they require annotating videos, which comes with a considerable labor cost. This…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Naoki Wake , Atsushi Kanehira , Kazuhiro Sasabuchi , Jun Takamatsu , Katsushi Ikeuchi

From just a short glance at a video, we can often tell whether a person's action is intentional or not. Can we train a model to recognize this? We introduce a dataset of in-the-wild videos of unintentional action, as well as a suite of…

Computer Vision and Pattern Recognition · Computer Science 2019-11-27 Dave Epstein , Boyuan Chen , Carl Vondrick