English
Related papers

Related papers: DreamPhysics: Learning Physics-Based 3D Dynamics w…

200 papers

Video prediction is a useful function for autonomous driving, enabling intelligent vehicles to reliably anticipate how driving scenes will evolve and thereby supporting reasoning and safer planning. However, existing models are constrained…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Ke Li , Tianjia Yang , Kaidi Liang , Xianbiao Hu , Ruwen Qin

For vision-language models (VLMs), understanding the dynamic properties of objects and their interactions in 3D scenes from videos is crucial for effective reasoning about high-level temporal and action semantics. Although humans are adept…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Xingrui Wang , Wufei Ma , Angtian Wang , Shuo Chen , Adam Kortylewski , Alan Yuille

We explore a novel video creation experience, namely Video Creation by Demonstration. Given a demonstration video and a context image from a different scene, we generate a physically plausible video that continues naturally from the context…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Yihong Sun , Hao Zhou , Liangzhe Yuan , Jennifer J. Sun , Yandong Li , Xuhui Jia , Hartwig Adam , Bharath Hariharan , Long Zhao , Ting Liu

Creating deformable 3D content has gained increasing attention with the rise of text-to-image and image-to-video generative models. While these models provide rich semantic priors for appearance, they struggle to capture the physical…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Jixuan He , Chieh Hubert Lin , Lu Qi , Ming-Hsuan Yang

In this paper, we aim to model 3D scene geometry, appearance, and the underlying physics purely from multi-view videos. By applying various governing PDEs as PINN losses or incorporating physics simulation into neural networks, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Jinxi Li , Ziyang Song , Siyuan Zhou , Bo Yang

We propose a method to learn explicit, class-conditioned spatial priors for object placement in natural scenes by distilling the implicit placement knowledge encoded in text-conditioned diffusion models. Prior work relies either on manually…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Marco Schouten , Ioannis Siglidis , Serge Belongie , Dim P. Papadopoulos

We present a method to simulate fluid flow on evolving surfaces, e.g., an oil film on a water surface. Given an animated surface (e.g., extracted from a particle-based fluid simulation) in three-dimensional space, we add a second simulation…

Graphics · Computer Science 2020-10-08 Dieter Morgenroth , Stefan Reinhardt , Daniel Weiskopf , Bernhard Eberhardt

This paper presents a novel approach to inpainting 3D regions of a scene, given masked multi-view images, by distilling a 2D diffusion model into a learned 3D scene representation (e.g. a NeRF). Unlike 3D generative methods that explicitly…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Kira Prabhu , Jane Wu , Lynn Tsai , Peter Hedman , Dan B Goldman , Ben Poole , Michael Broxton

Multiview diffusion models have rapidly emerged as a powerful tool for content creation with spatial consistency across viewpoints, offering rich visual realism without requiring explicit geometry and appearance representation. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Hubert Kompanowski , Varun Jampani , Aaryaman Vasishta , Binh-Son Hua

Can we learn the physics of matter in motion directly from images and video--and trust it? Answering this question requires integrating experiments, physics-based simulation, and data across traditionally separate disciplines. Much of this…

Computational Engineering, Finance, and Science · Computer Science 2026-04-21 Hagen Holthusen , Kevin Linka , Ellen Kuhl

Recent text-guided generation of individual 3D object has achieved great success using diffusion priors. However, these methods are not suitable for object insertion and replacement tasks as they do not consider the background, leading to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Hanyuan Xiao , Yingshu Chen , Huajian Huang , Haolin Xiong , Jing Yang , Pratusha Prasad , Yajie Zhao

A recent frontier in computer vision has been the task of 3D video generation, which consists of generating a time-varying 3D representation of a scene. To generate dynamic 3D scenes, current methods explicitly model 3D temporal dynamics by…

Computer Vision and Pattern Recognition · Computer Science 2024-08-01 Rishab Parthasarathy , Zachary Ankner , Aaron Gokaslan

We propose a technique that tackles action detection in multimodal videos under a realistic and challenging condition in which only limited training data and partially observed modalities are available. Common methods in transfer learning…

Computer Vision and Pattern Recognition · Computer Science 2018-07-31 Zelun Luo , Jun-Ting Hsieh , Lu Jiang , Juan Carlos Niebles , Li Fei-Fei

Modeling wind-driven object dynamics from video observations is highly challenging due to the invisibility and spatio-temporal variability of wind, as well as the complex deformations of objects. We present DiffWind, a physics-informed…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yuanhang Lei , Boming Zhao , Zesong Yang , Xingxuan Li , Tao Cheng , Haocheng Peng , Ru Zhang , Yang Yang , Siyuan Huang , Yujun Shen , Ruizhen Hu , Hujun Bao , Zhaopeng Cui

Self-supervised and language-supervised image models contain rich knowledge of the world that is important for generalization. Many robotic tasks, however, require a detailed understanding of 3D geometry, which is often lacking in 2D image…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 William Shen , Ge Yang , Alan Yu , Jansen Wong , Leslie Pack Kaelbling , Phillip Isola

While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deformation or spatial drift. We hypothesize that these failures…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Hongyang Du , Junjie Ye , Xiaoyan Cong , Runhao Li , Jingcheng Ni , Aman Agarwal , Zeqi Zhou , Zekun Li , Randall Balestriero , Yue Wang

We explore the problem of computationally generating special `prime' images that produce optical illusions when physically arranged and viewed in a certain way. First, we propose a formal definition for this problem. Next, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Ryan Burgert , Xiang Li , Abe Leite , Kanchana Ranasinghe , Michael S. Ryoo

Forecasting a typical object's future motion is a critical task for interpreting and interacting with dynamic environments in computer vision. Event-based sensors, which could capture changes in the scene with exceptional temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-10-14 Song Wu , Zhiyu Zhu , Junhui Hou , Guangming Shi , Jinjian Wu

Human mesh recovery (HMR) provides rich human body information for various real-world applications. While image-based HMR methods have achieved impressive results, they often struggle to recover humans in dynamic scenarios, leading to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Ce Zheng , Xianpeng Liu , Qucheng Peng , Tianfu Wu , Pu Wang , Chen Chen

Integrating motion into static images not only enhances visual expressiveness but also creates a sense of immersion and temporal depth, establishing it as a longstanding and impactful theme in artistic expression. Fluid elements such as…

Graphics · Computer Science 2025-10-21 Hao Jin , Haoran Xie