English
Related papers

Related papers: 3D-IntPhys: Towards More Generalized 3D-grounded V…

200 papers

Three-dimensional (3D) understanding of objects and scenes play a key role in humans' ability to interact with the world and has been an active area of research in computer vision, graphics, and robotics. Large scale synthetic and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Matthew Wallingford , Anand Bhattad , Aditya Kusupati , Vivek Ramanujan , Matt Deitke , Sham Kakade , Aniruddha Kembhavi , Roozbeh Mottaghi , Wei-Chiu Ma , Ali Farhadi

Forecasting future scenarios in dynamic environments is essential for intelligent decision-making and navigation, a challenge yet to be fully realized in computer vision and robotics. Traditional approaches like video prediction and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Boming Zhao , Yuan Li , Ziyu Sun , Lin Zeng , Yujun Shen , Rui Ma , Yinda Zhang , Hujun Bao , Zhaopeng Cui

A major element of depth perception and 3D understanding is the ability to predict the 3D layout of a scene and its contained objects for a novel pose. Indoor environments are particularly suitable for novel view prediction, since the set…

Computer Vision and Pattern Recognition · Computer Science 2018-08-13 Pulak Purkait , Ujwal Bonde , Christopher Zach

Physically plausible fluid simulations play an important role in modern computer graphics and engineering. However, in order to achieve real-time performance, computational speed needs to be traded-off with physical accuracy. Surrogate…

Fluid Dynamics · Physics 2021-05-19 Nils Wandel , Michael Weinmann , Reinhard Klein

We present a novel paradigm of building an animatable 3D human representation from a monocular video input, such that it can be rendered in any unseen poses and views. Our method is based on a dynamic Neural Radiance Field (NeRF) rigged by…

Computer Vision and Pattern Recognition · Computer Science 2022-08-19 Gusi Te , Xiu Li , Xiao Li , Jinglu Wang , Wei Hu , Yan Lu

Perceptual learning enables humans to recognize and represent stimuli invariant to various transformations and build a consistent representation of the self and physical world. Such representations preserve the invariant physical relations…

Neural and Evolutionary Computing · Computer Science 2020-07-02 Du Xiaorui , Yavuzhan Erdem , Immanuel Schweizer , Cristian Axenie

Current 3D inpainting and object removal methods are largely limited to front-facing scenes, facing substantial challenges when applied to diverse, "unconstrained" scenes where the camera orientation and trajectory are unrestricted. To…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Zhihao Shi , Dong Huo , Yuhongze Zhou , Kejia Yin , Yan Min , Juwei Lu , Xinxin Zuo

A natural approach to generative modeling of videos is to represent them as a composition of moving objects. Recent works model a set of 2D sprites over a slowly-varying background, but without considering the underlying 3D scene that gives…

Computer Vision and Pattern Recognition · Computer Science 2021-03-26 Paul Henderson , Christoph H. Lampert

Learning sensorimotor control policies from high-dimensional images crucially relies on the quality of the underlying visual representations. Prior works show that structured latent space such as visual keypoints often outperforms…

Machine Learning · Computer Science 2021-06-15 Boyuan Chen , Pieter Abbeel , Deepak Pathak

We present a physics-based inverse rendering method that learns the illumination, geometry, and materials of a scene from posed multi-view RGB images. To model the illumination of a scene, existing inverse rendering works either completely…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Youming Deng , Xueting Li , Sifei Liu , Ming-Hsuan Yang

Developing deep neural networks to generate 3D scenes is a fundamental problem in neural synthesis with immediate applications in architectural CAD, computer graphics, as well as in generating virtual robot training environments. This task…

Computer Vision and Pattern Recognition · Computer Science 2021-09-02 Haitao Yang , Zaiwei Zhang , Siming Yan , Haibin Huang , Chongyang Ma , Yi Zheng , Chandrajit Bajaj , Qixing Huang

Multi-view implicit scene reconstruction methods have become increasingly popular due to their ability to represent complex scene details. Recent efforts have been devoted to improving the representation of input information and to reducing…

Computer Vision and Pattern Recognition · Computer Science 2022-10-04 Edward J. Smith , Michal Drozdzal , Derek Nowrouzezahrai , David Meger , Adriana Romero-Soriano

The rapid advancement of Multimodal Large Language Models (MLLMs) has significantly impacted various multimodal tasks. However, these models face challenges in tasks that require spatial understanding within 3D environments. Efforts to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Duo Zheng , Shijia Huang , Liwei Wang

Neural networks have recently been used to analyze diverse physical systems and to identify the underlying dynamics. While existing methods achieve impressive results, they are limited by their strong demand for training data and their weak…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Florian Hofherr , Lukas Koestler , Florian Bernard , Daniel Cremers

Two-phase flow phenomena underpin critical technologies such as hydrogen fuel cells, spray cooling, and combustion, where droplet dynamics govern performance and efficiency. Conventional optical diagnostics, including shadowgraphy and…

Neural surrogate models for computational fluid dynamics (CFD) are typically trained as forward operators that map explicit problem specifications, such as geometry and boundary conditions, to solution fields. This ties the model to the…

Machine Learning · Computer Science 2026-05-29 Jonas Weidner , Yeray Martin-Ruisanchez , Daniel Rueckert , Benedikt Wiestler , Julian Suk

While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and physical dynamics. Existing solutions typically rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Xianjin Wu , Dingkang Liang , Tianrui Feng , Kui Xia , Yumeng Zhang , Xiaofan Li , Xiao Tan , Xiang Bai

We present a deep reinforcement learning method of progressive view inpainting for 3D point scene completion under volume guidance, achieving high-quality scene reconstruction from only a single depth image with severe occlusion. Our…

Computer Vision and Pattern Recognition · Computer Science 2019-03-13 Xiaoguang Han , Zhaoxuan Zhang , Dong Du , Mingdai Yang , Jingming Yu , Pan Pan , Xin Yang , Ligang Liu , Zixiang Xiong , Shuguang Cui

Our goal in this work is to generate realistic videos given just one initial frame as input. Existing unsupervised approaches to this task do not consider the fact that a video typically shows a 3D environment, and that this should remain…

Computer Vision and Pattern Recognition · Computer Science 2021-06-18 Paul Henderson , Christoph H. Lampert , Bernd Bickel

Predicting scene dynamics from visual observations is challenging. Existing methods capture dynamics only within observed boundaries failing to extrapolate far beyond the training sequence. Node-RF (Neural ODE-based NeRF) overcomes this…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Hiran Sarkar , Liming Kuang , Yordanka Velikova , Benjamin Busam