中文
相关论文

相关论文: 3D-IntPhys: Towards More Generalized 3D-grounded V…

200 篇论文

Physics driven image simulation allows for the modeling and creation of realistic imagery beyond what is afforded by typical rendering pipelines. We aim to automatically generate a physically realistic scene for simulation of a given region…

计算机视觉与模式识别 · 计算机科学 2025-04-23 Scott Sorensen , Wayne Treible , Robert Wagner , Andrew D. Gilliam , Todd Rovito , Joseph L. Mundy

Next-frame prediction is a useful and powerful method for modelling and understanding the dynamics of video data. Inspired by the empirical success of causal language modelling and next-token prediction in language modelling, we explore the…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Thomas Winterbottom , G. Thomas Hudson , Daniel Kluvanec , Dean Slack , Jamie Sterling , Junjie Shentu , Chenghao Xiao , Zheming Zhou , Noura Al Moubayed

We study the task of predicting dynamic physical properties from videos. More specifically, we consider physical properties that require temporal information to be inferred: elasticity of a bouncing object, viscosity of a flowing liquid,…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Guanqi Zhan , Xianzheng Ma , Weidi Xie , Andrew Zisserman

In this paper, we present a novel, scalable approach for constructing open set, instance-level 3D scene representations, advancing open world understanding of 3D environments. Existing methods require pre-constructed 3D scenes and face…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Rafay Mohiuddin , Sai Manoj Prakhya , Fiona Collins , Ziyuan Liu , André Borrmann

We investigate whether video generative models can exhibit visuospatial intelligence, a capability central to human cognition, using only visual data. To this end, we present Video4Spatial, a framework showing that video diffusion models…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Zeqi Xiao , Yiwei Zhao , Lingxiao Li , Yushi Lan , Ning Yu , Rahul Garg , Roshni Cooper , Mohammad H. Taghavi , Xingang Pan

Real objects that inhabit the physical world follow physical laws and thus behave plausibly during interaction with other physical objects. However, current methods that perform 3D reconstructions of real-world scenes from multi-view 2D…

We present a method, Neural Radiance Flow (NeRFlow),to learn a 4D spatial-temporal representation of a dynamic scene from a set of RGB images. Key to our approach is the use of a neural implicit representation that learns to capture the 3D…

计算机视觉与模式识别 · 计算机科学 2021-09-07 Yilun Du , Yinan Zhang , Hong-Xing Yu , Joshua B. Tenenbaum , Jiajun Wu

Humans have a strong intuitive understanding of physical processes such as fluid falling by just a glimpse of such a scene picture, i.e., quickly derived from our immersive visual experiences in memory. This work achieves such a…

计算机视觉与模式识别 · 计算机科学 2023-04-11 Jinxian Liu , Ye Chen , Bingbing Ni , Jiyao Mao , Zhenbo Yu

One major goal of vision is to infer physical models of objects, surfaces, and their layout from sensors. In this paper, we aim to interpret indoor scenes from one RGBD image. Our representation encodes the layout of orthogonal walls and…

计算机视觉与模式识别 · 计算机科学 2018-11-15 Chuhang Zou , Ruiqi Guo , Zhizhong Li , Derek Hoiem

Traditional fluid dynamics simulation pipelines combine numerical solvers with rendering, producing highly realistic results but at considerable computational cost. Diffusion-based generative video models offer a faster alternative, yet…

图形学 · 计算机科学 2026-03-18 Yang Bai , George Eskandar , Ziyuan Liu , Gitta Kutyniok

Physics perception very often faces the problem that only limited data or partial measurements on the scene are available. In this work, we propose a strategy to learn the full state of sloshing liquids from measurements of the free…

计算机视觉与模式识别 · 计算机科学 2022-02-25 Beatriz Moya , Alberto Badias , David Gonzalez , Francisco Chinesta , Elias Cueto

We introduce Consistent Instance Field, a continuous and probabilistic spatio-temporal representation for dynamic scene understanding. Unlike prior methods that rely on discrete tracking or view-dependent features, our approach disentangles…

In this paper, we focus on 3D scene inpainting, where parts of an input image set, captured from different viewpoints, are masked out. The main challenge lies in generating plausible image completions that are geometrically consistent…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Ahmad Salimi , Tristan Aumentado-Armstrong , Marcus A. Brubaker , Konstantinos G. Derpanis

Dynamic 3D interaction has been attracting a lot of attention recently. However, creating such 4D content remains challenging. One solution is to animate 3D scenes with physics-based simulation, which requires manually assigning precise…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Tianyu Huang , Haoze Zhang , Yihan Zeng , Zhilu Zhang , Hui Li , Wangmeng Zuo , Rynson W. H. Lau

Realistic simulation is critical for applications ranging from robotics to animation. Learned simulators have emerged as a possibility to capture real world physics directly from video data, but very often require privileged information…

图形学 · 计算机科学 2025-08-12 Mikel Zhobro , Andreas René Geist , Georg Martius

In many real-world settings, image observations of freely rotating 3D rigid bodies, such as satellites, may be available when low-dimensional measurements are not. However, the high-dimensionality of image data precludes the use of…

计算机视觉与模式识别 · 计算机科学 2023-08-24 Justice Mason , Christine Allen-Blanchette , Nicholas Zolman , Elizabeth Davison , Naomi Leonard

Learning robust visuomotor policies that generalize across diverse objects and interaction dynamics remains a central challenge in robotic manipulation. Most existing approaches rely on direct observation-to-action mappings or compress…

机器人学 · 计算机科学 2025-09-24 Sangjun Noh , Dongwoo Nam , Kangmin Kim , Geonhyup Lee , Yeonguk Yu , Raeyoung Kang , Kyoobin Lee

Deep learning has shown great potential for modeling the physical dynamics of complex particle systems such as fluids. Existing approaches, however, require the supervision of consecutive particle properties, including positions and…

机器学习 · 计算机科学 2022-06-22 Shanyan Guan , Huayu Deng , Yunbo Wang , Xiaokang Yang

We present a new point-based approach for modeling the appearance of real scenes. The approach uses a raw point cloud as the geometric representation of a scene, and augments each point with a learnable neural descriptor that encodes local…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Kara-Ali Aliev , Artem Sevastopolsky , Maria Kolos , Dmitry Ulyanov , Victor Lempitsky

Indoor scene understanding is central to applications such as robot navigation and human companion assistance. Over the last years, data-driven deep neural networks have outperformed many traditional approaches thanks to their…

计算机视觉与模式识别 · 计算机科学 2017-07-04 Yinda Zhang , Shuran Song , Ersin Yumer , Manolis Savva , Joon-Young Lee , Hailin Jin , Thomas Funkhouser