中文
相关论文

相关论文: ContactGaussian-WM: Learning Physics-Grounded Worl…

200 篇论文

Rigid body interactions are fundamental to numerous scientific disciplines, but remain challenging to simulate due to their abrupt nonlinear nature and sensitivity to complex, often unknown environmental factors. These challenges call for…

机器学习 · 计算机科学 2025-07-28 Amaury Wei , Olga Fink

Knowledge of the physical material properties governing the dynamics of a real-world object becomes necessary to accurately anticipate its response to unseen interactions. Existing methods for estimating such physical material parameters…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Norika Wada , Kohei Yamashita , Ryo Kawahara , Ko Nishino

We introduce Vysics, a vision-and-physics framework for a robot to build an expressive geometry and dynamics model of a single rigid body, using a seconds-long RGBD video and the robot's proprioception. While the computer vision community…

机器人学 · 计算机科学 2025-04-29 Bibit Bianchini , Minghan Zhu , Mengti Sun , Bowen Jiang , Camillo J. Taylor , Michael Posa

Extracting compact, physically interpretable representations from high-dimensional scientific data is a persistent challenge due to the complex, nonlinear structures inherent in physical systems. We propose a Gaussian Mixture Variational…

机器学习 · 计算机科学 2025-12-01 Tiffany Fan , Murray Cutforth , Marta D'Elia , Alexandre Cortiella , Alireza Doostan , Eric Darve

Interactive world models that simulate object dynamics are crucial for robotics, VR, and AR. However, it remains a significant challenge to learn physics-consistent dynamics models from limited real-world video data, especially for…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Yu Yang , Zhilu Zhang , Xiang Zhang , Yihan Zeng , Hui Li , Wangmeng Zuo

Action-conditioned video prediction models (often referred to as world models) have shown strong potential for robotics applications, but existing approaches are often slow and struggle to capture physically consistent interactions over…

Vision-based autonomous driving has gained much attention due to its low costs and excellent performance. Compared with dense BEV (Bird's Eye View) or sparse query models, Gaussian-centric method is a comprehensive yet sparse representation…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Yiyao Zhu , Ying Xue , Haiming Zhang , Guangfeng Jiang , Wending Zhou , Xu Yan , Jiantao Gao , Yingjie Cai , Bingbing Liu , Zhen Li , Shaojie Shen

Humans develop an understanding of intuitive physics through active interaction with the world. This approach is in stark contrast to current video models, such as Sora, which rely on passive observation and therefore struggle with grasping…

In this paper, we aim to model 3D scene geometry, appearance, and the underlying physics purely from multi-view videos. By applying various governing PDEs as PINN losses or incorporating physics simulation into neural networks, existing…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Jinxi Li , Ziyang Song , Siyuan Zhou , Bo Yang

Generating intelligent robot behavior in contact-rich settings is a research problem where zeroth-order methods currently prevail. A major contributor to the success of such methods is their robustness in the face of non-smooth and…

机器人学 · 计算机科学 2025-04-15 Onur Beker , Nico Gürtler , Ji Shi , A. René Geist , Amirreza Razmjoo , Georg Martius , Sylvain Calinon

Real objects that inhabit the physical world follow physical laws and thus behave plausibly during interaction with other physical objects. However, current methods that perform 3D reconstructions of real-world scenes from multi-view 2D…

Autonomous driving requires robust perception models trained on high-quality, large-scale multi-view driving videos for tasks like 3D object detection, segmentation and trajectory prediction. While world models provide a cost-effective…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Zhuoran Yang , Xi Guo , Chenjing Ding , Chiyu Wang , Wei Wu

Given a visual scene, humans have strong intuitions about how a scene can evolve over time under given actions. The intuition, often termed visual intuitive physics, is a critical ability that allows us to make effective plans to manipulate…

计算机视觉与模式识别 · 计算机科学 2023-04-25 Haotian Xue , Antonio Torralba , Joshua B. Tenenbaum , Daniel LK Yamins , Yunzhu Li , Hsiao-Yu Tung

Human motion synthesis is an important problem with applications in graphics, gaming and simulation environments for robotics. Existing methods require accurate motion capture data for training, which is costly to obtain. Instead, we…

计算机视觉与模式识别 · 计算机科学 2022-08-15 Kevin Xie , Tingwu Wang , Umar Iqbal , Yunrong Guo , Sanja Fidler , Florian Shkurti

Learning and predicting the dynamics of physical systems requires a profound understanding of the underlying physical laws. Recent works on learning physical laws involve generalizing the equation discovery frameworks to the discovery of…

机器学习 · 统计学 2023-10-11 Tapas Tripura , Souvik Chakraborty

Robots struggle to understand object properties like shape, material, and semantics due to limited prior knowledge, hindering manipulation in unstructured environments. In contrast, humans learn these properties through interactive…

机器人学 · 计算机科学 2025-07-09 Ho Jin Choi , Nadia Figueroa

We introduce a data-driven method for learning the equations of motion of mechanical systems directly from position measurements, without requiring access to velocity data. This is particularly relevant in system identification tasks where…

系统与控制 · 电气工程与系统科学 2025-05-28 Martine Dyring Hansen , Elena Celledoni , Benjamin Kwanen Tapley

Young children demonstrate early abilities to understand their physical world, estimating depth, motion, object coherence, interactions, and many other aspects of physical scene understanding. Children are both data-efficient and flexible…

A structured understanding of our world in terms of objects, relations, and hierarchies is an important component of human cognition. Learning such a structured world model from raw sensory data remains a challenge. As a step towards this…

机器学习 · 统计学 2020-01-07 Thomas Kipf , Elise van der Pol , Max Welling

World models derived from large-scale video generative pre-training have emerged as a promising paradigm for generalist robot policy learning. However, standard approaches often focus on high-fidelity RGB video prediction, this can result…