English

PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation

Robotics 2026-07-02 v1 Artificial Intelligence Computation and Language Computer Vision and Pattern Recognition Machine Learning

Abstract

Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI. Existing visual-language-action models and world models struggle with accurate 3D geometry and physically meaningful forecasting. We propose PhysMani, a framework that couples a physics-principled 3D Gaussian world model with a future-aware action policy model. The world model learns a divergence-free Gaussian velocity field via online optimization for fast and physically grounded future dynamics prediction. The policy model integrates the predicted 3D scene future dynamics through a learnable token based cross-attention module. We introduce PhysMani-Bench, a dynamic manipulation benchmark with 16 tasks, and demonstrate a superior success rate over strong baselines in both simulation and real-world robot experiments.

Cite

@article{arxiv.2607.01938,
  title  = {PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation},
  author = {Peng Yun and Shouwang Huang and Hao Li and Jinxi Li and Jianan Wang and Bo Yang},
  journal= {arXiv preprint arXiv:2607.01938},
  year   = {2026}
}

Comments

ECCV 2026. Code and data are available at: https://github.com/vLAR-group/PhysMani