中文
相关论文

相关论文: ContactGaussian-WM: Learning Physics-Grounded Worl…

200 篇论文

Understanding 3D scenes requires flexible combinations of visual reasoning tasks, including depth estimation, novel view synthesis, and object manipulation, all of which are essential for perception and interaction. Existing approaches have…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Wanhee Lee , Klemen Kotar , Rahul Mysore Venkatesh , Jared Watrous , Honglin Chen , Khai Loong Aw , Daniel L. K. Yamins

World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current embodied world models…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Yu Shang , Xin Zhang , Yinzhou Tang , Lei Jin , Chen Gao , Wei Wu , Yong Li

Learning predictive world models from visual observations is a core problem in embodied AI, with applications to model-based reinforcement learning and robotic planning. Existing latent world models typically generate future states with…

机器学习 · 计算机科学 2026-05-12 Qixin Xiao , Maani Ghaffari

In recent years, there has been rapid development in 3D generation models, opening up new possibilities for applications such as simulating the dynamic movements of 3D objects and customizing their behaviors. However, current 3D generative…

计算机视觉与模式识别 · 计算机科学 2024-06-12 Fangfu Liu , Hanyang Wang , Shunyu Yao , Shengjun Zhang , Jie Zhou , Yueqi Duan

For robots to robustly understand and interact with the physical world, it is highly beneficial to have a comprehensive representation - modelling geometry, physics, and visual observations - that informs perception, planning, and control…

机器人学 · 计算机科学 2024-06-18 Jad Abou-Chakra , Krishan Rana , Feras Dayoub , Niko Sünderhauf

Distilling interpretable physical laws from videos has led to expanded interest in the computer vision community recently thanks to the advances in deep learning, but still remains a great challenge. This paper introduces an end-to-end…

计算机视觉与模式识别 · 计算机科学 2022-05-04 Lele Luan , Yang Liu , Hao Sun

Can we learn the physics of matter in motion directly from images and video--and trust it? Answering this question requires integrating experiments, physics-based simulation, and data across traditionally separate disciplines. Much of this…

计算工程、金融与科学 · 计算机科学 2026-04-21 Hagen Holthusen , Kevin Linka , Ellen Kuhl

Understanding dynamic scenes from casual videos is critical for scalable robot learning, yet four-dimensional (4D) reconstruction under strictly monocular settings remains highly ill-posed. To address this challenge, our key insight is that…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Can Li , Jie Gu , Jingmin Chen , Fangzhou Qiu , Lei Sun

In this paper, we aim to model 3D scene geometry, appearance, and physical information just from dynamic multi-view videos in the absence of any human labels. By leveraging physics-informed losses as soft constraints or integrating simple…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Jinxi Li , Ziyang Song , Bo Yang

Estimating physical properties for visual data is a crucial task in computer vision, graphics, and robotics, underpinning applications such as augmented reality, physical simulation, and robotic grasping. However, this area remains…

We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language commands and images,…

We tackle the problem of learning complex, general behaviors directly in the real world. We propose an approach for robots to efficiently learn manipulation skills using only a handful of real-world interaction trajectories from many…

机器人学 · 计算机科学 2023-08-22 Russell Mendonca , Shikhar Bahl , Deepak Pathak

Predicting physical dynamics from raw visual data remains a major challenge in AI. While recent video generation models have achieved impressive visual quality, they still cannot consistently generate physically plausible videos due to a…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Shiqian Li , Ruihong Shen , Junfeng Ni , Chang Pan , Chi Zhang , Yixin Zhu

Imitation learning has emerged as a promising approach towards building generalist robots. However, scaling imitation learning for large robot foundation models remains challenging due to its reliance on high-quality expert demonstrations.…

机器人学 · 计算机科学 2025-05-26 Chuning Zhu , Raymond Yu , Siyuan Feng , Benjamin Burchfiel , Paarth Shah , Abhishek Gupta

Model-based planning in robotic domains is challenged by the hybrid nature of physical dynamics, where continuous motion is punctuated by discrete events such as contacts and impacts. Conventional latent world models typically employ…

人工智能 · 计算机科学 2026-05-14 Mingwei Li , Xiaoyuan Zhang , Chengwei Yang , Zilong Zheng , Yaodong Yang

World models for deformable objects should recover not only geometry and appearance, but also underlying physical dynamics, interaction grounding, and material behavior. Learning such a model from real videos is challenging because…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Can Li , Zhoujian Li , Ren Li , Jie Gu , Lei Lei , Jingmin Chen , Lei Sun

The scalability of embodied intelligence is fundamentally constrained by the scarcity of real-world interaction data. While simulation platforms provide a promising alternative, existing approaches often suffer from a substantial visual and…

Distilling analytical models from data has the potential to advance our understanding and prediction of nonlinear dynamics. Although discovery of governing equations based on observed system states (e.g., trajectory time series) has…

机器学习 · 计算机科学 2021-06-10 Lele Luan , Yang Liu , Hao Sun

Specifying tasks with videos is a powerful technique towards acquiring novel and general robot skills. However, reasoning over mechanics and dexterous interactions can make it challenging to scale learning contact-rich manipulation. In this…

机器人学 · 计算机科学 2021-11-10 Bernardo Aceituno , Alberto Rodriguez , Shubham Tulsiani , Abhinav Gupta , Mustafa Mukadam

We introduce PhysMotion, a novel framework that leverages principled physics-based simulations to guide intermediate 3D representations generated from a single image and input conditions (e.g., applied force and torque), producing…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Xiyang Tan , Ying Jiang , Xuan Li , Zeshun Zong , Tianyi Xie , Yin Yang , Chenfanfu Jiang