中文
相关论文

相关论文: Interpreting Physics in Video World Models

200 篇论文

Vision Language Action models (VLAs) trained with policy-based reinforcement learning (RL) encode complex behaviors without explicitly modeling environmental dynamics. However, it remains unclear whether VLAs implicitly learn world models,…

机器学习 · 计算机科学 2025-09-30 Marco Molinari , Leonardo Nevali , Saharsha Navani , Omar G. Younis

We propose a new video representation in terms of an over-segmentation of dense trajectories covering the whole video. Trajectories are often used to encode long-temporal information in several computer vision applications. Similar to…

计算机视觉与模式识别 · 计算机科学 2019-01-23 Ijaz Akhter , Cheong Loong Fah , Richard Hartley

Recent advances in large-scale pretraining have yielded visual foundation models with strong capabilities. Not only can recent models generalize to arbitrary images for their training task, their intermediate representations are useful for…

We study pre-training representations for decision-making using video data, which is abundantly available for tasks such as game agents and software testing. Even though significant empirical advances have been made on this problem, a…

机器学习 · 计算机科学 2024-03-21 Dipendra Misra , Akanksha Saran , Tengyang Xie , Alex Lamb , John Langford

Building on the momentum of image generation diffusion models, there is an increasing interest in video-based diffusion models. However, video generation poses greater challenges due to its higher-dimensional nature, the scarcity of…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Aimon Rahman , Malsha V. Perera , Vishal M. Patel

World models have made significant progress in modeling dynamic environments; however, most embodied world models are still restricted to 2D representations, lacking the comprehensive multi-view information essential for embodied spatial…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Peiyan Tu , Hanxin Zhu , Jingwen Sun , Shaojie Ren , Cong Wang , Jiayi Luo , Xiaoqian Cheng , Zhibo Chen

Perceiving the shape and material of an object from a single image is inherently ambiguous, especially when lighting is unknown and unconstrained. Despite this, humans can often disentangle shape and material, and when they are uncertain,…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Xinran Nicole Han , Ko Nishino , Todd Zickler

Recent advancements in diffusion frameworks have significantly enhanced video editing, achieving high fidelity and strong alignment with textual prompts. However, conventional approaches using image diffusion models fall short in handling…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Yixuan Zhu , Haolin Wang , Shilin Ma , Wenliang Zhao , Yansong Tang , Lei Chen , Jie Zhou

Video compression has always been a popular research area, where many traditional and deep video compression methods have been proposed. These methods typically rely on signal prediction theory to enhance compression performance by…

计算机视觉与模式识别 · 计算机科学 2023-08-31 Lv Tang , Xinfeng Zhang , Gai Zhang , Xiaoqi Ma

Predicting future frames of a video sequence has been a problem of high interest in the field of Computer Vision as it caters to a multitude of applications. The ability to predict, anticipate and reason about future events is the essence…

计算机视觉与模式识别 · 计算机科学 2020-09-04 Jasmeen Kaur , Sukhendu Das

Both text and video data are abundant on the internet and support large-scale self-supervised learning through next token or frame prediction. However, they have not been equally leveraged: language models have had significant real-world…

计算机视觉与模式识别 · 计算机科学 2024-02-28 Sherry Yang , Jacob Walker , Jack Parker-Holder , Yilun Du , Jake Bruce , Andre Barreto , Pieter Abbeel , Dale Schuurmans

World models have recently re-emerged as a central paradigm for embodied intelligence, robotics, autonomous driving, and model-based reinforcement learning. However, current world model research is often dominated by three partially…

人工智能 · 计算机科学 2026-05-27 Sen Cui , Jingheng Ma

We present a novel approach to estimating physical properties of objects from video. Our approach consists of a physics engine and a correction estimator. Starting from the initial observed state, object behavior is simulated forward in…

计算机视觉与模式识别 · 计算机科学 2022-06-03 Martin Link , Max Schwarz , Sven Behnke

In this paper we present a discussion of the basic aspects of the well-known problem of prediction and inference in physics, with specific attention to the role of models, the use of data and the application of recent developments in…

综合物理 · 物理学 2024-10-07 Luca Gammaitoni , Angelo Vulpiani

Imitation learning has emerged as a promising approach towards building generalist robots. However, scaling imitation learning for large robot foundation models remains challenging due to its reliance on high-quality expert demonstrations.…

机器人学 · 计算机科学 2025-05-26 Chuning Zhu , Raymond Yu , Siyuan Feng , Benjamin Burchfiel , Paarth Shah , Abhishek Gupta

We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language commands and images,…

Psychovisual models suggest human vision decouples low-level feature extraction from higher cognition by first forming intermediate abstractions. In contrast, deep learning-based vision models routinely extract and aggregate features using…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Wendi Ma , Aryaman Sharma , Wei Dai , Shekhar S. Chandra

Game-Based Learning has proven to be an effective method for enhancing engagement with educational material. However, gaining a deeper understanding of player strategies remains challenging. Sequential game-state and action-based tracking…

人机交互 · 计算机科学 2025-07-03 Braden Roper , William Thompson , Chris Weaver

A longstanding goal in computer vision is to model motions from videos, while the representations behind motions, i.e. the invisible physical interactions that cause objects to deform and move, remain largely unexplored. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Zhiyuan Gao , Jiageng Mao , Hong-Xing Yu , Haozhe Lou , Emily Yue-Ting Jia , Jernej Barbic , Jiajun Wu , Yue Wang

Video understanding represents the most challenging frontier in computer vision, requiring models to reason about complex spatiotemporal relationships, long-term dependencies, and multimodal evidence. The recent emergence of Video-Large…