English
Related papers

Related papers: LikePhys: Evaluating Intuitive Physics Understandi…

200 papers

Diffusion models have demonstrated strong generative capabilities across scientific domains, but often produce outputs that violate physical laws. We propose a new perspective by framing physics-informed generation as a sparse reward…

Machine Learning · Computer Science 2025-09-26 Mingze Yuan , Pengfei Jin , Na Li , Quanzheng Li

Learning rewards from expert videos offers an affordable and effective solution to specify the intended behaviors for reinforcement learning (RL) tasks. In this work, we propose Diffusion Reward, a novel framework that learns rewards from…

Machine Learning · Computer Science 2024-08-12 Tao Huang , Guangqi Jiang , Yanjie Ze , Huazhe Xu

World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. However, existing world models often require extensive…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Siqiao Huang , Jialong Wu , Qixing Zhou , Shangchen Miao , Mingsheng Long

The ability to predict future states is crucial to informed decision-making while interacting with dynamic environments. With cameras providing a prevalent and information-rich sensing modality, the problem of predicting future states from…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Nikhil U. Shinde , Xiao Liang , Florian Richter , Michael C. Yip

We propose Inner Loop Feedback (ILF), a novel approach to accelerate diffusion models' inference. ILF trains a lightweight module to predict future features in the denoising process by leveraging the outputs from a chosen diffusion backbone…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Matthew Gwilliam , Han Cai , Di Wu , Abhinav Shrivastava , Zhiyu Cheng

Understanding the 3D world from 2D images involves more than detection and segmentation of the objects within the scene. It also includes the interpretation of the structure and arrangement of the scene elements. Such understanding is often…

Computer Vision and Pattern Recognition · Computer Science 2022-12-07 Or Nachmias , Ohad Fried , Ariel Shamir

In this paper, we teach a machine to discover the laws of physics from video streams. We assume no prior knowledge of physics, beyond a temporal stream of bounding boxes. The problem is very difficult because a machine must learn not only a…

Computer Vision and Pattern Recognition · Computer Science 2019-11-28 Pradyumna Chari , Chinmay Talegaonkar , Yunhao Ba , Achuta Kadambi

Video-based representations have gained prominence in planning and decision-making due to their ability to encode rich spatiotemporal dynamics and geometric relationships. These representations enable flexible and generalizable solutions…

Robotics · Computer Science 2026-02-11 Po-Chen Ko , Jiayuan Mao , Yu-Hsiang Fu , Hsien-Jeng Yeh , Chu-Rong Chen , Wei-Chiu Ma , Yilun Du , Shao-Hua Sun

Autonomous driving requires robust perception models trained on high-quality, large-scale multi-view driving videos for tasks like 3D object detection, segmentation and trajectory prediction. While world models provide a cost-effective…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Zhuoran Yang , Xi Guo , Chenjing Ding , Chiyu Wang , Wei Wu

Fairness-awareness has emerged as an essential building block for the responsible use of artificial intelligence in real applications. In many cases, inequity in performance is due to the change in distribution over different regions. While…

Machine Learning · Computer Science 2024-02-07 Zhihao Wang , Yiqun Xie , Zhili Li , Xiaowei Jia , Zhe Jiang , Aolin Jia , Shuo Xu

Discovering physical laws directly from high-dimensional visual data is a long-standing human pursuit but remains a formidable challenge for machines, representing a fundamental goal of scientific intelligence. This task is inherently…

Computational Engineering, Finance, and Science · Computer Science 2026-02-24 Ruikun Li , Jun Yao , Yingfan Hua , Shixiang Tang , Biqing Qi , Bin Liu , Wanli Ouyang , Yan Lu

The ability to discover abstract physical concepts and understand how they work in the world through observing lies at the core of human intelligence. The acquisition of this ability is based on compositionally perceiving the environment in…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Qu Tang , XiangYu Zhu , Zhen Lei , ZhaoXiang Zhang

Video world models learn representations of physical dynamics, but controlling their physical expectations at inference time remains an open problem. Recent interpretability work identified a Physics Emergence Zone (PEZ), a group of middle…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Nahid Alam

In recent years, vision-language models (VLMs) have shown remarkable performance on visual reasoning tasks (e.g. attributes, location). While such tasks measure the requisite knowledge to ground and reason over a given visual instance, they…

Computation and Language · Computer Science 2022-09-16 Shikhar Singh , Ehsan Qasemi , Muhao Chen

We introduce SeePhys Pro, a fine-grained modality transfer benchmark that studies whether models preserve the same reasoning capability when critical information is progressively transferred from text to image. Unlike standard…

In this work, we discuss evaluating video foundation models in a fair and robust manner. Unlike language or image foundation models, many video foundation models are evaluated with differing parameters (such as sampling rate, number of…

Immersive video offers the freedom to navigate inside virtualized environment. Instead of streaming the bulky immersive videos entirely, a viewport (also referred to as field of view, FoV) adaptive streaming is preferred. We often stream…

Multimedia · Computer Science 2018-02-19 Shaowei Xie , Qiu Shen , Yiling Xu , Qiaojian Qian , Shaowei Wang , Zhan Ma , Wenjun Zhang

While deep convolutional neural networks frequently approach or exceed human-level performance at benchmark tasks involving static images, extending this success to moving images is not straightforward. Having models which can learn to…

Computer Vision and Pattern Recognition · Computer Science 2017-02-07 Tegan Maharaj , Nicolas Ballas , Anna Rohrbach , Aaron Courville , Christopher Pal

We introduce MultiPhys, a method designed for recovering multi-person motion from monocular videos. Our focus lies in capturing coherent spatial placement between pairs of individuals across varying degrees of engagement. MultiPhys, being…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Nicolas Ugrinovic , Boxiao Pan , Georgios Pavlakos , Despoina Paschalidou , Bokui Shen , Jordi Sanchez-Riera , Francesc Moreno-Noguer , Leonidas Guibas

Recent advancements in text-to-video (T2V) diffusion models have enabled high-fidelity and realistic video synthesis. However, current T2V models often struggle to generate physically plausible content due to their limited inherent ability…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Xiangdong Zhang , Jiaqi Liao , Shaofeng Zhang , Fanqing Meng , Xiangpeng Wan , Junchi Yan , Yu Cheng