English
Related papers

Related papers: ContPhy: Continuum Physical Concept Learning and R…

200 papers

Causality knowledge is vital to building robust AI systems. Deep learning models often perform poorly on tasks that require causal reasoning, which is often derived using some form of commonsense knowledge not immediately available in the…

Computer Vision and Pattern Recognition · Computer Science 2021-07-23 Aman Chadha , Vinija Jain

AI video generation is undergoing a revolution, with quality and realism advancing rapidly. These advances have led to a passionate scientific debate: Do video models learn "world models" that discover laws of physics -- or, alternatively,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Saman Motamed , Laura Culp , Kevin Swersky , Priyank Jaini , Robert Geirhos

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual…

Recent progress in text-to-video (T2V) generation has enabled the synthesis of visually compelling and temporally coherent videos from natural language. However, these models often fall short in basic physical commonsense, producing outputs…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Enes Sanli , Baris Sarper Tezcan , Aykut Erdem , Erkut Erdem

In this study, we create a CConS (Counter-commonsense Contextual Size comparison) dataset to investigate how physical commonsense affects the contextualized size comparison task; the proposed dataset consists of both contexts that fit…

Computation and Language · Computer Science 2023-06-06 Kazushi Kondo , Saku Sugawara , Akiko Aizawa

A fundamental component of human vision is our ability to parse complex visual scenes and judge the relations between their constituent objects. AI benchmarks for visual reasoning have driven rapid progress in recent years with…

Computer Vision and Pattern Recognition · Computer Science 2022-06-14 Aimen Zerroug , Mohit Vaishnav , Julien Colin , Sebastian Musslick , Thomas Serre

Video prediction is increasingly viewed as a path toward generalizable world models, yet it remains unclear whether these systems learn underlying causal structure or merely exploit superficial visual correlations for future prediction. We…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 León Begiristain , Olaf Dünkel , Adam Kortylewski

Recent advances in generative foundational models, often termed "world models," have propelled interest in applying them to critical tasks like robotic planning and autonomous system training. For reliable deployment, these models must…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Rishi Upadhyay , Howard Zhang , Jim Solomon , Ayush Agrawal , Pranay Boreddy , Shruti Satya Narayana , Yunhao Ba , Alex Wong , Celso M de Melo , Achuta Kadambi

As humans, we can modify our assumptions about a scene by imagining alternative objects or concepts in our minds. For example, we can easily anticipate the implications of the sun being overcast by rain clouds (e.g., the street will get…

Computation and Language · Computer Science 2022-07-11 Hyounghun Kim , Abhay Zala , Mohit Bansal

Many real-world tasks require an agent to reason jointly over text and visual objects, (e.g., navigating in public spaces), which we refer to as context-sensitive text-rich visual reasoning. Specifically, these tasks require an…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Rohan Wadhawan , Hritik Bansal , Kai-Wei Chang , Nanyun Peng

We introduce PointOdyssey, a large-scale synthetic dataset, and data generation framework, for the training and evaluation of long-term fine-grained tracking algorithms. Our goal is to advance the state-of-the-art by placing emphasis on…

Computer Vision and Pattern Recognition · Computer Science 2023-07-28 Yang Zheng , Adam W. Harley , Bokui Shen , Gordon Wetzstein , Leonidas J. Guibas

A vexing problem in artificial intelligence is reasoning about events that occur in complex, changing visual stimuli such as in video analysis or game play. Inspired by a rich tradition of visual reasoning and memory in cognitive psychology…

Artificial Intelligence · Computer Science 2018-07-23 Guangyu Robert Yang , Igor Ganichev , Xiao-Jing Wang , Jonathon Shlens , David Sussillo

Multimodal reasoning remains a fundamental challenge in artificial intelligence. Despite substantial advances in text-based reasoning, even state-of-the-art models such as GPT-o3 struggle to maintain strong performance in multimodal…

Computation and Language · Computer Science 2025-09-09 Hao Liang , Ruitao Wu , Bohan Zeng , Junbo Niu , Wentao Zhang , Bin Dong

Humans demonstrate remarkable abilities to predict physical events in complex scenes. Two classes of models for physical scene understanding have recently been proposed: "Intuitive Physics Engines", or IPEs, which posit that people make…

Artificial Intelligence · Computer Science 2016-10-05 Renqiao Zhang , Jiajun Wu , Chengkai Zhang , William T. Freeman , Joshua B. Tenenbaum

Recent advances in text-to-video generation have achieved impressive perceptual quality, yet generated content often violates fundamental principles of physical plausibility - manifesting as implausible object dynamics, incoherent…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Peiyao Wang , Weining Wang , Qi Li

While current methods have shown promising progress on estimating 3D human motion from monocular videos, their motion estimates are often physically unrealistic because they mainly consider kinematics. In this paper, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Yufei Zhang , Jeffrey O. Kephart , Zijun Cui , Qiang Ji

We introduce SeePhys Pro, a fine-grained modality transfer benchmark that studies whether models preserve the same reasoning capability when critical information is progressively transferred from text to image. Unlike standard…

Accurately predicting fluid dynamics and evolution has been a long-standing challenge in physical sciences. Conventional deep learning methods often rely on the nonlinear modeling capabilities of neural networks to establish mappings…

Machine Learning · Computer Science 2025-04-09 Huaguan Chen , Yang Liu , Hao Sun

Visual Commonsense Reasoning (VCR) predicts an answer with corresponding rationale, given a question-image input. VCR is a recently introduced visual scene understanding task with a wide range of applications, including visual question…

Computer Vision and Pattern Recognition · Computer Science 2023-12-11 Xuejiao Tang , Xin Huang , Wenbin Zhang , Travers B. Child , Qiong Hu , Zhen Liu , Ji Zhang

Recent advances in image and video generation raise hopes that these models possess world modeling capabilities, the ability to generate realistic, physically plausible videos. This could revolutionize applications in robotics, autonomous…