中文
相关论文

相关论文: CRONOS: Benchmarking Counterfactual Physical Consi…

200 篇论文

Understanding causes and effects in mechanical systems is an essential component of reasoning in the physical world. This work poses a new problem of counterfactual learning of object mechanics from visual input. We develop the CoPhy…

计算机视觉与模式识别 · 计算机科学 2020-04-08 Fabien Baradel , Natalia Neverova , Julien Mille , Greg Mori , Christian Wolf

Video anomaly detection is an essential yet challenging task in the multimedia community, with promising applications in smart cities and secure communities. Existing methods attempt to learn abstract representations of regular events with…

多媒体 · 计算机科学 2023-08-04 Yang Liu , Zhaoyang Xia , Mengyang Zhao , Donglai Wei , Yuzheng Wang , Liu Siao , Bobo Ju , Gaoyun Fang , Jing Liu , Liang Song

Forecasting how 3D medical scans evolve over time is important for disease progression, treatment planning, and developmental assessment. Yet existing models either rely on a single prior scan, fixed grid times, or target global labels,…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Nico Albert Disch , Saikat Roy , Constantin Ulrich , Yannick Kirchhoff , Maximilian Rokuss , Robin Peretzke , David Zimmerer , Klaus Maier-Hein

Generative AI has revolutionised visual content editing, empowering users to effortlessly modify images and videos. However, not all edits are equal. To perform realistic edits in domains such as natural image or medical imaging,…

Recent advances in image and video generation raise hopes that these models possess world modeling capabilities, the ability to generate realistic, physically plausible videos. This could revolutionize applications in robotics, autonomous…

In order to reach human performance on complexvisual tasks, artificial systems need to incorporate a sig-nificant amount of understanding of the world in termsof macroscopic objects, movements, forces, etc. Inspiredby work on intuitive…

Humans are able to perceive, understand and reason about causal events. Developing models with similar physical and causal understanding capabilities is a long-standing goal of artificial intelligence. As a step towards this direction, we…

Learning causal relationships in high-dimensional data (images, videos) is a hard task, as they are often defined on low dimensional manifolds and must be extracted from complex signals dominated by appearance, lighting, textures and also…

计算机视觉与模式识别 · 计算机科学 2022-06-30 Steeven Janny , Fabien Baradel , Natalia Neverova , Madiha Nadri , Greg Mori , Christian Wolf

CERBERUS is a synthetic benchmark designed to help train and evaluate AI models for detecting cracks and other defects in infrastructure. It includes a crack image generator and realistic 3D inspection scenarios built in Unity. The…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Justin Reinman , Sunwoong Choi

Counterfactual reasoning is crucial for robust video understanding but remains underexplored in existing multimodal benchmarks. In this paper, we introduce \textbf{COVER} (\textbf{\underline{CO}}unterfactual…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Qiji Zhou , Yifan Gong , Guangsheng Bao , Hongjie Qiu , Jinqiang Li , Xiangrong Zhu , Huajian Zhang , Yue Zhang

Recent advances in large generative models have greatly enhanced both image editing and in-context image generation, yet a critical gap remains in ensuring physical consistency, where edited objects must remain coherent. This capability is…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Jay Zhangjie Wu , Xuanchi Ren , Tianchang Shen , Tianshi Cao , Kai He , Yifan Lu , Ruiyuan Gao , Enze Xie , Shiyi Lan , Jose M. Alvarez , Jun Gao , Sanja Fidler , Zian Wang , Huan Ling

Video generation models are increasingly used as world simulators for tasks like driving and robotic manipulation. What matters in these settings is not whether a single video looks right, but whether the model's output changes when its…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Kunlin Cai , Rui Song , Jinghuai Zhang , Kaiyuan Zhang , Pranav Bodapati , Alicia Yu , Fnu Suya , Mohammad Rostami , Jiaqi Ma , Yuan Tian

Object-context shortcuts remain a persistent challenge in vision-language models, undermining zero-shot reliability when test-time scenes differ from familiar training co-occurrences. We recast this issue as a causal inference problem and…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Pei Peng , MingKun Xie , Hang Hao , Tong Jin , ShengJun Huang

Perceiving the world in terms of objects and tracking them through time is a crucial prerequisite for reasoning and scene understanding. Recently, several methods have been proposed for unsupervised learning of object-centric…

计算机视觉与模式识别 · 计算机科学 2021-08-18 Marissa A. Weis , Kashyap Chitta , Yash Sharma , Wieland Brendel , Matthias Bethge , Andreas Geiger , Alexander S. Ecker

Causal discovery is at the core of human cognition. It enables us to reason about the environment and make counterfactual predictions about unseen scenarios that can vastly differ from our previous experiences. We consider the task of…

机器学习 · 计算机科学 2020-12-01 Yunzhu Li , Antonio Torralba , Animashree Anandkumar , Dieter Fox , Animesh Garg

We propose an architecture for training generative models of counterfactual conditionals of the form, 'can we modify event A to cause B instead of C?', motivated by applications in robot control. Using an 'adversarial training' paradigm, an…

机器人学 · 计算机科学 2020-09-23 Simón C. Smith , Subramanian Ramamoorthy

Leading approaches in machine vision employ different architectures for different tasks, trained on costly task-specific labeled datasets. This complexity has held back progress in areas, such as robotics, where robust task-general…

计算机视觉与模式识别 · 计算机科学 2023-06-06 Daniel M. Bear , Kevin Feigelis , Honglin Chen , Wanhee Lee , Rahul Venkatesh , Klemen Kotar , Alex Durango , Daniel L. K. Yamins

Recent advances in video generation models demonstrate their potential as world simulators, but they often struggle with videos deviating from physical laws, a key concern overlooked by most text-to-video benchmarks. We introduce a…

计算机视觉与模式识别 · 计算机科学 2025-03-06 Yongfan Chen , Xiuwen Zhu , Tianyu Li

We introduce World Consistency Score (WCS), a novel unified evaluation metric for generative video models that emphasizes internal world consistency of the generated videos. WCS integrates four interpretable sub-components - object…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Akshat Rakheja , Aarsh Ashdhir , Aryan Bhattacharjee , Vanshika Sharma

While current vision algorithms excel at many challenging tasks, it is unclear how well they understand the physical dynamics of real-world environments. Here we introduce Physion, a dataset and benchmark for rigorously evaluating the…

‹ 上一页 1 2 3 10 下一页 ›