中文
相关论文

相关论文: CRONOS: Benchmarking Counterfactual Physical Consi…

200 篇论文

The ability to understand physical dynamics is critical for agents to act in the world. Here, we use Counterfactual World Modeling (CWM) to extract vision structures for dynamics understanding. CWM uses a temporally-factored masking policy…

The assessment of evaluation metrics (meta-evaluation) is crucial for determining the suitability of existing metrics in text-to-image (T2I) generation tasks. Human-based meta-evaluation is costly and time-intensive, and automated…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Christoph Leiter , Yuki M. Asano , Margret Keuper , Steffen Eger

Multimodal large language models (MLLMs) achieve strong performance on single-view spatial reasoning tasks, yet it remains unclear whether they maintain stable spatial state representations under counterfactual viewpoint changes. We…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Shanmukha Vellamcheti , Uday Kiran Kothapalli , Disharee Bhowmick , Sathyanarayanan N. Aakur

We introduce the Continuum Physical Dataset (ContPhy), a novel benchmark for assessing machine physical commonsense. ContPhy complements existing physical reasoning benchmarks by encompassing the inference of diverse physical properties,…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Zhicheng Zheng , Xin Yan , Zhenfang Chen , Jingzhou Wang , Qin Zhi Eddie Lim , Joshua B. Tenenbaum , Chuang Gan

CounterFactual (CF) visual explanations try to find images similar to the query image that change the decision of a vision system to a specified outcome. Existing methods either require inference-time optimization or joint training with a…

计算机视觉与模式识别 · 计算机科学 2022-03-30 Saeed Khorram , Li Fuxin

Generative world models are reshaping embodied AI, enabling agents to synthesize realistic 4D driving environments that look convincing but often fail physically or behaviorally. Despite rapid progress, the field still lacks a unified way…

Vision-Language Models (VLMs) excel at multimodal reasoning, yet it remains unclear whether their answers are grounded in visual evidence or driven by learned language and world priors. Counting provides a precise testbed: when visual…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Reem Alzahrani , Hassan Alshanqiti , Bushra Bin Hemid , Zaid Alyafeai , Abdelrahman Eldesokey , Bernard Ghanem

Video large language models (Video-LLMs) have made strong progress in general video understanding, but their ability to maintain temporal object consistency remains underexplored. Existing benchmarks often emphasize event recognition,…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Junzhe Chen , Siyuan Meng , Yuxi Chen , Man Zhao , Wenyao Gui , Xiaojie Guo

By providing substantial amounts of data and standardized evaluation protocols, datasets in computer vision have helped fuel advances across all areas of visual recognition. But even in light of breakthrough results on recent benchmarks, it…

计算机视觉与模式识别 · 计算机科学 2018-07-06 Brandon RichardWebster , Samuel E. Anthony , Walter J. Scheirer

Instructional videos are the dominant medium for learning physical tasks, yet they rarely match the user's real-world visual context. Motor simulation and cognitive load theories predict this mismatch should matter, but we do not know (1)…

人机交互 · 计算机科学 2026-05-19 Yayuan Li , Chenglin Li , Jingying Wang , Filippos Bellos , Anhong Guo , Jason J. Corso

Causal Representation Learning (CRL) aims to uncover the data-generating process and identify the underlying causal variables and relations, whose evaluation remains inherently challenging due to the requirement of known ground-truth causal…

机器学习 · 计算机科学 2025-10-20 Guangyi Chen , Yunlong Deng , Peiyuan Zhu , Yan Li , Yifan Shen , Zijian Li , Kun Zhang

Contextual information plays an important role in many computer vision tasks, such as object detection, video action detection, image classification, etc. Recognizing a single object or action out of context could be sometimes very…

计算机视觉与模式识别 · 计算机科学 2023-02-13 Xuan Wang , Zhigang Zhu

Contradictory multimodal inputs are common in real-world settings, yet existing benchmarks typically assume input consistency and fail to evaluate cross-modal contradiction detection - a fundamental capability for preventing hallucinations…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Teodora Popordanoska , Jiameng Li , Matthew B. Blaschko

Image captioning, which generates natural language descriptions of the visual information in an image, is a crucial task in vision-language research. Previous models have typically addressed this task by aligning the generative capabilities…

计算机视觉与模式识别 · 计算机科学 2024-09-02 Qian Cao , Xu Chen , Ruihua Song , Xiting Wang , Xinting Huang , Yuchen Ren

Multimodal large language models often struggle with faithful reasoning in complex visual scenes, where intricate entities and relations require precise visual grounding at each step. This reasoning unfaithfulness frequently manifests as…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Chuhan Wang , Xintong Li , Jennifer Yuntong Zhang , Junda Wu , Chengkai Huang , Lina Yao , Julian McAuley , Jingbo Shang

The ability to reason about temporal and causal events from videos lies at the core of human intelligence. Most video reasoning benchmarks, however, focus on pattern recognition from complex visual and language input, instead of on causal…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Kexin Yi , Chuang Gan , Yunzhu Li , Pushmeet Kohli , Jiajun Wu , Antonio Torralba , Joshua B. Tenenbaum

Visual counterfactual explanations identify modifications to an image that would change the prediction of a classifier. We propose a set of techniques based on generative models (VAE) and a classifier ensemble directly trained in the latent…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Claire Theobald , Frédéric Pennerath , Brieuc Conan-Guez , Miguel Couceiro , Amedeo Napoli

This presentation introduces a self-supervised learning approach to the synthesis of new video clips from old ones, with several new key elements for improved spatial resolution and realism: It conditions the synthesis process on contextual…

计算机视觉与模式识别 · 计算机科学 2021-10-27 Guillaume Le Moing , Jean Ponce , Cordelia Schmid

This work introduces a framework to diagnose the strengths and shortcomings of Autonomous Vehicle (AV) collision avoidance technology with synthetic yet realistic potential collision scenarios adapted from real-world, collision-free data.…

最优化与控制 · 数学 2024-09-18 Robert Dyro , Matthew Foutter , Ruolin Li , Luigi Di Lillo , Edward Schmerling , Xilin Zhou , Marco Pavone

There has been a recent surge in research on adversarial perturbations that defeat Deep Neural Networks (DNNs) in machine vision; most of these perturbation-based attacks target object classifiers. Inspired by the observation that humans…

计算机视觉与模式识别 · 计算机科学 2020-07-27 Shasha Li , Shitong Zhu , Sudipta Paul , Amit Roy-Chowdhury , Chengyu Song , Srikanth Krishnamurthy , Ananthram Swami , Kevin S Chan