中文
相关论文

相关论文: ContPhy: Continuum Physical Concept Learning and R…

200 篇论文

Video prediction models based on convolutional networks, recurrent networks, and their combinations often result in blurry predictions. We identify an important contributing factor for imprecise predictions that has not been studied…

计算机视觉与模式识别 · 计算机科学 2018-09-11 Wonmin Byeon , Qin Wang , Rupesh Kumar Srivastava , Petros Koumoutsakos

Machine comprehension of texts longer than a single sentence often requires coreference resolution. However, most current reading comprehension benchmarks do not contain complex coreferential phenomena and hence fail to evaluate the ability…

计算与语言 · 计算机科学 2019-09-06 Pradeep Dasigi , Nelson F. Liu , Ana Marasović , Noah A. Smith , Matt Gardner

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able…

What is the right way to reason about human activities? What directions forward are most promising? In this work, we analyze the current state of human activity understanding in videos. The goal of this paper is to examine datasets,…

计算机视觉与模式识别 · 计算机科学 2017-08-10 Gunnar A. Sigurdsson , Olga Russakovsky , Abhinav Gupta

Human cognition is deeply intertwined with a sense of time, known as Chronoception. This sense allows us to judge how long facts remain valid and when knowledge becomes outdated. Despite progress in vision, language, and motor control, AI…

计算与语言 · 计算机科学 2025-05-13 Krish Goel , Sanskar Pandey , KS Mahadevan , Harsh Kumar , Vishesh Khadaria

Understanding relations between objects is crucial for understanding the semantics of a visual scene. It is also an essential step in order to bridge visual and language models. However, current state-of-the-art computer vision models still…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Palaash Agrawal , Haidi Azaman , Cheston Tan

With the rapid development of large multimodal models, reliable judge and critic models have become essential for open-ended evaluation and preference alignment, providing pairwise preferences, numerical scores, and explanatory…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Tianyi Xiong , Shihao Wang , Guilin Liu , Yi Dong , Ming Li , Heng Huang , Jan Kautz , Zhiding Yu

Video generation techniques have achieved remarkable advancements in visual quality, yet faithfully reproducing real-world physics remains elusive. Preference-based model post-training may improve physical consistency, but requires costly…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Wenxu Qian , Chaoyue Wang , Hou Peng , Zhiyu Tan , Hao Li , Anxiang Zeng

Current evaluation protocols predominantly assess physical reasoning in stationary scenes, creating a gap in evaluating agents' abilities to interact with dynamic events. While contemporary methods allow agents to modify initial scene…

人工智能 · 计算机科学 2024-03-26 Shiqian Li , Kewen Wu , Chi Zhang , Yixin Zhu

We introduce a benchmark to evaluate the capability of AI to solve problems in theoretical physics, focusing on high-energy theory and cosmology. The first iteration of our benchmark consists of 57 problems of varying difficulty, from…

What does it mean for two videos to be similar? Videos may appear similar when judged by the actions they depict, yet entirely different if evaluated based on the locations where they were filmed. While humans naturally compare videos by…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Benedetta Liberatori , Alessandro Conti , Lorenzo Vaquero , Yiming Wang , Elisa Ricci , Paolo Rota

Multimodal reasoning is a critical component in the pursuit of artificial intelligence systems that exhibit human-like intelligence, especially when tackling complex tasks. While the chain-of-thought (CoT) technique has gained considerable…

人工智能 · 计算机科学 2023-09-26 Jingxuan Wei , Cheng Tan , Zhangyang Gao , Linzhuang Sun , Siyuan Li , Bihui Yu , Ruifeng Guo , Stan Z. Li

Physical dynamical systems can be viewed as natural information processors: their systems preserve, transform, and disperse input information. This perspective motivates learning not only from data generated by such systems, but also how to…

机器学习 · 计算机科学 2026-03-05 Felix Köster , Atsushi Uchida

To reach human performance on complex tasks, a key ability for artificial systems is to understand physical interactions between objects, and predict future outcomes of a situation. This ability, often referred to as intuitive physics, has…

计算机视觉与模式识别 · 计算机科学 2020-05-04 Ronan Riochet , Josef Sivic , Ivan Laptev , Emmanuel Dupoux

Human motion synthesis is an important problem with applications in graphics, gaming and simulation environments for robotics. Existing methods require accurate motion capture data for training, which is costly to obtain. Instead, we…

计算机视觉与模式识别 · 计算机科学 2022-08-15 Kevin Xie , Tingwu Wang , Umar Iqbal , Yunrong Guo , Sanja Fidler , Florian Shkurti

Physically Plausible Video Generation (PPVG) has emerged as a promising avenue for modeling real-world physical phenomena. PPVG requires an understanding of commonsense knowledge, which remains a challenge for video diffusion models.…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Zixuan Wang , Yixin Hu , Haolan Wang , Feng Chen , Yan Liu , Wen Li , Yinjie Lei

Recent studies have shown that long chain-of-thought (CoT) reasoning can significantly enhance the performance of large language models (LLMs) on complex tasks. However, this benefit is yet to be demonstrated in the domain of video…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Yuanxin Liu , Kun Ouyang , Haoning Wu , Yi Liu , Lin Sui , Xinhao Li , Yan Zhong , Y. Charles , Xinyu Zhou , Xu Sun

Commonsense reasoning, the ability to make logical assumptions about daily scenes, is one core intelligence of human beings. In this work, we present a novel task and dataset for evaluating the ability of text-to-image generative models to…

多媒体 · 计算机科学 2024-01-24 Mianzhi Pan , Jianfei Li , Mingyue Yu , Zheng Ma , Kanzhi Cheng , Jianbing Zhang , Jiajun Chen

While multimodal large language models excel at tasks that integrate visual perception with symbolic reasoning, their performance is often undermined by a critical vulnerability: perception-induced errors that propagate through the…

With the continuous advancement of large language models (LLMs), it is essential to create new benchmarks to effectively evaluate their expanding capabilities and identify areas for improvement. This work focuses on multi-image reasoning,…