English
Related papers

Related papers: Dynamic Visual Reasoning by Learning Differentiabl…

200 papers

We introduce latent intuitive physics, a transfer learning framework for physics simulation that can infer hidden properties of fluids from a single 3D video and simulate the observed fluid in novel scenes. Our key insight is to use latent…

Artificial Intelligence · Computer Science 2024-08-06 Xiangming Zhu , Huayu Deng , Haochen Yuan , Yunbo Wang , Xiaokang Yang

We present a central-peripheral vision-inspired framework (CVP), a simple yet effective multimodal model for spatial reasoning that draws inspiration from the two types of human visual fields -- central vision and peripheral vision.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Zeyuan Chen , Xiang Zhang , Haiyang Xu , Jianwen Xie , Zhuowen Tu

Multimodal Large Language Models (MLLMs) have achieved notable gains in various tasks by incorporating Chain-of-Thought (CoT) reasoning in language spaces. Recent work extends this direction by leveraging external tools for visual editing,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Bangzheng Li , Ximeng Sun , Jiang Liu , Ze Wang , Jialian Wu , Xiaodong Yu , Hao Chen , Emad Barsoum , Muhao Chen , Zicheng Liu

Despite exciting recent results showing vision-language systems' capacity to reason about images using natural language, their capacity for video reasoning remains under-explored. We motivate framing video reasoning as the sequential…

Computation and Language · Computer Science 2023-11-10 Vaishnavi Himakunthala , Andy Ouyang , Daniel Rose , Ryan He , Alex Mei , Yujie Lu , Chinmay Sonar , Michael Saxon , William Yang Wang

Can we learn the physics of matter in motion directly from images and video--and trust it? Answering this question requires integrating experiments, physics-based simulation, and data across traditionally separate disciplines. Much of this…

Computational Engineering, Finance, and Science · Computer Science 2026-04-21 Hagen Holthusen , Kevin Linka , Ellen Kuhl

Video understanding aims to enable models to perceive, reason about, and interact with the dynamic visual world. In contrast to image understanding, video understanding inherently requires modeling temporal dynamics and evolving visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Zhaochong An , Zirui Li , Mingqiao Ye , Feng Qiao , Jiaang Li , Zongwei Wu , Vishal Thengane , Chengzu Li , Lei Li , Luc Van Gool , Guolei Sun , Serge Belongie

The correlation between the vision and text is essential for video moment retrieval (VMR), however, existing methods heavily rely on separate pre-training feature extractors for visual and textual understanding. Without sufficient temporal…

Computer Vision and Pattern Recognition · Computer Science 2023-03-27 Dezhao Luo , Jiabo Huang , Shaogang Gong , Hailin Jin , Yang Liu

Humans are able to make rich predictions about the future dynamics of physical objects from a glance. On the other hand, most existing computer vision approaches require strong assumptions about the underlying system, ad-hoc modeling, or…

Neural and Evolutionary Computing · Computer Science 2018-09-19 Zhihua Wang , Stefano Rosa , Yishu Miao , Zihang Lai , Linhai Xie , Andrew Markham , Niki Trigoni

Deploying multiple machine learning models on resource-constrained robotic platforms for different perception tasks often results in redundant computations, large memory footprints, and complex integration challenges. In response, this work…

Robotics · Computer Science 2025-08-19 Jakub Łucki , Jonathan Becktor , Georgios Georgakis , Rob Royce , Shehryar Khattak

While Large Language Models (LLMs) excel at reasoning on text and Vision-Language Models (VLMs) are highly effective for visual perception, applying those models for visual instruction-based planning remains a widely open problem. In this…

Machine Learning · Computer Science 2025-09-11 Mohamed Salim Aissi , Clemence Grislain , Mohamed Chetouani , Olivier Sigaud , Laure Soulier , Nicolas Thome

Originally designed for applications in computer graphics, visual computing (VC) methods synthesize information about physical and virtual worlds, using prescribed algorithms optimized for spatial computing. VC is used to analyze geometry,…

Recent advances in vision-language reasoning underscore the importance of thinking with images, where models actively ground their reasoning in visual evidence. Yet, prevailing frameworks treat visual actions as optional tools, boosting…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Changpeng Wang , Haozhe Wang , Xi Chen , Junhan Liu , Taofeng Xue , Chong Peng , Donglian Qi , Fangzhen Lin , Yunfeng Yan

Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Chia-Hsiang Kao , Cong Phuoc Huynh , Chien-Yi Wang , Noranart Vesdapunt , Stefan Stojanov , Bharath Hariharan , Oleksandr Obiednikov , Ning Zhou

Learning informative representations from image-based observations is of fundamental concern in deep Reinforcement Learning (RL). However, data-inefficiency remains a significant barrier to this objective. To overcome this obstacle, we…

Machine Learning · Computer Science 2022-01-19 Tao Huang , Jiachen Wang , Xiao Chen

The crux of Referring Video Object Segmentation (RVOS) lies in modeling dense text-video relations to associate abstract linguistic concepts with dynamic visual contents at pixel-level. Current RVOS methods typically use vision and language…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Zikun Zhou , Wentao Xiong , Li Zhou , Xin Li , Zhenyu He , Yaowei Wang

MLLMs have demonstrated significant visual understanding capabilities, yet their fine-grained visual perception in complex real-world scenarios, such as densely crowded public areas, remains limited. Inspired by the recent success of RL in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Sungjune Park , Hyunjun Kim , Junho Kim , Seongho Kim , Yong Man Ro

Distilling interpretable physical laws from videos has led to expanded interest in the computer vision community recently thanks to the advances in deep learning, but still remains a great challenge. This paper introduces an end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2022-05-04 Lele Luan , Yang Liu , Hao Sun

Current visual reasoning methods mainly focus on exploring specific reasoning modes. Although improvements can be achieved in particular domains, they struggle to develop general reasoning capabilities. Inspired by this, we propose a novel…

Artificial Intelligence · Computer Science 2026-05-15 Zejun Li , Yingxiu Zhao , Jiwen Zhang , Siyuan Wang , Yang Yao , Runzhou Zhao , Jun Song , Bo Zheng , Zhongyu Wei

Video understanding is fundamental to tasks such as action recognition, video reasoning, and robotic control. Early video understanding methods based on large vision-language models (LVLMs) typically adopt a single-pass reasoning paradigm…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Yiyang Zhou , Yangfan He , Yaofeng Su , Siwei Han , Joel Jang , Gedas Bertasius , Mohit Bansal , Huaxiu Yao

We propose a novel framework for the task of object-centric video prediction, i.e., extracting the compositional structure of a video sequence, as well as modeling objects dynamics and interactions from visual observations in order to…

Computer Vision and Pattern Recognition · Computer Science 2023-08-01 Angel Villar-Corrales , Ismail Wahdan , Sven Behnke