English
Related papers

Related papers: PhysicsSolutionAgent: Towards Multimodal Explanati…

200 papers

Simulations are widely used to teach science in grade schools. These simulations are often augmented with a conversational artificial intelligence (AI) agent to provide real-time scaffolding support for students conducting experiments using…

There is a growing interest in applying large language models (LLMs) in robotic tasks, due to their remarkable reasoning ability and extensive knowledge learned from vast training corpora. Grounding LLMs in the physical world remains an…

Robotics · Computer Science 2024-04-11 Wenqiang Lai , Yuan Gao , Tin Lun Lam

Computational fluid dynamics (CFD) has been the main workhorse of computational physics. Yet its steep learning curve and fragmented, multi-stage workflow create significant barriers. To address these challenges, we present Foam-Agent, a…

Artificial Intelligence · Computer Science 2026-03-06 Ling Yue , Nithin Somasekharan , Tingwen Zhang , Yadi Cao , Zhangze Chen , Shimin Di , Shaowu Pan

Vision-Language Models (VLMs) exhibit remarkable common-sense and semantic reasoning capabilities. However, they lack a grounded understanding of physical dynamics. This limitation arises from training VLMs on static internet-scale…

Robotics · Computer Science 2026-04-01 Haowen Liu , Shaoxiong Yao , Haonan Chen , Jiawei Gao , Jiayuan Mao , Jia-Bin Huang , Yilun Du

Large vision-language models (LVLMs) have achieved impressive results in visual question-answering and reasoning tasks through vision instruction tuning on specific datasets. However, there remains significant room for improvement in…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Xiyao Wang , Jiuhai Chen , Zhaoyang Wang , Yuhang Zhou , Yiyang Zhou , Huaxiu Yao , Tianyi Zhou , Tom Goldstein , Parminder Bhatia , Furong Huang , Cao Xiao

World models have emerged as a powerful paradigm for building interactive simulation environments, with recent video-based approaches demonstrating impressive progress in generating visually plausible dynamics. However, because these models…

Artificial Intelligence · Computer Science 2026-05-15 Hongyu Wang , Jingquan Wang , Bocheng Zou , Radu Serban , Dan Negrut

Automated discovery of physical laws from observational data in the real world is a grand challenge in AI. Current methods, relying on symbolic regression or LLMs, are limited to uni-modal data and overlook the rich, visual phenomenological…

Artificial Intelligence · Computer Science 2025-08-26 Jiaqi Liu , Songning Lai , Pengze Li , Di Yu , Wenjie Zhou , Yiyang Zhou , Peng Xia , Zijun Wang , Xi Chen , Shixiang Tang , Lei Bai , Wanli Ouyang , Mingyu Ding , Huaxiu Yao , Aoran Wang

Vision-Language Models (VLMs) have demonstrated strong performance on textbook-style physics problems, yet they frequently fail when confronted with dynamic real-world scenarios that require temporal consistency and causal reasoning across…

Artificial Intelligence · Computer Science 2026-04-28 Sinin Zhang , Yunfei Xie , Yuxuan Cheng , Haoyu Zhang , Tong Zhang

While foundation models (FMs), such as diffusion models and large vision-language models (LVLMs), have been widely applied in educational contexts, their ability to generate pedagogically effective visual explanations remains limited. Most…

Artificial Intelligence · Computer Science 2025-05-29 Haonian Ji , Shi Qiu , Siyang Xin , Siwei Han , Zhaorun Chen , Dake Zhang , Hongyi Wang , Huaxiu Yao

Video agentic models have advanced challenging video-language tasks. However, most agentic approaches still heavily rely on greedy parsing over densely sampled video frames, resulting in high computational cost. We present VideoSeek, a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Jingyang Lin , Jialian Wu , Jiang Liu , Ximeng Sun , Ze Wang , Xiaodong Yu , Jiebo Luo , Zicheng Liu , Emad Barsoum

Video Question Answering is a challenging task, which requires the model to reason over multiple frames and understand the interaction between different objects to answer questions based on the context provided within the video, especially…

Artificial Intelligence · Computer Science 2024-07-31 Bhanu Prakash Reddy Guda , Tanmay Kulkarni , Adithya Sampath , Swarnashree Mysore Sathyendra

Training AI agents to proactively assist humans in daily activities, from routine household tasks to urgent safety situations, requires large-scale visual data. However, capturing such scenarios in the real world is often difficult, costly,…

Computation and Language · Computer Science 2026-05-12 Yu-Hsiang Liu , Yu-Chien Tang , An-Zi Yen

Understanding risk in autonomous driving requires not only perception and prediction, but also high-level reasoning about agent behavior and context. Current Vision Language Model (VLM)-based methods primarily ground agents in static images…

Artificial Intelligence · Computer Science 2026-04-21 Yuan Gao , Mattia Piccinini , Roberto Brusnicki , Yuchen Zhang , Johannes Betz

Remote work and online courses have become important methods of knowledge dissemination, leading to a large number of document-based instructional videos. Unlike traditional video datasets, these videos mainly feature rich-text images and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Haochen Wang , Kai Hu , Liangcai Gao

Computed Tomography (CT) scan, which produces 3D volumetric medical data that can be viewed as hundreds of cross-sectional images (a.k.a. slices), provides detailed anatomical information for diagnosis. For radiologists, creating CT…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Yuren Mao , Wenyi Xu , Yuyang Qin , Yunjun Gao

Deep research has revolutionized data analysis, yet data scientists still devote substantial time to manually crafting visualizations, highlighting the need for robust automation from natural language queries. However, current systems…

Artificial Intelligence · Computer Science 2025-10-06 Zichen Chen , Jiefeng Chen , Sercan Ö. Arik , Misha Sra , Tomas Pfister , Jinsung Yoon

We introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving. The dataset contains 28K unique short video scenarios, and 419K annotations. Evaluating state-of-the-art vision-language models on our…

Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of shortcut solutions based on superficial visual or textual cues.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Benno Krojer , Mojtaba Komeili , Candace Ross , Quentin Garrido , Koustuv Sinha , Nicolas Ballas , Mahmoud Assran

Building robust vision systems for high-stakes domains such as remote sensing requires stronger visual reasoning than what single-pass inference typically provides; yet, retraining large models is often computationally expensive and data…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Chung-En Johnny Yu , Brian Jalaian , Nathaniel D. Bastian

Vision-Language Models (VLMs) enable on-demand visual assistance, yet current applications for people with visual impairments (PVI) impose high cognitive load and exhibit task drift, limiting real-world utility. We first conducted a…

Human-Computer Interaction · Computer Science 2025-11-04 Yi Zhao , Siqi Wang , Qiqun Geng , Erxin Yu , Jing Li