English
Related papers

Related papers: TUMTraffic-VideoQA: A Benchmark for Unified Spatio…

200 papers

Multimodal information, together with our knowledge, help us to understand the complex and dynamic world. Large language models (LLM) and large multimodal models (LMM), however, still struggle to emulate this capability. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Yuanhan Zhang , Kaichen Zhang , Bo Li , Fanyi Pu , Christopher Arif Setiadharma , Jingkang Yang , Ziwei Liu

Recent advancements in multimodal large language models (MLLMs) have shown strong understanding of driving scenes, drawing interest in their application to autonomous driving. However, high-level reasoning in safety-critical scenarios,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Seungjun Yu , Seonho Lee , Namho Kim , Jaeyo Shin , Junsung Park , Wonjeong Ryu , Raehyuk Jung , Hyunjung Shim

Video Question Answering (VideoQA) is a complex video-language task that demands a sophisticated understanding of both visual content and temporal dynamics. Traditional Transformer-style architectures, while effective in integrating…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Zijie Song , Zhenzhen Hu , Yixiao Ma , Jia Li , Richang Hong

Traffic data imputation is a critical preprocessing step in intelligent transportation systems, underpinning the reliability of downstream transportation services. Despite substantial progress in imputation models, model selection and…

Machine Learning · Computer Science 2025-10-21 Shengnan Guo , Tonglong Wei , Yiheng Huang , Yan Lin , Zekai Shen , Yujuan Dong , Junliang Lin , Youfang Lin , Huaiyu Wan

Visual events are a composition of temporal actions involving actors spatially interacting with objects. When developing computer vision models that can reason about compositional spatio-temporal events, we need benchmarks that can analyze…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Madeleine Grunde-McLaughlin , Ranjay Krishna , Maneesh Agrawala

Reasoning about causal and temporal event relations in videos is a new destination of Video Question Answering (VideoQA).The major stumbling block to achieve this purpose is the semantic gap between language and video since they are at…

Computer Vision and Pattern Recognition · Computer Science 2022-11-03 Shaoning Xiao , Long Chen , Kaifeng Gao , Zhao Wang , Yi Yang , Zhimeng Zhang , Jun Xiao

This paper strives to solve complex video question answering (VideoQA) which features long video containing multiple objects and events at different time. To tackle the challenge, we highlight the importance of identifying question-critical…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Yicong Li , Junbin Xiao , Chun Feng , Xiang Wang , Tat-Seng Chua

The rise of Visual-Language Models (LVLMs) has unlocked new possibilities for seamlessly integrating visual and textual information. However, their ability to interpret cartographic maps remains largely unexplored. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Huy Quang Ung , Guillaume Habault , Yasutaka Nishimura , Hao Niu , Roberto Legaspi , Tomoki Oya , Ryoichi Kojima , Masato Taya , Chihiro Ono , Atsunori Minamikawa , Yan Liu

Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benchmarks mostly resemble static image benchmarks and are…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Mu Cai , Reuben Tan , Jianrui Zhang , Bocheng Zou , Kai Zhang , Feng Yao , Fangrui Zhu , Jing Gu , Yiwu Zhong , Yuzhang Shang , Yao Dou , Jaden Park , Jianfeng Gao , Yong Jae Lee , Jianwei Yang

Time series data are integral to critical applications across domains such as finance, healthcare, transportation, and environmental science. While recent work has begun to explore multi-task time series question answering (QA), current…

Videos are unique in their integration of temporal elements, including camera, scene, action, and attribute, along with their dynamic relationships over time. However, existing benchmarks for video understanding often treat these properties…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Fanheng Kong , Jingyuan Zhang , Hongzhi Zhang , Shi Feng , Daling Wang , Linhao Yu , Xingguang Ji , Yu Tian , Victoria W. , Fuzheng Zhang

Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills. We introduce FlowVQA, a novel benchmark aimed at assessing the capabilities of visual…

Computation and Language · Computer Science 2024-07-01 Shubhankar Singh , Purvi Chaurasia , Yerram Varun , Pranshu Pandya , Vatsal Gupta , Vivek Gupta , Dan Roth

The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2025. Its primary goal is to benchmark state-of-the-art video models and measure the progress…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Joseph Heyward , Nikhil Parthasarathy , Tyler Zhu , Aravindh Mahendran , João Carreira , Dima Damen , Andrew Zisserman , Viorica Pătrăucean

Situation awareness is essential for understanding and reasoning about 3D scenes in embodied AI agents. However, existing datasets and benchmarks for situated understanding are limited in data modality, diversity, scale, and task scope. To…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Xiongkun Linghu , Jiangyong Huang , Xuesong Niu , Xiaojian Ma , Baoxiong Jia , Siyuan Huang

Understanding accurate atomic temporal event is essential for video comprehension. However, current video-language benchmarks often fall short to evaluate Large Multi-modal Models' (LMMs) temporal event understanding capabilities, as they…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Yuqi Liu , Qin Jin , Tianyuan Qu , Xuan Liu , Yang Du , Bei Yu , Jiaya Jia

Recent advancements in Vision-Language Models (VLMs) have demonstrated strong potential for autonomous driving tasks. However, their spatial understanding and reasoning-key capabilities for autonomous driving-still exhibit significant…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Kexin Tian , Jingrui Mao , Yunlong Zhang , Jiwan Jiang , Yang Zhou , Zhengzhong Tu

We present Scene-Graph Based Multi-Modal Traffic Agent (SGTA), a modular framework for traffic video understanding that combines structured scene graphs with multi-modal reasoning. It constructs a traffic scene graph from roadside videos…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Xingcheng Zhou , Mingyu Liu , Walter Zimmer , Jiajie Zhang , Alois Knoll

With the rapid development of multimedia processing and deep learning technologies, especially in the field of video understanding, video quality assessment (VQA) has achieved significant progress. Although researchers have moved from…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Jiebin Yan , Lei Wu , Yuming Fang , Xuelin Liu , Xue Xia , Weide Liu

Multimodal systems have great potential to assist humans in procedural activities, where people follow instructions to achieve their goals. Despite diverse application scenarios, systems are typically evaluated on traditional classification…

Computation and Language · Computer Science 2025-11-05 Kimihiro Hasegawa , Wiradee Imrattanatrai , Zhi-Qi Cheng , Masaki Asada , Susan Holm , Yuran Wang , Ken Fukuda , Teruko Mitamura

Cooperative autonomous driving requires traffic scene understanding from both vehicle and infrastructure perspectives. While vision-language models (VLMs) show strong general reasoning capabilities, their performance in safety-critical…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Rui Gan , Junyi Ma , Pei Li , Xingyou Yang , Kai Chen , Sikai Chen , Bin Ran