English
Related papers

Related papers: From Segments to Scenes: Temporal Understanding in…

200 papers

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current…

Multimedia · Computer Science 2025-05-29 Yong Ren , Chenxing Li , Le Xu , Hao Gu , Duzhen Zhang , Yujie Chen , Manjie Xu , Ruibo Fu , Shan Yang , Dong Yu

Multimodal Large Language Models (MLLMs) have achieved significant advancements in tasks like Visual Question Answering (VQA) by leveraging foundational Large Language Models (LLMs). However, their abilities in specific areas such as visual…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Mohamed Fazli Imam , Chenyang Lyu , Alham Fikri Aji

The integration of Vision-Language Models (VLMs) into autonomous driving systems has shown promise in addressing key challenges such as learning complexity, interpretability, and common-sense reasoning. However, existing approaches often…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Xuesong Chen , Linjiang Huang , Tao Ma , Rongyao Fang , Shaoshuai Shi , Hongsheng Li

TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-world scenarios due to…

Multimedia · Computer Science 2026-03-20 Xinyuan Qian , Xinjia Zhu , Alessio Brutti , Dong Liang

While autonomous driving (AD) stacks struggle with decision making under partial observability and real-world complexity, human drivers are capable of applying commonsense reasoning to make near-optimal decisions with limited information.…

Time series anomaly detection (TSAD) is essential for maintaining the reliability and security of IoT-enabled service systems. Existing methods require training one specific model for each dataset, which exhibits limited generalization…

Machine Learning · Computer Science 2026-04-23 PengYu Chen , Shang Wan , Xiaohou Shi , Yuan Chang , Yan Sun , Sajal K. Das

Autonomous driving requires a comprehensive understanding of the surrounding environment for reliable trajectory planning. Previous works rely on dense rasterized scene representation (e.g., agent occupancy and semantic map) to perform…

Adapting Multimodal Large Language Models (MLLMs) for hour-long videos is bottlenecked by context limits. Dense visual streams saturate token budgets and exacerbate the lost-in-the-middle phenomenon. Existing heuristics, like sparse…

Intelligent Traffic Monitoring (ITMo) technologies hold the potential for improving road safety/security and for enabling smart city infrastructure. Understanding traffic situations requires a complex fusion of perceptual information with…

Computation and Language · Computer Science 2023-07-18 Jiarui Zhang , Filip Ilievski , Kaixin Ma , Aravinda Kollaa , Jonathan Francis , Alessandro Oltramari

The reliability of a machine vision system for autonomous driving depends heavily on its training data distribution. When a vehicle encounters significantly different conditions, such as atypical obstacles, its perceptual capabilities can…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Fabrizio Genilotti , Arianna Stropeni , Gionata Grotto , Francesco Borsatti , Manuel Barusco , Davide Dalle Pezze , Gian Antonio Susto

In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at both the 7th and 8th Dialog System Technology Challenges…

Computation and Language · Computer Science 2021-10-14 Ankit P. Shah , Shijie Geng , Peng Gao , Anoop Cherian , Takaaki Hori , Tim K. Marks , Jonathan Le Roux , Chiori Hori

Large Vision Language Models (LVLMs) have shown strong capabilities in understanding and analyzing visual scenes across various domains. However, in the context of autonomous driving, their limited comprehension of 3D environments restricts…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Jannik Lübberstedt , Esteban Rivera , Nico Uhlemann , Markus Lienkamp

The evolution of autonomous driving towards full automation demands robust interactive capabilities; however, the development of Vision-Language-Action (VLA) models is constrained by the sparsity of interactive scenarios and inadequate…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Haojie Feng , Peizhi Zhang , Mengjie Tian , Xinrui Zhang , Zhuoren Li , Junpeng Huang , Xiurong Wang , Junfan Zhu , Jianzhou Wang , Dongxiao Yin , Lu Xiong

Self-attention learns pairwise interactions to model long-range dependencies, yielding great improvements for video action recognition. In this paper, we seek a deeper understanding of self-attention for temporal modeling in videos. We…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Bo He , Xitong Yang , Zuxuan Wu , Hao Chen , Ser-Nam Lim , Abhinav Shrivastava

Large Vision-Language Models (LVLMs) have received widespread attention for advancing the interpretable self-driving. Existing evaluations of LVLMs primarily focus on multi-faceted capabilities in natural circumstances, lacking automated…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Kai Chen , Yanze Li , Wenhua Zhang , Yanxin Liu , Pengxiang Li , Ruiyuan Gao , Lanqing Hong , Meng Tian , Xinhai Zhao , Zhenguo Li , Dit-Yan Yeung , Huchuan Lu , Xu Jia

Video Anomaly Detection (VAD) aims to localize abnormal events on the timeline of long-range surveillance videos. Anomaly-scoring-based methods have been prevailing for years but suffer from the high complexity of thresholding and low…

Computer Vision and Pattern Recognition · Computer Science 2024-01-12 Hui Lv , Qianru Sun

While autonomous driving technologies continue to advance, current Advanced Driver Assistance Systems (ADAS) remain limited in their ability to interpret scene context or engage with drivers through natural language. These systems typically…

Robotics · Computer Science 2025-07-15 Kyungtae Han , Yitao Chen , Rohit Gupta , Onur Altintas

Personalization, while extensively studied in conventional autonomous driving pipelines, has been largely overlooked in the context of end-to-end autonomous driving (E2EAD), despite its critical role in fostering user trust, safety…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Ruiyang Hao , Bowen Jing , Haibao Yu , Zaiqing Nie

Vision Language Models (VLMs) demonstrate significant potential as embodied AI agents for various mobility applications. However, a standardized, closed-loop benchmark for evaluating their spatial reasoning and sequential decision-making…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Weizhen Wang , Chenda Duan , Zhenghao Peng , Yuxin Liu , Bolei Zhou

Reasoning about motion and space is a fundamental cognitive capability that is required by multiple real-world applications. While many studies highlight that large multimodal language models (MLMs) struggle to reason about space, they only…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Arijit Ray , Jiafei Duan , Ellis Brown , Reuben Tan , Dina Bashkirova , Rose Hendrix , Kiana Ehsani , Aniruddha Kembhavi , Bryan A. Plummer , Ranjay Krishna , Kuo-Hao Zeng , Kate Saenko
‹ Prev 1 4 5 6 7 8 10 Next ›