中文
相关论文

相关论文: CCTVBench: Contrastive Consistency Traffic VideoQA…

200 篇论文

While multimodal large language models (MLLMs) exhibit strong performance on single-video tasks (e.g., video question answering), their capability for spatiotemporal pattern reasoning across multiple videos remains a critical gap in pattern…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Nannan Zhu , Yonghao Dong , Teng Wang , Xueqian Li , Shengjun Deng , Yijia Wang , Zheng Hong , Tiantian Geng , Guo Niu , Hanyan Huang , Xiongfei Yao , Shuaiwei Jiao

Multimodal large language models (MLLMs) achieve strong performance on single-view spatial reasoning tasks, yet it remains unclear whether they maintain stable spatial state representations under counterfactual viewpoint changes. We…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Shanmukha Vellamcheti , Uday Kiran Kothapalli , Disharee Bhowmick , Sathyanarayanan N. Aakur

We introduce ACCIDENT, a benchmark dataset for traffic accident detection in CCTV footage, designed to evaluate models in supervised (IID and OOD) and zero-shot settings, reflecting both data-rich and data-scarce scenarios. The benchmark…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Lukas Picek , Michal Čermák , Marek Hanzl , Vojtěch Čermák

Cooperative autonomous driving requires traffic scene understanding from both vehicle and infrastructure perspectives. While vision-language models (VLMs) show strong general reasoning capabilities, their performance in safety-critical…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Rui Gan , Junyi Ma , Pei Li , Xingyou Yang , Kai Chen , Sikai Chen , Bin Ran

Video language models (Video-LLMs) are prone to hallucinations, often generating plausible but ungrounded content when visual evidence is weak, ambiguous, or biased. Existing decoding methods, such as contrastive decoding (CD), rely on…

人工智能 · 计算机科学 2026-02-10 Qixin Xiao

Traffic accidents result in millions of injuries and fatalities globally, with a significant number occurring at intersections each year. Traffic Signal Control (TSC) is an effective strategy for enhancing safety at these urban junctures.…

机器学习 · 计算机科学 2025-12-17 Mingyuan Li , Chunyu Liu , Zhuojun Li , Xiao Liu , Guangsheng Yu , Bo Du , Jun Shen , Qiang Wu

Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largely unexplored. In open-domain applications visual and textual…

Safety-critical planning in complex environments, particularly at urban intersections, remains a fundamental challenge for autonomous driving. Existing methods, whether rule-based or data-driven, frequently struggle to capture complex scene…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Kefei Tian , Yuansheng Lian , Kai Yang , Xiangdong Chen , Shen Li

Vision Language Models (VLMs) have recently shown significant advancements in video understanding, especially in feature alignment, event reasoning, and instruction-following tasks. However, their capability for counterfactual reasoning,…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Yuefei Chen , Jiang Liu , Xiaodong Lin , Ruixiang Tang

Vision-language models (VLMs) achieve strong performance on many benchmarks, yet a basic reliability question remains underexplored: when visual evidence conflicts with commonsense, do models follow what is shown or what commonsense…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Kesheng Chen , Yamin Hu , Qi Zhou , Zhenqian Zhu , Wenjian Luo

Video large language models (Video-LLMs) have made strong progress in general video understanding, but their ability to maintain temporal object consistency remains underexplored. Existing benchmarks often emphasize event recognition,…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Junzhe Chen , Siyuan Meng , Yuxi Chen , Man Zhao , Wenyao Gui , Xiaojie Guo

Ensuring validation for highly automated driving poses significant obstacles to the widespread adoption of highly automated vehicles. Scenario-based testing offers a potential solution by reducing the homologation effort required for these…

机器学习 · 计算机科学 2023-09-19 Maximilian Zipfl , Moritz Jarosch , J. Marius Zöllner

Traffic cameras are essential in urban areas, playing a crucial role in intelligent transportation systems. Multiple cameras at intersections enhance law enforcement capabilities, traffic management, and pedestrian safety. However,…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Md Adnan Arefeen , Biplob Debnath , Srimat Chakradhar

Accident detection using Closed Circuit Television (CCTV) footage is one of the most imperative features for enhancing transport safety and efficient traffic control. To this end, this research addresses the issues of supervised monitoring…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Zhenghao Xi , Xiang Liu , Yaqi Liu , Yitong Cai , Yangyu Zheng

Computational Colour Constancy (CCC) consists of estimating the colour of one or more illuminants in a scene and using them to remove unwanted chromatic distortions. Much research has focused on illuminant estimation for CCC on single…

计算机视觉与模式识别 · 计算机科学 2023-03-03 Matteo Rizzo , Cristina Conati , Daesik Jang , Hui Hu

Rapid advances in multimodal models demand benchmarks that rigorously evaluate understanding and reasoning in safety-critical, dynamic real-world settings. We present AccidentBench, a large-scale benchmark that combines vehicle accident…

Despite their impressive generative capabilities, LLMs are hindered by fact-conflicting hallucinations in real-world applications. The accurate identification of hallucinations in texts generated by LLMs, especially in complex inferential…

计算与语言 · 计算机科学 2024-05-28 Xiang Chen , Duanzheng Song , Honghao Gui , Chenxi Wang , Ningyu Zhang , Yong Jiang , Fei Huang , Chengfei Lv , Dan Zhang , Huajun Chen

While sequential reasoning enhances the capability of Vision-Language Models (VLMs) to execute complex multimodal tasks, their reliability in grounding these reasoning chains within actual visual evidence remains insufficiently explored. We…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Rory Driscoll , Alexandros Christoforos , Chadbourne Davis

Automating crash video analysis is essential to leverage the growing availability of driving video data for traffic safety research and accountability attribution in autonomous driving. Crash video analysis is a challenging multitask…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Kaidi Liang , Ke Li , Xianbiao Hu , Ruwen Qin

Large Vision-Language Model (LVLM) systems have demonstrated impressive vision-language reasoning capabilities but suffer from pervasive and severe hallucination issues, posing significant risks in critical domains such as healthcare and…

计算机视觉与模式识别 · 计算机科学 2024-11-20 Zhehan Kan , Ce Zhang , Zihan Liao , Yapeng Tian , Wenming Yang , Junyuan Xiao , Xu Li , Dongmei Jiang , Yaowei Wang , Qingmin Liao
‹ 上一页 1 2 3 10 下一页 ›