English
Related papers

Related papers: MITS: A Large-Scale Multimodal Benchmark Dataset f…

200 papers

We introduce the first very large detection dataset for event cameras. The dataset is composed of more than 39 hours of automotive recordings acquired with a 304x240 ATIS sensor. It contains open roads and very diverse driving scenarios,…

Computer Vision and Pattern Recognition · Computer Science 2020-02-03 Pierre de Tournemire , Davide Nitti , Etienne Perot , Davide Migliore , Amos Sironi

Recently, Vision-Language Models (VLMs) have achieved remarkable progress in multimodal tasks, and multimodal instruction data serves as the foundation for enhancing VLM capabilities. Despite the availability of several open-source…

Urban traffic optimization using traffic cameras as sensors is driving the need to advance state-of-the-art multi-target multi-camera (MTMC) tracking. This work introduces CityFlow, a city-scale traffic camera dataset consisting of more…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Zheng Tang , Milind Naphade , Ming-Yu Liu , Xiaodong Yang , Stan Birchfield , Shuo Wang , Ratnesh Kumar , David Anastasiu , Jenq-Neng Hwang

Intelligent transportation systems (ITSs) and other smart-city technologies are increasingly advancing in capability and complexity. While simulation environments continue to improve, their fidelity and ease of use can quickly degrade as…

Physics and Society · Physics 2019-07-31 Adam Morrissett , Roja Eini , Mostafa Zaman , Nasibeh Zohrabi , Sherif Abdelwahed

Evaluation is important for multimodal generation tasks, while traditional multimodal evaluation metrics suffer from several limitations. With the rapid progress of MLLMs, there is growing interest in applying MLLMs to build general…

Computation and Language · Computer Science 2026-04-30 Junzhe Zhang , Huixuan Zhang , Xinyu Hu , Li Lin , Mingqi Gao , Shi Qiu , Xiaojun Wan

With the rapid proliferation of autonomous driving, there has been a heightened focus on the research of lidar-based 3D semantic segmentation and object detection methodologies, aiming to ensure the safety of traffic participants. In recent…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Jiahua Xu , Si Zuo , Chenfeng Wei , Wei Zhou

We propose L2T, an advancement of visual instruction tuning (VIT). While VIT equips Multimodal LLMs (MLLMs) with promising multimodal capabilities, the current design choices for VIT often result in overfitting and shortcut learning,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Zhihan Zhou , Feng Hong , Jiaan Luo , Jiangchao Yao , Dongsheng Li , Bo Han , Ya Zhang , Yanfeng Wang

In autonomous driving community, numerous benchmarks have been established to assist the tasks of 3D/2D object detection, stereo vision, semantic/instance segmentation. However, the more meaningful dynamic evolution of the surrounding…

Computer Vision and Pattern Recognition · Computer Science 2019-03-18 Jianru Xue , Jianwu Fang , Tao Li , Bohua Zhang , Pu Zhang , Zhen Ye , Jian Dou

In intelligent transportation systems (ITSs), incorporating pedestrians and vehicles in-the-loop is crucial for developing realistic and safe traffic management solutions. However, there is falls short of simulating complex real-world ITS…

Emerging Technologies · Computer Science 2025-03-07 Xiaolong Li , Jianhao Wei , Haidong Wang , Li Dong , Ruoyang Chen , Changyan Yi , Jun Cai , Dusit Niyato , Xuemin , Shen

During the process of driving, humans usually rely on multiple senses to gather information and make decisions. Analogously, in order to achieve embodied intelligence in autonomous driving, it is essential to integrate multidimensional…

The perception system in personalized mobile agents requires developing indoor scene understanding models, which can understand 3D geometries, capture objectiveness, analyze human behaviors, etc. Nonetheless, this direction has not been…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Yao-Hung Hubert Tsai , Hanlin Goh , Ali Farhadi , Jian Zhang

Recent advances in clinical AI have enabled remarkable progress across many clinical domains. However, existing benchmarks and models are primarily limited to a small set of modalities and tasks, which hinders the development of large-scale…

Machine Learning · Computer Science 2025-03-21 Wei Dai , Peilin Chen , Malinda Lu , Daniel Li , Haowen Wei , Hejie Cui , Paul Pu Liang

Recent advancements in Large Multimodal Models (LMMs) have significantly improved multimodal understanding and generation. However, these models still struggle to generate tightly interleaved image-text outputs, primarily due to the limited…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yukang Feng , Jianwen Sun , Chuanhao Li , Zizhen Li , Jiaxin Ai , Fanrui Zhang , Yifan Chang , Sizhuo Zhou , Shenglin Zhang , Yu Dai , Kaipeng Zhang

The remarkable potential of multi-modal large language models (MLLMs) in comprehending both vision and language information has been widely acknowledged. However, the scarcity of 3D scenes-language pairs in comparison to their 2D…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Zeju Li , Chao Zhang , Xiaoyan Wang , Ruilong Ren , Yifan Xu , Ruifei Ma , Xiangde Liu

Traffic signboards are vital for road safety and intelligent transportation systems, enabling navigation and autonomous driving. Yet, recognizing traffic signs at night remains underexplored due to the scarcity of realistic public datasets…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Aditya Mishra , Akshay Agarwal , Haroon Lone

We introduce Long-VITA, a simple yet effective large multi-modal model for long-context visual-language understanding tasks. It is adept at concurrently processing and analyzing modalities of image, video, and text over 4K frames or 1M…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Yunhang Shen , Chaoyou Fu , Shaoqi Dong , Xiong Wang , Yi-Fan Zhang , Peixian Chen , Mengdan Zhang , Haoyu Cao , Ke Li , Shaohui Lin , Xiawu Zheng , Yan Zhang , Yiyi Zhou , Ran He , Caifeng Shan , Rongrong Ji , Xing Sun

Large language models (LLMs) show strong potential for Intelligent Transportation Systems (ITS), particularly in tasks requiring situational reasoning and multi-agent coordination. These capabilities make them well suited for cooperative…

Artificial Intelligence · Computer Science 2026-05-21 Abderrahmane Lakas , Mohamed Amine Ferrag , Merouane Debbah

With the increasing integration of visual and textual content in Social Networking Services (SNS), evaluating the multimodal capabilities of Large Language Models (LLMs) is crucial for enhancing user experience, content understanding, and…

Computation and Language · Computer Science 2025-12-16 Hongcheng Guo , Zheyong Xie , Shaosheng Cao , Boyang Wang , Weiting Liu , Anjie Le , Lei Li , Zhoujun Li

We present a large-scale, longitudinal visual dataset of urban streetlights captured by 22 fixed-angle cameras deployed across Bristol, U.K., from 2021 to 2025. The dataset contains over 526,000 images, collected hourly under diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Peizheng Li , Ioannis Mavromatis , Ajith Sahadevan , Tim Farnham , Adnan Aijaz , Aftab Khan

Semantic scene understanding is crucial for robust and safe autonomous navigation, particularly so in off-road environments. Recent deep learning advances for 3D semantic segmentation rely heavily on large sets of training data, however…

Computer Vision and Pattern Recognition · Computer Science 2022-05-26 Peng Jiang , Philip Osteen , Maggie Wigness , Srikanth Saripalli
‹ Prev 1 3 4 5 6 7 10 Next ›