English
Related papers

Related papers: GPT-4V Explorations: Mining Autonomous Driving

200 papers

The rapidly evolving sector of Multi-modal Large Language Models (MLLMs) is at the forefront of integrating linguistic and visual processing in artificial intelligence. This paper presents an in-depth comparative study of two pioneering…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Zhangyang Qi , Ye Fang , Mengchen Zhang , Zeyi Sun , Tong Wu , Ziwei Liu , Dahua Lin , Jiaqi Wang , Hengshuang Zhao

We propose a novel and challenging benchmark, AutoEval-Video, to comprehensively evaluate large vision-language models in open-ended video question answering. The comprehensiveness of AutoEval-Video is demonstrated in two aspects: 1)…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Xiuyuan Chen , Yuan Lin , Yuchen Zhang , Weiran Huang

The emergence of Large Multimodal Models (LMMs) marks a significant milestone in the development of artificial intelligence. Insurance, as a vast and complex discipline, involves a wide variety of data forms in its operational processes,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Chenwei Lin , Hanjia Lyu , Jiebo Luo , Xian Xu

The remote sensing image intelligence understanding model is undergoing a new profound paradigm shift which has been promoted by multi-modal large language model (MLLM), i.e. from the paradigm learning a domain model (LaDM) shifts to…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Linrui Xu , Ling Zhao , Wang Guo , Qiujun Li , Kewang Long , Kaiqi Zou , Yuhan Wang , Haifeng Li

Large language models (LLMs) have demonstrated a powerful ability to answer various queries as a general-purpose assistant. The continuous multi-modal large language models (MLLM) empower LLMs with the ability to perceive visual signals.…

Computation and Language · Computer Science 2024-01-05 Ziqiang Zheng , Yiwei Chen , Jipeng Zhang , Tuan-Anh Vu , Huimin Zeng , Yue Him Wong Tim , Sai-Kit Yeung

Large language models have seen widespread adoption in math problem-solving. However, in geometry problems that usually require visual aids for better understanding, even the most advanced multi-modal models currently still face challenges…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Shihao Cai , Keqin Bao , Hangyu Guo , Jizhi Zhang , Jun Song , Bo Zheng

Safely navigating street intersections is a complex challenge for blind and low-vision individuals, as it requires a nuanced understanding of the surrounding context - a task heavily reliant on visual cues. Traditional methods for assisting…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Hochul Hwang , Sunjae Kwon , Yekyung Kim , Donghyun Kim

Autonomous driving is a challenging task that requires perceiving and understanding the surrounding environment for safe trajectory planning. While existing vision-based end-to-end models have achieved promising results, these methods are…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Tengpeng Li , Hanli Wang , Xianfei Li , Wenlong Liao , Tao He , Pai Peng

OpenAI's latest large vision-language model (LVLM), GPT-4V(ision), has piqued considerable interest for its potential in medical applications. Despite its promise, recent studies and internal reviews highlight its underperformance in…

Computation and Language · Computer Science 2023-12-13 Pengcheng Chen , Ziyan Huang , Zhongying Deng , Tianbin Li , Yanzhou Su , Haoyu Wang , Jin Ye , Yu Qiao , Junjun He

The integration of artificial intelligence into scientific research has reached a new pinnacle with GPT-4V, a large language model featuring enhanced vision capabilities, accessible through ChatGPT or an API. This study demonstrates the…

Artificial Intelligence · Computer Science 2025-08-19 Zhiling Zheng , Zhiguo He , Omar Khattab , Nakul Rampal , Matei A. Zaharia , Christian Borgs , Jennifer T. Chayes , Omar M. Yaghi

Traffic control in unsignalized urban intersections presents significant challenges due to the complexity, frequent conflicts, and blind spots. This study explores the capability of leveraging Multimodal Large Language Models (MLLMs), such…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Sari Masri , Huthaifa I. Ashqar , Mohammed Elhenawy

We present MM-VID, an integrated system that harnesses the capabilities of GPT-4V, combined with specialized tools in vision, audio, and speech, to facilitate advanced video understanding. MM-VID is designed to address the challenges posed…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Kevin Lin , Faisal Ahmed , Linjie Li , Chung-Ching Lin , Ehsan Azarnasab , Zhengyuan Yang , Jianfeng Wang , Lin Liang , Zicheng Liu , Yumao Lu , Ce Liu , Lijuan Wang

Large Multimodal Model (LMM) GPT-4V(ision) endows GPT-4 with visual grounding capabilities, making it possible to handle certain tasks through the Visual Question Answering (VQA) paradigm. This paper explores the potential of VQA-oriented…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Jiangning Zhang , Haoyang He , Xuhai Chen , Zhucun Xue , Yabiao Wang , Chengjie Wang , Lei Xie , Yong Liu

Recent studies indicate that Generative Pre-trained Transformer 4 with Vision (GPT-4V) outperforms human physicians in medical challenge tasks. However, these evaluations primarily focused on the accuracy of multi-choice questions alone.…

Large language models (LLMs) have undergone significant expansion and have been increasingly integrated across various domains. Notably, in the realm of robot task planning, LLMs harness their advanced reasoning and language comprehension…

Traditional autonomous driving systems often struggle with reasoning in complex, unexpected scenarios due to limited comprehension of spatial relationships. In response, this study introduces a Large Language Model (LLM)-based Autonomous…

Computer Vision and Pattern Recognition · Computer Science 2025-02-12 Namhee Kim , Woojin Park

Recent developments in multimodal methodologies have marked the beginning of an exciting era for models adept at processing diverse data types, encompassing text, audio, and visual content. Models like GPT-4V, which merge computer vision…

Computation and Language · Computer Science 2024-11-15 Xiang Zhang , Senyu Li , Ning Shi , Bradley Hauer , Zijun Wu , Grzegorz Kondrak , Muhammad Abdul-Mageed , Laks V. S. Lakshmanan

Educational scholars have analyzed various image data acquired from teaching and learning situations, such as photos that shows classroom dynamics, students' drawings with regard to the learning content, textbook illustrations, etc.…

Physics Education · Physics 2024-05-14 Gyeong-Geon Lee , Xiaoming Zhai

Recently, it has been recognized that large language models demonstrate high performance on various intellectual tasks. However, few studies have investigated alignment with humans in behaviors that involve sensibility, such as aesthetic…

Artificial Intelligence · Computer Science 2024-03-07 Yoshia Abe , Tatsuya Daikoku , Yasuo Kuniyoshi

Deep learning models for autonomous driving, encompassing perception, planning, and control, depend on vast datasets to achieve their high performance. However, their generalization often suffers due to domain-specific data distributions,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Esteban Rivera , Jannik Lübberstedt , Nico Uhlemann , Markus Lienkamp