English
Related papers

Related papers: Examining the Commitments and Difficulties Inheren…

200 papers

The advent of large language models (LLMs) has heightened interest in their potential for multimodal applications that integrate language and vision. This paper explores the capabilities of GPT-4V in the realms of geography, environmental…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Chenjiao Tan , Qian Cao , Yiwei Li , Jielu Zhang , Xiao Yang , Huaqin Zhao , Zihao Wu , Zhengliang Liu , Hao Yang , Nemin Wu , Tao Tang , Xinyue Ye , Lilong Chai , Ninghao Liu , Changying Li , Lan Mu , Tianming Liu , Gengchen Mai

Multimodal foundation models (MFMs), such as GPT-4o, have recently made remarkable progress. However, their detailed visual understanding beyond question answering remains unclear. In this paper, we benchmark popular MFMs (GPT-4o, o4-mini,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Rahul Ramachandran , Ali Garjani , Roman Bachmann , Andrei Atanov , Oğuzhan Fatih Kar , Amir Zamir

Recent developments in multimodal large language models (MLLMs) have spurred significant interest in their potential applications across various medical imaging domains. On the one hand, there is a temptation to use these generative models…

Image and Video Processing · Electrical Eng. & Systems 2024-06-05 Sulaiman Khan , Md. Rafiul Biswas , Alina Murad , Hazrat Ali , Zubair Shah

While there is much excitement about the potential of large multimodal models (LMM), a comprehensive evaluation is critical to establish their true capabilities and limitations. In support of this aim, we evaluate two state-of-the-art LMMs,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-15 Mengchen Liu , Chongyan Chen , Danna Gurari

The task of image captioning demands an algorithm to generate natural language descriptions of visual inputs. Recent advancements have seen a convergence between image captioning research and the development of Large Language Models (LLMs)…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Davide Bucciarelli , Nicholas Moratelli , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara

The surge of interest towards Multi-modal Large Language Models (MLLMs), e.g., GPT-4V(ision) from OpenAI, has marked a significant trend in both academia and industry. They endow Large Language Models (LLMs) with powerful capabilities in…

Neighborhood environments include physical and environmental conditions such as housing quality, roads, and sidewalks, which significantly influence human health and well-being. Traditional methods for assessing these environments,…

Artificial Intelligence · Computer Science 2025-05-14 Andrew Cart , Shaohu Zhang , Melanie Escue , Xugui Zhou , Haitao Zhao , Prashanth BusiReddyGari , Beiyu Lin , Shuang Li

Recent advancements in multimodal techniques open exciting possibilities for models excelling in diverse tasks involving text, audio, and image processing. Models like GPT-4V, blending computer vision and language modeling, excel in complex…

Computation and Language · Computer Science 2023-10-20 Xiang Zhang , Senyu Li , Zijun Wu , Ning Shi

Large language models are effective at few-shot in-context learning (ICL). Recent advancements in multimodal foundation models have enabled unprecedentedly long context windows, presenting an opportunity to explore their capability to…

Machine Learning · Computer Science 2024-10-08 Yixing Jiang , Jeremy Irvin , Ji Hun Wang , Muhammad Ahmed Chaudhry , Jonathan H. Chen , Andrew Y. Ng

Recent advances in multimodal large language models enable new possibilities for image-based decision support. However, their reliability and operational trade-offs in neuroimaging remain insufficiently understood. We present a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Katarina Trojachanec Dineva , Stefan Andonov , Ilinka Ivanoska , Ivan Kitanovski , Sasho Gramatikov , Tamara Kostova , Monika Simjanoska Misheva , Kostadin Mishev

Recent advancements in generative AI systems have raised concerns about academic integrity among educators. Beyond excelling at solving programming problems and text-based multiple-choice questions, recent research has also found that large…

Artificial Intelligence · Computer Science 2024-12-17 Sebastian Gutierrez , Irene Hou , Jihye Lee , Kenneth Angelikas , Owen Man , Sophia Mettille , James Prather , Paul Denny , Stephen MacNeil

Vision language models (VLMs) have recently emerged and gained the spotlight for their ability to comprehend the dual modality of image and textual data. VLMs such as LLaVA, ChatGPT-4, and Gemini have recently shown impressive performance…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Prateek Verma , Minh-Hao Van , Xintao Wu

The success of Large Language Models (LLMs) has led to a parallel rise in the development of Large Multimodal Models (LMMs), which have begun to transform a variety of applications. These sophisticated multimodal models are designed to…

Artificial Intelligence · Computer Science 2025-05-20 Fouad Trad , Ali Chehab

The rapidly evolving sector of Multi-modal Large Language Models (MLLMs) is at the forefront of integrating linguistic and visual processing in artificial intelligence. This paper presents an in-depth comparative study of two pioneering…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Zhangyang Qi , Ye Fang , Mengchen Zhang , Zeyi Sun , Tong Wu , Ziwei Liu , Dahua Lin , Jiaqi Wang , Hengshuang Zhao

Multimodal large language models (MLLMs) have shown remarkable capabilities across a broad range of tasks but their knowledge and abilities in the geographic and geospatial domains are yet to be explored, despite potential wide-ranging…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Jonathan Roberts , Timo Lüddecke , Rehan Sheikh , Kai Han , Samuel Albanie

Traffic safety remains a critical global concern, with timely and accurate accident detection essential for hazard reduction and rapid emergency response. Infrastructure-based vision sensors offer scalable and efficient solutions for…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Ilhan Skender , Kailin Tong , Selim Solmaz , Daniel Watzenig

We present a novel framework for automatically evaluating building conditions nationwide in the United States by leveraging large language models (LLMs) and Google Street View (GSV) imagery. By fine-tuning Gemma 3 27B on a modest…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Siyuan Yao , Siavash Ghorbany , Kuangshi Ai , Arnav Cherukuthota , Meghan Forstchen , Alexis Korotasz , Matthew Sisk , Ming Hu , Chaoli Wang

The emergence of multimodal large models (MLMs) has significantly advanced the field of visual understanding, offering remarkable capabilities in the realm of visual question answering (VQA). Yet, the true challenge lies in the domain of…

Computation and Language · Computer Science 2024-08-27 Yunxin Li , Longyue Wang , Baotian Hu , Xinyu Chen , Wanqi Zhong , Chenyang Lyu , Wei Wang , Min Zhang

Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior quality and capabilities across diverse tasks. However, their…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Xuelu Feng , Yunsheng Li , Dongdong Chen , Mei Gao , Mengchen Liu , Junsong Yuan , Chunming Qiao

This study investigates the potential of a multimodal large language model (LLM), specifically ChatGPT-4o, to perform human-like interpretations of traffic scenes using static dashcam images. Herein, we focus on three judgment tasks…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Yuki Yoshihara , Linjing Jiang , Nihan Karatas , Hitoshi Kanamori , Asuka Harada , Takahiro Tanaka
‹ Prev 1 2 3 10 Next ›