中文
相关论文

相关论文: Examining the Commitments and Difficulties Inheren…

200 篇论文

The advent of large language models (LLMs) has heightened interest in their potential for multimodal applications that integrate language and vision. This paper explores the capabilities of GPT-4V in the realms of geography, environmental…

Multimodal foundation models (MFMs), such as GPT-4o, have recently made remarkable progress. However, their detailed visual understanding beyond question answering remains unclear. In this paper, we benchmark popular MFMs (GPT-4o, o4-mini,…

计算机视觉与模式识别 · 计算机科学 2026-05-04 Rahul Ramachandran , Ali Garjani , Roman Bachmann , Andrei Atanov , Oğuzhan Fatih Kar , Amir Zamir

Recent developments in multimodal large language models (MLLMs) have spurred significant interest in their potential applications across various medical imaging domains. On the one hand, there is a temptation to use these generative models…

图像与视频处理 · 电气工程与系统科学 2024-06-05 Sulaiman Khan , Md. Rafiul Biswas , Alina Murad , Hazrat Ali , Zubair Shah

While there is much excitement about the potential of large multimodal models (LMM), a comprehensive evaluation is critical to establish their true capabilities and limitations. In support of this aim, we evaluate two state-of-the-art LMMs,…

计算机视觉与模式识别 · 计算机科学 2024-02-15 Mengchen Liu , Chongyan Chen , Danna Gurari

The task of image captioning demands an algorithm to generate natural language descriptions of visual inputs. Recent advancements have seen a convergence between image captioning research and the development of Large Language Models (LLMs)…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Davide Bucciarelli , Nicholas Moratelli , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara

The surge of interest towards Multi-modal Large Language Models (MLLMs), e.g., GPT-4V(ision) from OpenAI, has marked a significant trend in both academia and industry. They endow Large Language Models (LLMs) with powerful capabilities in…

Neighborhood environments include physical and environmental conditions such as housing quality, roads, and sidewalks, which significantly influence human health and well-being. Traditional methods for assessing these environments,…

人工智能 · 计算机科学 2025-05-14 Andrew Cart , Shaohu Zhang , Melanie Escue , Xugui Zhou , Haitao Zhao , Prashanth BusiReddyGari , Beiyu Lin , Shuang Li

Recent advancements in multimodal techniques open exciting possibilities for models excelling in diverse tasks involving text, audio, and image processing. Models like GPT-4V, blending computer vision and language modeling, excel in complex…

计算与语言 · 计算机科学 2023-10-20 Xiang Zhang , Senyu Li , Zijun Wu , Ning Shi

Large language models are effective at few-shot in-context learning (ICL). Recent advancements in multimodal foundation models have enabled unprecedentedly long context windows, presenting an opportunity to explore their capability to…

机器学习 · 计算机科学 2024-10-08 Yixing Jiang , Jeremy Irvin , Ji Hun Wang , Muhammad Ahmed Chaudhry , Jonathan H. Chen , Andrew Y. Ng

Recent advances in multimodal large language models enable new possibilities for image-based decision support. However, their reliability and operational trade-offs in neuroimaging remain insufficiently understood. We present a…

Recent advancements in generative AI systems have raised concerns about academic integrity among educators. Beyond excelling at solving programming problems and text-based multiple-choice questions, recent research has also found that large…

Vision language models (VLMs) have recently emerged and gained the spotlight for their ability to comprehend the dual modality of image and textual data. VLMs such as LLaVA, ChatGPT-4, and Gemini have recently shown impressive performance…

计算机视觉与模式识别 · 计算机科学 2024-05-03 Prateek Verma , Minh-Hao Van , Xintao Wu

The success of Large Language Models (LLMs) has led to a parallel rise in the development of Large Multimodal Models (LMMs), which have begun to transform a variety of applications. These sophisticated multimodal models are designed to…

人工智能 · 计算机科学 2025-05-20 Fouad Trad , Ali Chehab

The rapidly evolving sector of Multi-modal Large Language Models (MLLMs) is at the forefront of integrating linguistic and visual processing in artificial intelligence. This paper presents an in-depth comparative study of two pioneering…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Zhangyang Qi , Ye Fang , Mengchen Zhang , Zeyi Sun , Tong Wu , Ziwei Liu , Dahua Lin , Jiaqi Wang , Hengshuang Zhao

Multimodal large language models (MLLMs) have shown remarkable capabilities across a broad range of tasks but their knowledge and abilities in the geographic and geospatial domains are yet to be explored, despite potential wide-ranging…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Jonathan Roberts , Timo Lüddecke , Rehan Sheikh , Kai Han , Samuel Albanie

Traffic safety remains a critical global concern, with timely and accurate accident detection essential for hazard reduction and rapid emergency response. Infrastructure-based vision sensors offer scalable and efficient solutions for…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Ilhan Skender , Kailin Tong , Selim Solmaz , Daniel Watzenig

We present a novel framework for automatically evaluating building conditions nationwide in the United States by leveraging large language models (LLMs) and Google Street View (GSV) imagery. By fine-tuning Gemma 3 27B on a modest…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Siyuan Yao , Siavash Ghorbany , Kuangshi Ai , Arnav Cherukuthota , Meghan Forstchen , Alexis Korotasz , Matthew Sisk , Ming Hu , Chaoli Wang

The emergence of multimodal large models (MLMs) has significantly advanced the field of visual understanding, offering remarkable capabilities in the realm of visual question answering (VQA). Yet, the true challenge lies in the domain of…

计算与语言 · 计算机科学 2024-08-27 Yunxin Li , Longyue Wang , Baotian Hu , Xinyu Chen , Wanqi Zhong , Chenyang Lyu , Wei Wang , Min Zhang

Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior quality and capabilities across diverse tasks. However, their…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Xuelu Feng , Yunsheng Li , Dongdong Chen , Mei Gao , Mengchen Liu , Junsong Yuan , Chunming Qiao

This study investigates the potential of a multimodal large language model (LLM), specifically ChatGPT-4o, to perform human-like interpretations of traffic scenes using static dashcam images. Herein, we focus on three judgment tasks…

计算机视觉与模式识别 · 计算机科学 2025-07-14 Yuki Yoshihara , Linjing Jiang , Nihan Karatas , Hitoshi Kanamori , Asuka Harada , Takahiro Tanaka
‹ 上一页 1 2 3 10 下一页 ›