English
Related papers

Related papers: MiniGPT-4: Enhancing Vision-Language Understanding…

200 papers

Large Language Models (LLMs) have revolutionized the field of Natural Language Processing thanks to their ability to reuse knowledge acquired on massive text corpora on a wide variety of downstream tasks, with minimal (if any) tuning steps.…

Computation and Language · Computer Science 2024-07-12 Flavio Petruzzellis , Alberto Testolin , Alessandro Sperduti

Multi-modality promises to unlock further uses for large language models. Recently, the state-of-the-art language model GPT-4 was enhanced with vision capabilities. We carry out a prompting evaluation of GPT-4V and five other baselines on…

Computation and Language · Computer Science 2023-12-20 Mukul Singh , José Cambronero , Sumit Gulwani , Vu Le , Gust Verbruggen

Multimodal Large Language Models (MLLMs), such as GPT4o, have shown strong capabilities in visual reasoning and explanation generation. However, despite these strengths, they face significant challenges in the increasingly critical task of…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Fanrui Zhang , Jiawei Liu , Jiaying Zhu , Esther Sun , Dong Li , Qiang Zhang , Zheng-Jun Zha

We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs. While less capable than humans in many real-world scenarios, GPT-4 exhibits human-level performance on…

Computation and Language · Computer Science 2024-03-11 OpenAI , Josh Achiam , Steven Adler , Sandhini Agarwal , Lama Ahmad , Ilge Akkaya , Florencia Leoni Aleman , Diogo Almeida , Janko Altenschmidt , Sam Altman , Shyamal Anadkat , Red Avila , Igor Babuschkin , Suchir Balaji , Valerie Balcom , Paul Baltescu , Haiming Bao , Mohammad Bavarian , Jeff Belgum , Irwan Bello , Jake Berdine , Gabriel Bernadett-Shapiro , Christopher Berner , Lenny Bogdonoff , Oleg Boiko , Madelaine Boyd , Anna-Luisa Brakman , Greg Brockman , Tim Brooks , Miles Brundage , Kevin Button , Trevor Cai , Rosie Campbell , Andrew Cann , Brittany Carey , Chelsea Carlson , Rory Carmichael , Brooke Chan , Che Chang , Fotis Chantzis , Derek Chen , Sully Chen , Ruby Chen , Jason Chen , Mark Chen , Ben Chess , Chester Cho , Casey Chu , Hyung Won Chung , Dave Cummings , Jeremiah Currier , Yunxing Dai , Cory Decareaux , Thomas Degry , Noah Deutsch , Damien Deville , Arka Dhar , David Dohan , Steve Dowling , Sheila Dunning , Adrien Ecoffet , Atty Eleti , Tyna Eloundou , David Farhi , Liam Fedus , Niko Felix , Simón Posada Fishman , Juston Forte , Isabella Fulford , Leo Gao , Elie Georges , Christian Gibson , Vik Goel , Tarun Gogineni , Gabriel Goh , Rapha Gontijo-Lopes , Jonathan Gordon , Morgan Grafstein , Scott Gray , Ryan Greene , Joshua Gross , Shixiang Shane Gu , Yufei Guo , Chris Hallacy , Jesse Han , Jeff Harris , Yuchen He , Mike Heaton , Johannes Heidecke , Chris Hesse , Alan Hickey , Wade Hickey , Peter Hoeschele , Brandon Houghton , Kenny Hsu , Shengli Hu , Xin Hu , Joost Huizinga , Shantanu Jain , Shawn Jain , Joanne Jang , Angela Jiang , Roger Jiang , Haozhun Jin , Denny Jin , Shino Jomoto , Billie Jonn , Heewoo Jun , Tomer Kaftan , Łukasz Kaiser , Ali Kamali , Ingmar Kanitscheider , Nitish Shirish Keskar , Tabarak Khan , Logan Kilpatrick , Jong Wook Kim , Christina Kim , Yongjik Kim , Jan Hendrik Kirchner , Jamie Kiros , Matt Knight , Daniel Kokotajlo , Łukasz Kondraciuk , Andrew Kondrich , Aris Konstantinidis , Kyle Kosic , Gretchen Krueger , Vishal Kuo , Michael Lampe , Ikai Lan , Teddy Lee , Jan Leike , Jade Leung , Daniel Levy , Chak Ming Li , Rachel Lim , Molly Lin , Stephanie Lin , Mateusz Litwin , Theresa Lopez , Ryan Lowe , Patricia Lue , Anna Makanju , Kim Malfacini , Sam Manning , Todor Markov , Yaniv Markovski , Bianca Martin , Katie Mayer , Andrew Mayne , Bob McGrew , Scott Mayer McKinney , Christine McLeavey , Paul McMillan , Jake McNeil , David Medina , Aalok Mehta , Jacob Menick , Luke Metz , Andrey Mishchenko , Pamela Mishkin , Vinnie Monaco , Evan Morikawa , Daniel Mossing , Tong Mu , Mira Murati , Oleg Murk , David Mély , Ashvin Nair , Reiichiro Nakano , Rajeev Nayak , Arvind Neelakantan , Richard Ngo , Hyeonwoo Noh , Long Ouyang , Cullen O'Keefe , Jakub Pachocki , Alex Paino , Joe Palermo , Ashley Pantuliano , Giambattista Parascandolo , Joel Parish , Emy Parparita , Alex Passos , Mikhail Pavlov , Andrew Peng , Adam Perelman , Filipe de Avila Belbute Peres , Michael Petrov , Henrique Ponde de Oliveira Pinto , Michael , Pokorny , Michelle Pokrass , Vitchyr H. Pong , Tolly Powell , Alethea Power , Boris Power , Elizabeth Proehl , Raul Puri , Alec Radford , Jack Rae , Aditya Ramesh , Cameron Raymond , Francis Real , Kendra Rimbach , Carl Ross , Bob Rotsted , Henri Roussez , Nick Ryder , Mario Saltarelli , Ted Sanders , Shibani Santurkar , Girish Sastry , Heather Schmidt , David Schnurr , John Schulman , Daniel Selsam , Kyla Sheppard , Toki Sherbakov , Jessica Shieh , Sarah Shoker , Pranav Shyam , Szymon Sidor , Eric Sigler , Maddie Simens , Jordan Sitkin , Katarina Slama , Ian Sohl , Benjamin Sokolowsky , Yang Song , Natalie Staudacher , Felipe Petroski Such , Natalie Summers , Ilya Sutskever , Jie Tang , Nikolas Tezak , Madeleine B. Thompson , Phil Tillet , Amin Tootoonchian , Elizabeth Tseng , Preston Tuggle , Nick Turley , Jerry Tworek , Juan Felipe Cerón Uribe , Andrea Vallone , Arun Vijayvergiya , Chelsea Voss , Carroll Wainwright , Justin Jay Wang , Alvin Wang , Ben Wang , Jonathan Ward , Jason Wei , CJ Weinmann , Akila Welihinda , Peter Welinder , Jiayi Weng , Lilian Weng , Matt Wiethoff , Dave Willner , Clemens Winter , Samuel Wolrich , Hannah Wong , Lauren Workman , Sherwin Wu , Jeff Wu , Michael Wu , Kai Xiao , Tao Xu , Sarah Yoo , Kevin Yu , Qiming Yuan , Wojciech Zaremba , Rowan Zellers , Chong Zhang , Marvin Zhang , Shengjia Zhao , Tianhao Zheng , Juntang Zhuang , William Zhuk , Barret Zoph

Large Language Models (LLMs) have garnered considerable interest within both academic and industrial. Yet, the application of LLMs to graph data remains under-explored. In this study, we evaluate the capabilities of four LLMs in addressing…

Artificial Intelligence · Computer Science 2023-09-12 Chang Liu , Bo Wu

Recent developments in multimodal large language models (MLLMs) have spurred significant interest in their potential applications across various medical imaging domains. On the one hand, there is a temptation to use these generative models…

Image and Video Processing · Electrical Eng. & Systems 2024-06-05 Sulaiman Khan , Md. Rafiul Biswas , Alina Murad , Hazrat Ali , Zubair Shah

The advent of large language models (LLMs) has heightened interest in their potential for multimodal applications that integrate language and vision. This paper explores the capabilities of GPT-4V in the realms of geography, environmental…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Chenjiao Tan , Qian Cao , Yiwei Li , Jielu Zhang , Xiao Yang , Huaqin Zhao , Zihao Wu , Zhengliang Liu , Hao Yang , Nemin Wu , Tao Tang , Xinyue Ye , Lilong Chai , Ninghao Liu , Changying Li , Lan Mu , Tianming Liu , Gengchen Mai

Building on the advances of language models, Large Multimodal Models (LMMs) have contributed significant improvements in video understanding. While the current video LMMs utilize advanced Large Language Models (LLMs), they rely on either…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Muhammad Maaz , Hanoona Rasheed , Salman Khan , Fahad Khan

Large language models (LLMs) have undergone significant expansion and have been increasingly integrated across various domains. Notably, in the realm of robot task planning, LLMs harness their advanced reasoning and language comprehension…

In today's visually dominated social media landscape, predicting the perceived credibility of visual content and understanding what drives human judgment are crucial for countering misinformation. However, these tasks are challenging due to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Yilang Peng , Sijia Qian , Yingdan Lu , Cuihua Shen

This tutorial note summarizes the presentation on ``Large Multimodal Models: Towards Building and Surpassing Multimodal GPT-4'', a part of CVPR 2023 tutorial on ``Recent Advances in Vision Foundation Models''. The tutorial consists of three…

Computer Vision and Pattern Recognition · Computer Science 2023-06-27 Chunyuan Li

LLMs have demonstrated remarkable abilities at interacting with humans through language, especially with the usage of instruction-following data. Recent advancements in LLMs, such as MiniGPT-4, LLaVA, and X-LLM, further enlarge their…

Computer Vision and Pattern Recognition · Computer Science 2023-07-18 Yang Zhao , Zhijie Lin , Daquan Zhou , Zilong Huang , Jiashi Feng , Bingyi Kang

Visual storytelling is an emerging field that combines images and narratives to create engaging and contextually rich stories. Despite its potential, generating coherent and emotionally resonant visual stories remains challenging due to the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-04 Xiaochuan Lin , Xiangyong Chen

OpenAI's multimodal GPT-4o has demonstrated remarkable capabilities in image generation and editing, yet its ability to achieve world knowledge-informed semantic synthesis--seamlessly integrating domain knowledge, contextual reasoning, and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Ning Li , Jingran Zhang , Justin Cui

Recently, the visual generation ability by GPT-4o(mni) has been unlocked by OpenAI. It demonstrates a very remarkable generation capability with excellent multimodal condition understanding and varied task instructions. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Pu Cao , Feng Zhou , Junyi Ji , Qingye Kong , Zhixiang Lv , Mingjian Zhang , Xuekun Zhao , Siqi Wu , Yinghui Lin , Qing Song , Lu Yang

Large language models (LLMs) have increased interest in vision language models (VLMs), which process image-text pairs as input. Studies investigating the visual understanding ability of VLMs have been proposed, but such studies are still…

Computation and Language · Computer Science 2024-06-25 Jesse Atuhurra , Iqra Ali , Tatsuya Hiraoka , Hidetaka Kamigaito , Tomoya Iwakura , Taro Watanabe

Concepts, both abstract and concrete, elicit a distribution of association strengths across perceptual color space, which influence aspects of visual cognition ranging from object recognition to interpretation of information visualizations.…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Kushin Mukherjee , Timothy T. Rogers , Karen B. Schloss

GPT-4o, an all-encompassing model, represents a milestone in the development of large multi-modal language models. It can understand visual, auditory, and textual modalities, directly output audio, and support flexible duplex interaction.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-06 Zhifei Xie , Changqiao Wu

The pursuit of autonomous driving technology hinges on the sophisticated integration of perception, decision-making, and control systems. Traditional approaches, both data-driven and rule-based, have been hindered by their inability to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Licheng Wen , Xuemeng Yang , Daocheng Fu , Xiaofeng Wang , Pinlong Cai , Xin Li , Tao Ma , Yingxuan Li , Linran Xu , Dengke Shang , Zheng Zhu , Shaoyan Sun , Yeqi Bai , Xinyu Cai , Min Dou , Shuanglu Hu , Botian Shi , Yu Qiao

In this paper we address image classification tasks leveraging knowledge encoded in Large Multimodal Models (LMMs). More specifically, we use the MiniGPT-4 model to extract semantic descriptions for the images, in a multimodal prompting…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Maria Tzelepi , Vasileios Mezaris