English
Related papers

Related papers: Preliminary Explorations with GPT-4o(mni) Native I…

200 papers

While the capabilities of generative models heavily improved in different domains (images, text, graphs, molecules, etc.), their evaluation metrics largely remain based on simplified quantities or manual inspection with limited…

Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their potential in supporting computer-use planning tasks, which are closely related to our…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Junxian Li , Kai Liu , Leyang Chen , Weida Wang , Zhixin Wang , Jiaqi Xu , Fan Li , Renjing Pei , Linghe Kong , Yulun Zhang

Recent advances in learning multi-modal representation have witnessed the success in biomedical domains. While established techniques enable handling multi-modal information, the challenges are posed when extended to various clinical…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Chenxin Li , Xinyu Liu , Cheng Wang , Yifan Liu , Weihao Yu , Jing Shao , Yixuan Yuan

Semantic and Cross-language code clone generation may be useful for code reuse, code comprehension, refactoring and benchmarking. OpenAI's GPT model has potential in such clone generation as GPT is used for text generation. When developers…

Software Engineering · Computer Science 2023-09-13 Palash R. Roy , Ajmain I. Alam , Farouq Al-omari , Banani Roy , Chanchal K. Roy , Kevin A. Schneider

Vision-Language Models (VLMs), such as GPT-4V and Llama 3.2 vision, have garnered significant research attention for their ability to leverage Large Language Models (LLMs) in multimodal tasks. However, their potential is constrained by…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Mukund Agarwalla , Himanshu Kumar , Raj Dandekar , Rajat Dandekar , Sreedath Panat

Large Language Models have shown prominent capabilities in generating functional code from natural language descriptions. However, a standardized way to evaluate these capabilities in an objective and unbiased manner is still to be found.…

Software Engineering · Computer Science 2024-10-23 Álvaro Barbero Jiménez

Many real-world tasks require an agent to reason jointly over text and visual objects, (e.g., navigating in public spaces), which we refer to as context-sensitive text-rich visual reasoning. Specifically, these tasks require an…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Rohan Wadhawan , Hritik Bansal , Kai-Wei Chang , Nanyun Peng

Objective: Radiology residents require timely, personalized feedback to develop accurate image analysis and reporting skills. Increasing clinical workload often limits attendings' ability to provide guidance. This study evaluates a…

Generating novel and useful concepts is essential during the early design stage to explore a large variety of design opportunities, which usually requires advanced design thinking ability and a wide range of knowledge from designers.…

Computation and Language · Computer Science 2022-11-08 Qihao Zhu , Jianxi Luo

Many animal species can approximately judge the number of objects in a visual scene at a single glance, and humans can further determine the exact cardinality of a set by deploying systematic counting procedures. In contrast, it has been…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Alberto Testolin , Kuinan Hou , Marco Zorzi

Recent years witness the tremendous success of generative adversarial networks (GANs) in synthesizing photo-realistic images. GAN generator learns to compose realistic images and reproduce the real data distribution. Through that, a…

Computer Vision and Pattern Recognition · Computer Science 2023-01-16 Yinghao Xu , Yujun Shen , Jiapeng Zhu , Ceyuan Yang , Bolei Zhou

This paper presents a comprehensive evaluation of the Optical Character Recognition (OCR) capabilities of the recently released GPT-4V(ision), a Large Multimodal Model (LMM). We assess the model's performance across a range of OCR tasks,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Yongxin Shi , Dezhi Peng , Wenhui Liao , Zening Lin , Xinhong Chen , Chongyu Liu , Yuyi Zhang , Lianwen Jin

As information exists in various modalities in real world, effective interaction and fusion among multimodal information plays a key role for the creation and perception of multimodal data in computer vision and deep learning research. With…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Fangneng Zhan , Yingchen Yu , Rongliang Wu , Jiahui Zhang , Shijian Lu , Lingjie Liu , Adam Kortylewski , Christian Theobalt , Eric Xing

Benchmarks for large multimodal language models (MLMs) now serve to simultaneously assess the general capabilities of models instead of evaluating for a specific capability. As a result, when a developer wants to identify which models to…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Jieyu Zhang , Weikai Huang , Zixian Ma , Oscar Michel , Dong He , Tanmay Gupta , Wei-Chiu Ma , Ali Farhadi , Aniruddha Kembhavi , Ranjay Krishna

Generative models have shown a giant leap in synthesizing photo-realistic images with minimal expertise, sparking concerns about the authenticity of online information. This study aims to develop a universal AI-generated image detector…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Zihan Liu , Hanyi Wang , Yaoyu Kang , Shilin Wang

Emotion recognition capabilities in multimodal AI systems are crucial for developing culturally responsive educational technologies, yet remain underexplored for Arabic language contexts where culturally appropriate learning tools are…

Computation and Language · Computer Science 2025-09-05 Bushra Asseri , Estabraq Abdelaziz , Maha Al Mogren , Tayef Alhefdhi , Areej Al-Wabil

Generative AI, such as OpenAI's GPT-4V large-language model, has rapidly entered mainstream discourse. Novel capabilities in image processing and natural-language communication may augment existing forecasting methods. Large language models…

Multimodal Large Language Models (MLLMs) like GPT-4V are capable of reasoning across text and image modalities, showing promise in a variety of complex vision-language tasks. In this preliminary study, we investigate the out-of-the-box…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Souradip Nath

We evaluated the capability of generative pre-trained transformers~(GPT-4) in analysis of textual data in tasks that require highly specialized domain expertise. Specifically, we focused on the task of analyzing court opinions to interpret…

Computation and Language · Computer Science 2023-10-05 Jaromir Savelka , Kevin D. Ashley , Morgan A Gray , Hannes Westermann , Huihui Xu

While large language models (LLMs) have revolutionized natural language processing with their task-agnostic capabilities, visual generation tasks such as image translation, style transfer, and character customization still rely heavily on…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Lianghua Huang , Wei Wang , Zhi-Fan Wu , Huanzhang Dou , Yupeng Shi , Yutong Feng , Chen Liang , Yu Liu , Jingren Zhou