中文
相关论文

相关论文: Exploring OCR Capabilities of GPT-4V(ision) : A Qu…

200 篇论文

This paper explores capabilities of Vision Language Models on spreadsheet comprehension. We propose three self-supervised challenges with corresponding evaluation metrics to comprehensively evaluate VLMs on Optical Character Recognition…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Shiyu Xia , Junyu Xiong , Haoyu Dong , Jianbo Zhao , Yuzhang Tian , Mengyu Zhou , Yeye He , Shi Han , Dongmei Zhang

Recent developments in multimodal methodologies have marked the beginning of an exciting era for models adept at processing diverse data types, encompassing text, audio, and visual content. Models like GPT-4V, which merge computer vision…

Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their…

This work conducts an evaluation of GPT-4V's multimodal capability for medical image analysis, with a focus on three representative tasks of radiology report generation, medical visual question answering, and medical visual grounding. For…

计算机视觉与模式识别 · 计算机科学 2024-02-01 Yingshu Li , Yunyi Liu , Zhanyu Wang , Xinyu Liang , Lei Wang , Lingqiao Liu , Leyang Cui , Zhaopeng Tu , Longyue Wang , Luping Zhou

In this paper, we critically evaluate the capabilities of the state-of-the-art multimodal large language model, i.e., GPT-4 with Vision (GPT-4V), on Visual Question Answering (VQA) task. Our experiments thoroughly assess GPT-4V's…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Zhiling Yan , Kai Zhang , Rong Zhou , Lifang He , Xiang Li , Lichao Sun

The advent of large language models (LLMs) has heightened interest in their potential for multimodal applications that integrate language and vision. This paper explores the capabilities of GPT-4V in the realms of geography, environmental…

With the rise of multimodal large language models, accurately extracting and understanding textual information from video content, referred to as video based optical character recognition (Video OCR), has become a crucial capability. This…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Yulin Fei , Yuhui Gao , Xingyuan Xian , Xiaojin Zhang , Tao Wu , Wei Chen

Driven by the large foundation models, the development of artificial intelligence has witnessed tremendous progress lately, leading to a surge of general interest from the public. In this study, we aim to assess the performance of OpenAI's…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Chaoyi Wu , Jiayu Lei , Qiaoyu Zheng , Weike Zhao , Weixiong Lin , Xiaoman Zhang , Xiao Zhou , Ziheng Zhao , Ya Zhang , Yanfeng Wang , Weidi Xie

Optical character recognition (OCR) and multilingual text understanding remain major failure modes of multimodal large language models (MLLMs), particularly in real-world images containing cluttered layouts, small fonts, blur, occlusion,…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Qinwu Xu , Yifan Jiang , Haoyu Ren

We perform a missing, reproducible evaluation of all publicly available GPT-4 family models concerning the Document Understanding field, where it is frequently required to comprehend text spacial arrangement and visual clues in addition to…

计算与语言 · 计算机科学 2024-05-29 Łukasz Borchmann

The recognition and understanding of traffic incidents, particularly traffic accidents, is a topic of paramount importance in the realm of intelligent transportation systems and intelligent vehicles. This area has continually captured the…

计算机视觉与模式识别 · 计算机科学 2024-02-08 Xingcheng Zhou , Alois C. Knoll

Due to their high versatility in tasks such as image captioning, document analysis, and automated content generation, multimodal Large Language Models (LLMs) have attracted significant attention across various industrial fields. In…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Kotaro Inoue

This study investigates the potential of Large Language Models (LLMs), particularly GPT-4o, for Optical Character Recognition (OCR) in low-resource scripts such as Urdu, Albanian, and Tajik, with English serving as a benchmark. Using a…

机器学习 · 计算机科学 2024-12-23 Muhammad Abdullah Sohail , Salaar Masood , Hamza Iqbal

Optical Character Recognition (OCR) for data extraction from documents is essential to intelligent informatics, such as digitizing medical records and recognizing road signs. Multi-modal Large Language Models (LLMs) can solve this task and…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Hyakka Nakada , Yoshiyasu Tanaka

While OCR has been used in various applications, its output is not always accurate, leading to misfit words. This research work focuses on improving the optical character recognition (OCR) with ML techniques with integration of OCR with…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Abhishek Bamotra , Phani Krishna Uppala

Multimodal large language models (MLLMs) are designed to process and integrate information from multiple sources, such as text, speech, images, and videos. Despite its success in language understanding, it is critical to evaluate the…

计算机视觉与模式识别 · 计算机科学 2024-04-11 Hao Lu , Xuesong Niu , Jiyao Wang , Yin Wang , Qingyong Hu , Jiaqi Tang , Yuting Zhang , Kaishen Yuan , Bin Huang , Zitong Yu , Dengbo He , Shuiguang Deng , Hao Chen , Yingcong Chen , Shiguang Shan

Large Multimodal Models (LMMs) have become increasingly versatile, accompanied by impressive Optical Character Recognition (OCR) related capabilities. Existing OCR-related benchmarks emphasize evaluating LMMs' abilities of relatively simple…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Haibin He , Maoyuan Ye , Jing Zhang , Xiantao Cai , Juhua Liu , Bo Du , Dacheng Tao

While large multi-modal models (LMM) have shown notable progress in multi-modal tasks, their capabilities in tasks involving dense textual content remains to be fully explored. Dense text, which carries important information, is often found…

计算与语言 · 计算机科学 2024-05-14 Shuo Zhang , Biao Yang , Zhang Li , Zhiyin Ma , Yuliang Liu , Xiang Bai

The advent of large vision-language models (LVLMs) represents a remarkable advance in the quest for artificial general intelligence. However, the model's effectiveness in both specialized and general tasks warrants further investigation.…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Yao Jiang , Xinyu Yan , Ge-Peng Ji , Keren Fu , Meijun Sun , Huan Xiong , Deng-Ping Fan , Fahad Shahbaz Khan

OpenAI's latest large vision-language model (LVLM), GPT-4V(ision), has piqued considerable interest for its potential in medical applications. Despite its promise, recent studies and internal reviews highlight its underperformance in…

计算与语言 · 计算机科学 2023-12-13 Pengcheng Chen , Ziyan Huang , Zhongying Deng , Tianbin Li , Yanzhou Su , Haoyu Wang , Jin Ye , Yu Qiao , Junjun He