English
Related papers

Related papers: Exploring OCR Capabilities of GPT-4V(ision) : A Qu…

200 papers

There is a growing interest in applying large language models (LLMs) in robotic tasks, due to their remarkable reasoning ability and extensive knowledge learned from vast training corpora. Grounding LLMs in the physical world remains an…

Robotics · Computer Science 2024-04-11 Wenqiang Lai , Yuan Gao , Tin Lun Lam

Large Multimodal Model (LMM) GPT-4V(ision) endows GPT-4 with visual grounding capabilities, making it possible to handle certain tasks through the Visual Question Answering (VQA) paradigm. This paper explores the potential of VQA-oriented…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Jiangning Zhang , Haoyang He , Xuhai Chen , Zhucun Xue , Yabiao Wang , Chengjie Wang , Lei Xie , Yong Liu

Multimodal large language models (MLLMs) have shown remarkable capabilities across a broad range of tasks but their knowledge and abilities in the geographic and geospatial domains are yet to be explored, despite potential wide-ranging…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Jonathan Roberts , Timo Lüddecke , Rehan Sheikh , Kai Han , Samuel Albanie

The recent development on large multimodal models (LMMs), especially GPT-4V(ision) and Gemini, has been quickly expanding the capability boundaries of multimodal models beyond traditional tasks like image captioning and visual question…

Information Retrieval · Computer Science 2024-03-14 Boyuan Zheng , Boyu Gou , Jihyung Kil , Huan Sun , Yu Su

We present Multimodal OCR (MOCR), a document parsing paradigm that jointly parses text and graphics into unified textual representations. Unlike conventional OCR systems that focus on text recognition and leave graphical regions as cropped…

While Large Vision-Language Models (LVLMs) demonstrate promising multilingual capabilities, their evaluation is currently hindered by two critical limitations: (1) the use of non-parallel corpora, which conflates inherent language…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Junyuan Gao , Jiahe Song , Jiang Wu , Runchuan Zhu , Guanlin Shen , Shasha Wang , Xingjian Wei , Haote Yang , Songyang Zhang , Weijia Li , Bin Wang , Dahua Lin , Lijun Wu , Conghui He

Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur,…

This study presents a comprehensive evaluation of GPT-4's translation capabilities compared to human translators of varying expertise levels. Through systematic human evaluation using the MQM schema, we assess translations across three…

Computation and Language · Computer Science 2024-11-22 Jianhao Yan , Pingchuan Yan , Yulong Chen , Jing Li , Xianchao Zhu , Yue Zhang

PDF documents have the potential to provide trillions of novel, high-quality tokens for training language models. However, these documents come in a diversity of types with differing formats and visual layouts that pose a challenge when…

Retrieving accurate details from documents is a crucial task, especially when handling a combination of scanned images and native digital formats. This document presents a combined framework for text extraction that merges Optical Character…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Rasha Sinha , Rekha B S

Information extraction from copy-heavy documents, characterized by massive volumes of structurally similar content, represents a critical yet understudied challenge in enterprise document processing. We present a systematic framework that…

Computation and Language · Computer Science 2025-10-14 Zilong Wang , Xiaoyu Shen

We present TextMonkey, a large multimodal model (LMM) tailored for text-centric tasks. Our approach introduces enhancement across several dimensions: By adopting Shifted Window Attention with zero-initialization, we achieve cross-window…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Yuliang Liu , Biao Yang , Qiang Liu , Zhang Li , Zhiyin Ma , Shuo Zhang , Xiang Bai

Conventional Optical Character Recognition (OCR) systems are challenged by variant invoice layouts, handwritten text, and low-quality scans, which are often caused by strong template dependencies that restrict their flexibility across…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Khushi Khanchandani , Advait Thakur , Akshita Shetty , Chaitravi Reddy , Ritisa Behera

Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior quality and capabilities across diverse tasks. However, their…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Xuelu Feng , Yunsheng Li , Dongdong Chen , Mei Gao , Mengchen Liu , Junsong Yuan , Chunming Qiao

Recent developments in multimodal large language models (MLLMs) have spurred significant interest in their potential applications across various medical imaging domains. On the one hand, there is a temptation to use these generative models…

Image and Video Processing · Electrical Eng. & Systems 2024-06-05 Sulaiman Khan , Md. Rafiul Biswas , Alina Murad , Hazrat Ali , Zubair Shah

Safely navigating street intersections is a complex challenge for blind and low-vision individuals, as it requires a nuanced understanding of the surrounding context - a task heavily reliant on visual cues. Traditional methods for assisting…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Hochul Hwang , Sunjae Kwon , Yekyung Kim , Donghyun Kim

Natural language is a powerful complementary modality of communication for data visualizations, such as bar and line charts. To facilitate chart-based reasoning using natural language, various downstream tasks have been introduced recently…

Computation and Language · Computer Science 2024-10-07 Mohammed Saidul Islam , Raian Rahman , Ahmed Masry , Md Tahmid Rahman Laskar , Mir Tafseer Nayeem , Enamul Hoque

Leveraging Large Multimodal Models (LMMs) to simulate human behaviors when processing multimodal information, especially in the context of social media, has garnered immense interest due to its broad potential and far-reaching implications.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Hanjia Lyu , Weihong Qi , Zhongyu Wei , Jiebo Luo

Optical character recognition (OCR) technology has been widely used in various scenes, as shown in Figure 1. Designing a practical OCR system is still a meaningful but challenging task. In previous work, considering the efficiency and…

Computer Vision and Pattern Recognition · Computer Science 2022-06-15 Chenxia Li , Weiwei Liu , Ruoyu Guo , Xiaoting Yin , Kaitao Jiang , Yongkun Du , Yuning Du , Lingfeng Zhu , Baohua Lai , Xiaoguang Hu , Dianhai Yu , Yanjun Ma

This study utilizes the advanced capabilities of the GPT-4 multimodal Large Language Model (LLM) to explore its potential in iris recognition - a field less common and more specialized than face recognition. By focusing on this niche yet…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Parisa Farmanifard , Arun Ross