中文
相关论文

相关论文: Detecting Legend Items on Historical Maps Using GP…

200 篇论文

Historical maps provide valuable information and knowledge about the past. However, as they often feature non-standard projections, hand-drawn styles, and artistic elements, it is challenging for non-experts to identify and interpret them.…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Ziyi Liu , Claudio Affolter , Sidi Wu , Yizi Chen , Lorenz Hurni

Text on historical maps provides valuable information for studies in history, economics, geography, and other related fields. Unlike structured or semi-structured documents, text on maps varies significantly in orientation, reading order,…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Yijun Lin , Rhett Olson , Junhan Wu , Yao-Yi Chiang , Jerod Weinman

Historical maps are essential resources that provide insights into the geographical landscapes of the past. They serve as valuable tools for researchers across disciplines such as history, geography, and urban studies, facilitating the…

计算机视觉与模式识别 · 计算机科学 2025-04-16 Yunshuang Yuan , Monika Sester

Many historical map sheets are publicly available for studies that require long-term historical geographic data. The cartographic design of these maps includes a combination of map symbols and text labels. Automatically reading text labels…

计算机视觉与模式识别 · 计算机科学 2021-12-14 Zekun Li , Runyu Guan , Qianmu Yu , Yao-Yi Chiang , Craig A. Knoblock

A method of finding and classifying various components and objects in a design diagram, drawing, or planning layout is proposed. The method automatically finds the objects present in a legend table and finds their position, count and…

计算机视觉与模式识别 · 计算机科学 2022-04-29 Sourish Sarkar , Pranav Pandey , Sibsambhu Kar

Effective building pattern recognition is critical for understanding urban form, automating map generalization, and visualizing 3D city models. Most existing studies use object-independent methods based on visual perception rules and…

计算机视觉与模式识别 · 计算机科学 2023-04-20 Zhiwei Wei , Yi Xiao , Wenjia Xu , Mi Shu , Lu Cheng , Yang Wang , Chunbo Liu

Humans effortlessly identify objects by leveraging a rich understanding of the surrounding scene, including spatial relationships, material properties, and the co-occurrence of other objects. In contrast, most computational object…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Ciprian Constantinescu , Marius Leordeanu

Large language models (LLMs) demonstrate extraordinary abilities in a wide range of natural language processing (NLP) tasks. In this paper, we show that, beyond text understanding capability, LLMs are capable of processing text layouts that…

计算与语言 · 计算机科学 2024-08-29 Weiming Li , Manni Duan , Dong An , Yan Shao

Urban traffic management faces significant challenges due to the dynamic environments, and traditional algorithms fail to quickly adapt to this environment in real-time and predict possible conflicts. This study explores the ability of a…

计算与语言 · 计算机科学 2024-08-05 Sari Masri , Huthaifa I. Ashqar , Mohammed Elhenawy

Extraction of text regions and individual text lines from historic documents is necessary for automatic transcription. We propose extending a CNN-based text baseline detection system by adding line height and text block boundary predictions…

计算机视觉与模式识别 · 计算机科学 2021-02-24 Oldřich Kodym , Michal Hradiš

Document layout analysis is a critical preprocessing step in document intelligence, enabling the detection and localization of structural elements such as titles, text blocks, tables, and formulas. Despite its importance, existing layout…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Ting Sun , Cheng Cui , Yuning Du , Yi Liu

We perform a missing, reproducible evaluation of all publicly available GPT-4 family models concerning the Document Understanding field, where it is frequently required to comprehend text spacial arrangement and visual clues in addition to…

计算与语言 · 计算机科学 2024-05-29 Łukasz Borchmann

In this research short, we examine the potential of using GPT-4o, a state-of-the-art large language model (LLM) to undertake evidence synthesis and systematic assessment tasks. Traditional workflows for such tasks involve large groups of…

计算与语言 · 计算机科学 2024-07-19 Elphin Tom Joe , Sai Dileep Koneru , Christine J Kirchhoff

Text on historical maps contains valuable information providing georeferenced historical, political, and cultural contexts. However, text extraction from historical maps is challenging due to the lack of (1) effective methods and (2)…

计算机视觉与模式识别 · 计算机科学 2025-06-19 Yijun Lin , Yao-Yi Chiang

While document layout analysis for Latin scripts has advanced significantly, driven by the advent of large multimodal models (LMMs), progress for the Khmer language remains constrained because of the scarcity of annotated training data.…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Marry Kong , Rina Buoy , Sovisal Chenda , Nguonly Taing , Masakazu Iwamura , Koichi Kise

Applying AI foundation models directly to geospatial datasets remains challenging due to their limited ability to represent and reason with geographical entities, specifically vector-based geometries and natural language descriptions of…

计算与语言 · 计算机科学 2025-05-26 Yuhan Ji , Song Gao , Ying Nie , Ivan Majić , Krzysztof Janowicz

The detection of semantic relationships between objects represented in an image is one of the fundamental challenges in image interpretation. Neural-Symbolic techniques, such as Logic Tensor Networks (LTNs), allow the combination of…

计算机视觉与模式识别 · 计算机科学 2021-09-21 Francesco Manigrasso , Filomeno Davide Miro , Lia Morra , Fabrizio Lamberti

Large language models (LLMs) have shown remarkable capabilities across a broad range of tasks involving question answering and the generation of coherent text and code. Comprehensively understanding the strengths and weaknesses of LLMs is…

计算与语言 · 计算机科学 2023-06-02 Jonathan Roberts , Timo Lüddecke , Sowmen Das , Kai Han , Samuel Albanie

Document layout analysis involves understanding the arrangement of elements within a document. This paper navigates the complexities of understanding various elements within document images, such as text, images, tables, and headings. The…

计算机视觉与模式识别 · 计算机科学 2024-05-02 Tahira Shehzadi , Didier Stricker , Muhammad Zeshan Afzal

Traffic control in unsignalized urban intersections presents significant challenges due to the complexity, frequent conflicts, and blind spots. This study explores the capability of leveraging Multimodal Large Language Models (MLLMs), such…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Sari Masri , Huthaifa I. Ashqar , Mohammed Elhenawy
‹ 上一页 1 2 3 10 下一页 ›