English
Related papers

Related papers: GutenOCR: A Grounded Vision-Language Front-End for…

200 papers

We focus on the confounding bias between language and location in the visual grounding pipeline, where we find that the bias is the major visual reasoning bottleneck. For example, the grounding process is usually a trivial language-location…

Computer Vision and Pattern Recognition · Computer Science 2022-01-03 Jianqiang Huang , Yu Qin , Jiaxin Qi , Qianru Sun , Hanwang Zhang

Accurate and interpretable plant disease diagnosis remains a major challenge for vision-language models (VLMs) in real-world agriculture. We introduce AgriChain, a dataset of approximately 11,000 expert-curated leaf images spanning diverse…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Hazza Mahmood , Yongqiang Yu , Rao Anwer

To effectively apply robots in working environments and assist humans, it is essential to develop and evaluate how visual grounding (VG) can affect machine performance on occluded objects. However, current VG works are limited in working…

Computation and Language · Computer Science 2021-04-15 Ke-Jyun Wang , Yun-Hsuan Liu , Hung-Ting Su , Jen-Wei Wang , Yu-Siang Wang , Winston H. Hsu , Wen-Chin Chen

Clinical documentation can contain emotionally charged language with stigmatizing or privileging valences. We present a framework for detecting and classifying such language as stigmatizing, privileging, or neutral. We constructed a curated…

Optical character recognition (OCR) has advanced rapidly with deep learning and multimodal models, yet most methods focus on well-resourced scripts such as Latin and Chinese. Ethnic minority languages remain underexplored due to complex…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Bonan Liu , Zeyu Zhang , Bingbing Meng , Han Wang , Hanshuo Zhang , Chengping Wang , Daji Ergu , Ying Cai

Video Temporal Grounding (VTG) aims to localize temporal segments in long, untrimmed videos that align with a given natural language query. This task typically comprises two subtasks: Moment Retrieval (MR) and Highlight Detection (HD).…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Minseok Kang , Minhyeok Lee , Minjung Kim , Donghyeong Kim , Sangyoun Lee

Recently, 3D vision-and-language tasks have attracted increasing research interest. Compared to other vision-and-language tasks, the 3D visual question answering (VQA) task is less exploited and is more susceptible to language priors and…

Computer Vision and Pattern Recognition · Computer Science 2022-09-27 Lichen Zhao , Daigang Cai , Jing Zhang , Lu Sheng , Dong Xu , Rui Zheng , Yinjie Zhao , Lipeng Wang , Xibo Fan

We propose a learning system in which language is grounded in visual percepts without specific pre-defined categories of terms. We present a unified generative method to acquire a shared semantic/visual embedding that enables the learning…

Computation and Language · Computer Science 2021-08-02 Nisha Pillai , Cynthia Matuszek , Francis Ferraro

Visual grounding, a crucial vision-language task involving the understanding of the visual context based on the query expression, necessitates the model to capture the interactions between objects, as well as various spatial and attribute…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Haozhan Shen , Tiancheng Zhao , Mingwei Zhu , Jianwei Yin

This paper presents our methodology and findings from three tasks across Optical Character Recognition (OCR) and Document Layout Analysis using advanced deep learning techniques. First, for the historical Hebrew fragments of the Dead Sea…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Hylke Westerdijk , Ben Blankenborg , Khondoker Ittehadul Islam

We investigate OCR-augmented generation with Vision Language Models (VLMs), exploring tasks in Korean and English toward multilingualism. To support research in this domain, we train and release KLOCR, a strong bilingual OCR baseline…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 JoonHo Lee , Sunho Park

Vision-and-Language Navigation (VLN) requires an embodied agent to traverse complex environments by following natural language instructions, demanding accurate alignment between visual observations and linguistic guidance. Despite recent…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Yaohua Liu , Xinyuan Song , Yunfu Deng , Yifan Xie , Binkai Ou , Yan Zhong

Video Anomaly Detection (VAD) has traditionally been framed as binary classification or outlier detection, providing neither interpretable reasoning nor precise spatial localization of anomalous events. While Vision-Language Models (VLMs)…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Sakshi Agarwal , Aishik Konwer , Ankit Parag Shah

Optical Coherence Tomography (OCT) is essential for diagnosing conditions such as glaucoma, diabetic retinopathy, and age-related macular degeneration. Accurate retinal layer segmentation enables quantitative biomarkers critical for…

Image and Video Processing · Electrical Eng. & Systems 2025-09-10 S M Asiful Islam Saky , Ugyen Tshering

We introduce VARCO-VISION-2.0, an open-weight bilingual vision-language model (VLM) for Korean and English with improved capabilities compared to the previous model VARCO-VISION-14B. The model supports multi-image understanding for complex…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Young-rok Cha , Jeongho Ju , SunYoung Park , Jong-Hyeon Lee , Younghyun Yu , Youngjune Kim

Visual grounding in 3D is the key for embodied agents to localize language-referred objects in open-world environments. However, existing benchmarks are limited to indoor focus, single-platform constraints, and small scale. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Rong Li , Yuhao Dong , Tianshuai Hu , Ao Liang , Youquan Liu , Dongyue Lu , Liang Pan , Lingdong Kong , Junwei Liang , Ziwei Liu

This paper introduces ChineseErrorCorrector3-4B, a unified model for Chinese spelling and grammatical error correction based on Qwen3-4B. The model demonstrates outstanding performance in general text correction tasks and achieves…

Computation and Language · Computer Science 2025-11-25 Wei Tian , YuhaoZhou

Outdoor advertisements remain a critical medium for modern marketing, yet accurately verifying billboard text visibility under real-world conditions is still challenging. Traditional Optical Character Recognition (OCR) pipelines excel at…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Maciej Szankin , Vidhyananth Venkatasamy , Lihang Ying

Multimodal large language models (MLLMs) have achieved remarkable performance across a wide range of vision language tasks. However, their ability in low-level visual perception, particularly in detecting fine-grained visual discrepancies,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Tengjin Weng , Wenhao Jiang , Jingyi Wang , Ming Li , Lin Ma , Zhong Ming

Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and understanding, i.e., traditional methods only scan images, while…