English
Related papers

Related papers: ICDAR 2025 Competition on End-to-End Document Imag…

200 papers

In this paper, we present the Multi-Forgery Detection Challenge held concurrently with the IEEE Computer Society Workshop on Biometrics at CVPR 2022. Our Multi-Forgery Detection Challenge aims to detect automatic image manipulations…

Computer Vision and Pattern Recognition · Computer Science 2022-07-28 Jianshu Li , Man Luo , Jian Liu , Tao Chen , Chengjie Wang , Ziwei Liu , Shuo Liu , Kewei Yang , Xuning Shao , Kang Chen , Boyuan Liu , Mingyu Guo , Ying Guo , Yingying Ao , Pengfei Gao

Text recognition in natural images remains a challenging yet essential task, with broad applications spanning computer vision and natural language processing. This paper introduces a novel end-to-end framework that combines ResNet and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Naphat Nithisopa , Teerapong Panboonyuen

The Image Difference Captioning (IDC) task aims to describe the visual differences between two similar images with natural language. The major challenges of this task lie in two aspects: 1) fine-grained visual differences that require…

Multimedia · Computer Science 2022-02-10 Linli Yao , Weiying Wang , Qin Jin

This paper reviews the NTIRE 2024 low light image enhancement challenge, highlighting the proposed solutions and results. The aim of this challenge is to discover an effective network design or solution capable of generating brighter,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Xiaoning Liu , Zongwei Wu , Ao Li , Florin-Alexandru Vasluianu , Yulun Zhang , Shuhang Gu , Le Zhang , Ce Zhu , Radu Timofte , Zhi Jin , Hongjun Wu , Chenxi Wang , Haitao Ling , Yuanhao Cai , Hao Bian , Yuxin Zheng , Jing Lin , Alan Yuille , Ben Shao , Jin Guo , Tianli Liu , Mohao Wu , Yixu Feng , Shuo Hou , Haotian Lin , Yu Zhu , Peng Wu , Wei Dong , Jinqiu Sun , Yanning Zhang , Qingsen Yan , Wenbin Zou , Weipeng Yang , Yunxiang Li , Qiaomu Wei , Tian Ye , Sixiang Chen , Zhao Zhang , Suiyi Zhao , Bo Wang , Yan Luo , Zhichao Zuo , Mingshen Wang , Junhu Wang , Yanyan Wei , Xiaopeng Sun , Yu Gao , Jiancheng Huang , Hongming Chen , Xiang Chen , Hui Tang , Yuanbin Chen , Yuanbo Zhou , Xinwei Dai , Xintao Qiu , Wei Deng , Qinquan Gao , Tong Tong , Mingjia Li , Jin Hu , Xinyu He , Xiaojie Guo , Sabarinathan , K Uma , A Sasithradevi , B Sathya Bama , S. Mohamed Mansoor Roomi , V. Srivatsav , Jinjuan Wang , Long Sun , Qiuying Chen , Jiahong Shao , Yizhi Zhang , Marcos V. Conde , Daniel Feijoo , Juan C. Benito , Alvaro García , Jaeho Lee , Seongwan Kim , Sharif S M A , Nodirkhuja Khujaev , Roman Tsoy , Ali Murtaza , Uswah Khairuddin , Ahmad 'Athif Mohd Faudzi , Sampada Malagi , Amogh Joshi , Nikhil Akalwadi , Chaitra Desai , Ramesh Ashok Tabib , Uma Mudenagudi , Wenyi Lian , Wenjing Lian , Jagadeesh Kalyanshetti , Vijayalaxmi Ashok Aralikatti , Palani Yashaswini , Nitish Upasi , Dikshit Hegde , Ujwala Patil , Sujata C , Xingzhuo Yan , Wei Hao , Minghan Fu , Pooja choksy , Anjali Sarvaiya , Kishor Upla , Kiran Raja , Hailong Yan , Yunkai Zhang , Baiang Li , Jingyi Zhang , Huan Zheng

We propose In-Context Translation (ICT), a general learning framework to unify visual recognition (e.g., semantic segmentation), low-level image processing (e.g., denoising), and conditional image generation (e.g., edge-to-image synthesis).…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Han Xue , Qianru Sun , Li Song , Wenjun Zhang , Zhiwu Huang

In this paper, we present our solution to a Multi-modal Algorithmic Reasoning Task: SMART-101 Challenge. Different from the traditional visual question-answering datasets, this challenge evaluates the abstraction, deduction, and…

Computer Vision and Pattern Recognition · Computer Science 2023-10-11 Xiangyu Wu , Yang Yang , Shengdong Xu , Yifeng Wu , Qingguo Chen , Jianfeng Lu

In-Image Machine Translation (IIMT) powers cross-border e-commerce product listings; existing research focuses on machine translation evaluation, while visual rendering quality is critical for user engagement. When facing context-dense…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Qingyu Wu , Yuxuan Han , Haijun Li , Zhao Xu , Jianshan Zhao , Xu Jin , Longyue Wang , Weihua Luo

Over the past few years, image-to-image (I2I) translation methods have been proposed to translate a given image into diverse outputs. Despite the impressive results, they mainly focus on the I2I translation between two domains, so the…

Computer Vision and Pattern Recognition · Computer Science 2022-02-08 Somi Jeong , Jiyoung Lee , Kwanghoon Sohn

We introduce INQUIRE, a text-to-image retrieval benchmark designed to challenge multimodal vision-language models on expert-level queries. INQUIRE includes iNaturalist 2024 (iNat24), a new dataset of five million natural world images, along…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Edward Vendrow , Omiros Pantazis , Alexander Shepard , Gabriel Brostow , Kate E. Jones , Oisin Mac Aodha , Sara Beery , Grant Van Horn

Image-text matching (ITM) aims to address the fundamental challenge of aligning visual and textual modalities, which inherently differ in their representations, continuous, high-dimensional image features vs. discrete, structured text. We…

Multimedia · Computer Science 2025-07-14 Junyu Chen , Yihua Gao , Mingyong Li

Document parsing from scanned images into structured formats remains a significant challenge due to its complexly intertwined elements such as text paragraphs, figures, formulas, and tables. Existing supervised fine-tuning methods often…

Computation and Language · Computer Science 2025-10-21 Baode Wang , Biao Wu , Weizhen Li , Meng Fang , Zuming Huang , Jun Huang , Haozhe Wang , Yanjie Liang , Ling Chen , Wei Chu , Yuan Qi

This paper introduces the system submitted by dun_oscar team for the ICPR MSR Challenge. Three subsystems for task1-task3 are descripted respectively. In task1, we develop a visual system which includes a OCR model, a text tracker, and a…

Computation and Language · Computer Science 2023-03-14 Binbin Du , Rui Deng , Yingxin Zhang

This paper describes the experimental framework and results of the ICDAR 2021 Competition on On-Line Signature Verification (SVC 2021). The goal of SVC 2021 is to evaluate the limits of on-line signature verification systems on popular…

This competition focus on Urban-Sense Segmentation based on the vehicle camera view. Class highly unbalanced Urban-Sense images dataset challenge the existing solutions and further studies. Deep Conventional neural network-based semantic…

Computer Vision and Pattern Recognition · Computer Science 2022-07-05 Akide Liu , Zihan Wang

In this paper, the solution of HYU MLLAB KT Team to the Multimodal Algorithmic Reasoning Task: SMART-101 CVPR 2024 Challenge is presented. Beyond conventional visual question-answering problems, the SMART-101 challenge aims to achieve…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Jinwoo Ahn , Junhyeok Park , Min-Jun Kim , Kang-Hyeon Kim , So-Yeong Sohn , Yun-Ji Lee , Du-Seong Chang , Yu-Jung Heo , Eun-Sol Kim

Multimodal Machine Translation (MMT) has demonstrated the significant help of visual information in machine translation. However, existing MMT methods face challenges in leveraging the modality gap by enforcing rigid visual-linguistic…

Computation and Language · Computer Science 2025-10-09 Jiafeng Xiong , Yuting Zhao

Ultra Light OCR Competition is a Chinese scene text recognition competition jointly organized by CSIG (China Society of Image and Graphics) and Baidu, Inc. In addition to focusing on common problems in Chinese scene text recognition, such…

Computer Vision and Pattern Recognition · Computer Science 2021-10-26 Shuhan Zhang , Yuxin Zou , Tianhe Wang , Yichao Xiong

Understanding documents with rich layouts and multi-modal components is a long-standing and practical task. Recent Large Vision-Language Models (LVLMs) have made remarkable strides in various tasks, particularly in single-page document…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Yubo Ma , Yuhang Zang , Liangyu Chen , Meiqi Chen , Yizhu Jiao , Xinze Li , Xinyuan Lu , Ziyu Liu , Yan Ma , Xiaoyi Dong , Pan Zhang , Liangming Pan , Yu-Gang Jiang , Jiaqi Wang , Yixin Cao , Aixin Sun

Previous works have shown that contextual information can improve the performance of neural machine translation (NMT). However, most existing document-level NMT methods only consider a few number of previous sentences. How to make use of…

Computation and Language · Computer Science 2021-09-15 Mingzhou Xu , Liangyou Li , Derek. F. Wong , Qun Liu , Lidia S. Chao

In recent years, text-image joint pre-training techniques have shown promising results in various tasks. However, in Optical Character Recognition (OCR) tasks, aligning text instances with their corresponding text regions in images poses a…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Chen Duan , Pei Fu , Shan Guo , Qianyi Jiang , Xiaoming Wei
‹ Prev 1 8 9 10 Next ›