English
Related papers

Related papers: ICDAR 2025 Competition on End-to-End Document Imag…

200 papers

This paper provides a review of the NTIRE 2025 challenge on real-world face restoration, highlighting the proposed solutions and the resulting outcomes. The challenge focuses on generating natural, realistic outputs while maintaining…

Large Vision-Language Models (LVLMs) have achieved remarkable performance in many vision-language tasks, yet their capabilities in fine-grained visual understanding remain insufficiently evaluated. Existing benchmarks either contain limited…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Fengbin Zhu , Ziyang Liu , Xiang Yao Ng , Haohui Wu , Wenjie Wang , Fuli Feng , Chao Wang , Huanbo Luan , Tat Seng Chua

Document image retrieval (DIR) aims to retrieve document images from a gallery according to a given query. Existing DIR methods are primarily based on image queries that retrieve documents within the same coarse semantic category, e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Hao Guo , Xugong Qin , Jun Jie Ou Yang , Peng Zhang , Gangyan Zeng , Yubo Li , Hailun Lin

Multimodal entity linking plays a crucial role in a wide range of applications. Recent advances in large language model-based methods have become the dominant paradigm for this task, effectively leveraging both textual and visual modalities…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Ziyan Liu , Junwen Li , Kaiwen Li , Tong Ruan , Chao Wang , Xinyan He , Zongyu Wang , Xuezhi Cao , Jingping Liu

Document images can be affected by many degradation scenarios, which cause recognition and processing difficulties. In this age of digitization, it is important to denoise them for proper usage. To address this challenge, we present a new…

Computer Vision and Pattern Recognition · Computer Science 2022-01-26 Mohamed Ali Souibgui , Sanket Biswas , Sana Khamekhem Jemni , Yousri Kessentini , Alicia Fornés , Josep Lladós , Umapada Pal

Existing document-level neural machine translation (NMT) models have sufficiently explored different context settings to provide guidance for target generation. However, little attention is paid to inaugurate more diverse context for…

Computation and Language · Computer Science 2022-01-06 Xu Zhang , Jian Yang , Haoyang Huang , Shuming Ma , Dongdong Zhang , Jinlong Li , Furu Wei

In this paper, we offer a preliminary investigation into the task of in-image machine translation: transforming an image containing text in one language into an image containing the same text in another language. We propose an end-to-end…

Computation and Language · Computer Science 2020-10-22 Elman Mansimov , Mitchell Stern , Mia Chen , Orhan Firat , Jakob Uszkoreit , Puneet Jain

This paper reviews the challenge on Sparse Neural Rendering that was part of the Advances in Image Manipulation (AIM) workshop, held in conjunction with ECCV 2024. This manuscript focuses on the competition set-up, the proposed methods and…

Text image machine translation (TIMT) has been widely used in various real-world applications, which translates source language texts in images into another target language sentence. Existing methods on TIMT are mainly divided into two…

Computation and Language · Computer Science 2023-05-11 Cong Ma , Yaping Zhang , Mei Tu , Yang Zhao , Yu Zhou , Chengqing Zong

Idiomatic expressions present a unique challenge in NLP, as their meanings are often not directly inferable from their constituent words. Despite recent advancements in Large Language Models (LLMs), idiomaticity remains a significant…

Computation and Language · Computer Science 2025-06-05 Thomas Pickard , Aline Villavicencio , Maggie Mi , Wei He , Dylan Phelps , Marco Idiart

Understanding document images (e.g., invoices) is a core but challenging task since it requires complex functions such as reading text and a holistic understanding of the document. Current Visual Document Understanding (VDU) methods…

The goal of visual word sense disambiguation is to find the image that best matches the provided description of the word's meaning. It is a challenging problem, requiring approaches that combine language and image understanding. In this…

Computation and Language · Computer Science 2023-04-17 Sławomir Dadas

This work summarizes the IJCB Occluded Face Recognition Competition 2022 (IJCB-OCFR-2022) embraced by the 2022 International Joint Conference on Biometrics (IJCB 2022). OCFR-2022 attracted a total of 3 participating teams, from academia.…

We present ForMaT (Format-Preserving Multilingual Translation), a parallel corpus of 3,956 PDFs across 15 language pairs that preserves original layout metadata proposed for multimodal machine translation. To ensure structural diversity in…

Computation and Language · Computer Science 2026-05-18 Michał Ciesiółka , Dawid Wiśniewski , Adrian Charkiewicz , Kamil Guttmann

In today's technological era, document images play an important and integral part in our day to day life, and specifically with the surge of Covid-19, digitally scanned documents have become key source of communication, thus avoiding any…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 Dikshit Sharma , Mohammed Javed

An automatic document classification system is presented that detects textual content in images and classifies documents into four predefined categories (Invoice, Report, Letter, and Form). The system supports both offline images (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Aya Kaysan Bahjat

This paper presents a comprehensive review of the NTIRE 2025 Challenge on Single-Image Efficient Super-Resolution (ESR). The challenge aimed to advance the development of deep models that optimize key computational metrics, i.e., runtime,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Bin Ren , Hang Guo , Lei Sun , Zongwei Wu , Radu Timofte , Yawei Li , Yao Zhang , Xinning Chai , Zhengxue Cheng , Yingsheng Qin , Yucai Yang , Li Song , Hongyuan Yu , Pufan Xu , Cheng Wan , Zhijuan Huang , Peng Guo , Shuyuan Cui , Chenjun Li , Xuehai Hu , Pan Pan , Xin Zhang , Heng Zhang , Qing Luo , Linyan Jiang , Haibo Lei , Qifang Gao , Yaqing Li , Weihua Luo , Tsing Li , Qing Wang , Yi Liu , Yang Wang , Hongyu An , Liou Zhang , Shijie Zhao , Lianhong Song , Long Sun , Jinshan Pan , Jiangxin Dong , Jinhui Tang , Jing Wei , Mengyang Wang , Ruilong Guo , Qian Wang , Qingliang Liu , Yang Cheng , Davinci , Enxuan Gu , Pinxin Liu , Yongsheng Yu , Hang Hua , Yunlong Tang , Shihao Wang , Yukun Yang , Zhiyu Zhang , Yukun Yang , Jiyu Wu , Jiancheng Huang , Yifan Liu , Yi Huang , Shifeng Chen , Rui Chen , Yi Feng , Mingxi Li , Cailu Wan , Xiangji Wu , Zibin Liu , Jinyang Zhong , Kihwan Yoon , Ganzorig Gankhuyag , Shengyun Zhong , Mingyang Wu , Renjie Li , Yushen Zuo , Zhengzhong Tu , Zongang Gao , Guannan Chen , Yuan Tian , Wenhui Chen , Weijun Yuan , Zhan Li , Yihang Chen , Yifan Deng , Ruting Deng , Yilin Zhang , Huan Zheng , Yanyan Wei , Wenxuan Zhao , Suiyi Zhao , Fei Wang , Kun Li , Yinggan Tang , Mengjie Su , Jae-hyeon Lee , Dong-Hyeop Son , Ui-Jin Choi , Tiancheng Shao , Yuqing Zhang , Mengcheng Ma , Donggeun Ko , Youngsang Kwak , Jiun Lee , Jaehwa Kwak , Yuxuan Jiang , Qiang Zhu , Siyue Teng , Fan Zhang , Shuyuan Zhu , Bing Zeng , David Bull , Jing Hu , Hui Deng , Xuan Zhang , Lin Zhu , Qinrui Fan , Weijian Deng , Junnan Wu , Wenqin Deng , Yuquan Liu , Zhaohong Xu , Jameer Babu Pinjari , Kuldeep Purohit , Zeyu Xiao , Zhuoyuan Li , Surya Vashisth , Akshay Dudhane , Praful Hambarde , Sachin Chaudhary , Satya Naryan Tazi , Prashant Patil , Santosh Kumar Vipparthi , Subrahmanyam Murala , Wei-Chen Shen , I-Hsiang Chen , Yunzhe Xu , Chen Zhao , Zhizhou Chen , Akram Khatami-Rizi , Ahmad Mahmoudi-Aznaveh , Alejandro Merino , Bruno Longarela , Javier Abad , Marcos V. Conde , Simone Bianco , Luca Cogo , Gianmarco Corti

Understanding documents with rich layouts is an essential step towards information extraction. Business intelligence processes often require the extraction of useful semantic content from documents at a large scale for subsequent…

Computer Vision and Pattern Recognition · Computer Science 2022-09-22 Sanket Biswas , Ayan Banerjee , Josep Lladós , Umapada Pal

Currently, Large Language Models (LLMs) have achieved remarkable results in machine translation. However, their performance in multi-domain translation (MDT) is less satisfactory, the meanings of words can vary across different domains,…

Computation and Language · Computer Science 2026-03-17 Zhibo Man , Yuanmeng Chen , Yujie Zhang , Jinan Xu

While Vision-Language Models (VLMs) achieve near-perfect scores on digital document benchmarks like OmniDocBench, their performance in the unpredictable physical world remains largely unknown due to the lack of controlled yet realistic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Changda Zhou , Ziyue Gao , Xueqing Wang , Tingquan Gao , Cheng Cui , Jing Tang , Yi Liu