English
Related papers

Related papers: ICDAR 2025 Competition on End-to-End Document Imag…

200 papers

In-Image Machine Translation (IIMT) aims to translate texts within images from one language to another. Previous research on IIMT was primarily conducted on simplified scenarios such as images of one-line text with black font in white…

Computation and Language · Computer Science 2025-05-22 Yanzhi Tian , Zeming Liu , Zhengyang Liu , Yuhang Guo

Motion blur is one of the most common degradation artifacts in dynamic scene photography. This paper reviews the NTIRE 2020 Challenge on Image and Video Deblurring. In this challenge, we present the evaluation results from 3 competition…

Computer Vision and Pattern Recognition · Computer Science 2020-05-12 Seungjun Nah , Sanghyun Son , Radu Timofte , Kyoung Mu Lee

This work presents a comparative evaluation of machine translation systems applied to images containing textual information, a task that lies at the intersection of computer vision and natural language processing. The study compares three…

Computation and Language · Computer Science 2026-05-29 Blai Puchol , Sergio Gómez González , Miguel Domingo , Francisco Casacuberta

We introduce ABot-OCR, an end-to-end vision-language model that transcribes a page image directly into clean Markdown in a single forward pass. By doing so, our approach completely eliminates the need for brittle modular orchestration. To…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Kaitao Jiang , Ruiyan Gong , Xiaolong Cheng , Kangning Niu , Tianlun Li , Mu Xu

There has been a growing interest in developing multimodal machine translation (MMT) systems that enhance neural machine translation (NMT) with visual knowledge. This problem setup involves using images as auxiliary information during…

Computer Vision and Pattern Recognition · Computer Science 2023-08-30 Devaansh Gupta , Siddhant Kharbanda , Jiawei Zhou , Wanhua Li , Hanspeter Pfister , Donglai Wei

Recent studies on machine reading comprehension have focused on text-level understanding but have not yet reached the level of human understanding of the visual layout and content of real-world documents. In this study, we introduce a new…

Computation and Language · Computer Science 2021-05-11 Ryota Tanaka , Kyosuke Nishida , Sen Yoshida

Chinese scene text reading is one of the most challenging problems in computer vision and has attracted great interest. Different from English text, Chinese has more than 6000 commonly used characters and Chinesecharacters can be arranged…

Computer Vision and Pattern Recognition · Computer Science 2019-12-23 Xi Liu , Rui Zhang , Yongsheng Zhou , Qianyi Jiang , Qi Song , Nan Li , Kai Zhou , Lei Wang , Dong Wang , Minghui Liao , Mingkun Yang , Xiang Bai , Baoguang Shi , Dimosthenis Karatzas , Shijian Lu , C. V. Jawahar

This paper reviews the NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images. This challenge received a wide range of impressive solutions, which are developed and evaluated using our collected real-world Raindrop…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Xin Li , Yeying Jin , Xin Jin , Zongwei Wu , Bingchen Li , Yufei Wang , Wenhan Yang , Yu Li , Zhibo Chen , Bihan Wen , Robby T. Tan , Radu Timofte , Qiyu Rong , Hongyuan Jing , Mengmeng Zhang , Jinglong Li , Xiangyu Lu , Yi Ren , Yuting Liu , Meng Zhang , Xiang Chen , Qiyuan Guan , Jiangxin Dong , Jinshan Pan , Conglin Gou , Qirui Yang , Fangpu Zhang , Yunlong Lin , Sixiang Chen , Guoxi Huang , Ruirui Lin , Yan Zhang , Jingyu Yang , Huanjing Yue , Jiyuan Chen , Qiaosi Yi , Hongjun Wang , Chenxi Xie , Shuai Li , Yuhui Wu , Kaiyi Ma , Jiakui Hu , Juncheng Li , Liwen Pan , Guangwei Gao , Wenjie Li , Zhenyu Jin , Heng Guo , Zhanyu Ma , Yubo Wang , Jinghua Wang , Wangzhi Xing , Anjusree Karnavar , Diqi Chen , Mohammad Aminul Islam , Hao Yang , Ruikun Zhang , Liyuan Pan , Qianhao Luo , XinCao , Han Zhou , Yan Min , Wei Dong , Jun Chen , Taoyi Wu , Weijia Dou , Yu Wang , Shengjie Zhao , Yongcheng Huang , Xingyu Han , Anyan Huang , Hongtao Wu , Hong Wang , Yefeng Zheng , Abhijeet Kumar , Aman Kumar , Marcos V. Conde , Paula Garrido , Daniel Feijoo , Juan C. Benito , Guanglu Dong , Xin Lin , Siyuan Liu , Tianheng Zheng , Jiayu Zhong , Shouyi Wang , Xiangtai Li , Lanqing Guo , Lu Qi , Chao Ren , Shuaibo Wang , Shilong Zhang , Wanyu Zhou , Yunze Wu , Qinzhong Tan , Jieyuan Pei , Zhuoxuan Li , Jiayu Wang , Haoyu Bian , Haoran Sun , Subhajit Paul , Ni Tang , Junhao Huang , Zihan Cheng , Hongyun Zhu , Yuehan Wu , Kaixin Deng , Hang Ouyang , Tianxin Xiao , Fan Yang , Zhizun Luo , Zeyu Xiao , Zhuoyuan Li , Nguyen Pham Hoang Le , An Dinh Thien , Son T. Luu , Kiet Van Nguyen , Ronghua Xu , Xianmin Tian , Weijian Zhou , Jiacheng Zhang , Yuqian Chen , Yihang Duan , Yujie Wu , Suresh Raikwar , Arsh Garg , Kritika , Jianhua Zheng , Xiaoshan Ma , Ruolin Zhao , Yongyu Yang , Yongsheng Liang , Guiming Huang , Qiang Li , Hongbin Zhang , Xiangyu Zheng , A. N. Rajagopalan

In this paper, we design and train a Generative Image-to-text Transformer, GIT, to unify vision-language tasks such as image/video captioning and question answering. While generative models provide a consistent network architecture between…

Computer Vision and Pattern Recognition · Computer Science 2022-12-19 Jianfeng Wang , Zhengyuan Yang , Xiaowei Hu , Linjie Li , Kevin Lin , Zhe Gan , Zicheng Liu , Ce Liu , Lijuan Wang

This competition investigates the performance of large-scale retrieval of historical document images based on writing style. Based on large image data sets provided by cultural heritage institutions and digital libraries, providing a total…

Computer Vision and Pattern Recognition · Computer Science 2019-12-10 Vincent Christlein , Anguelos Nicolaou , Mathias Seuret , Dominique Stutzmann , Andreas Maier

This paper presents the NTIRE 2025 image super-resolution ($\times$4) challenge, one of the associated competitions of the 10th NTIRE Workshop at CVPR 2025. The challenge aims to recover high-resolution (HR) images from low-resolution (LR)…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Zheng Chen , Kai Liu , Jue Gong , Jingkai Wang , Lei Sun , Zongwei Wu , Radu Timofte , Yulun Zhang , Xiangyu Kong , Xiaoxuan Yu , Hyunhee Park , Suejin Han , Hakjae Jeon , Dafeng Zhang , Hyung-Ju Chun , Donghun Ryou , Inju Ha , Bohyung Han , Lu Zhao , Yuyi Zhang , Pengyu Yan , Jiawei Hu , Pengwei Liu , Fengjun Guo , Hongyuan Yu , Pufan Xu , Zhijuan Huang , Shuyuan Cui , Peng Guo , Jiahui Liu , Dongkai Zhang , Heng Zhang , Huiyuan Fu , Huadong Ma , Yanhui Guo , Sisi Tian , Xin Liu , Jinwen Liang , Jie Liu , Jie Tang , Gangshan Wu , Zeyu Xiao , Zhuoyuan Li , Yinxiang Zhang , Wenxuan Cai , Vijayalaxmi Ashok Aralikatti , Nikhil Akalwadi , G Gyaneshwar Rao , Chaitra Desai , Ramesh Ashok Tabib , Uma Mudenagudi , Marcos V. Conde , Alejandro Merino , Bruno Longarela , Javier Abad , Weijun Yuan , Zhan Li , Zhanglu Chen , Boyang Yao , Aagam Jain , Milan Kumar Singh , Ankit Kumar , Shubh Kawa , Divyavardhan Singh , Anjali Sarvaiya , Kishor Upla , Raghavendra Ramachandra , Chia-Ming Lee , Yu-Fan Lin , Chih-Chung Hsu , Risheek V Hiremath , Yashaswini Palani , Yuxuan Jiang , Qiang Zhu , Siyue Teng , Fan Zhang , Shuyuan Zhu , Bing Zeng , David Bull , Jingwei Liao , Yuqing Yang , Wenda Shao , Junyi Zhao , Qisheng Xu , Kele Xu , Sunder Ali Khowaja , Ik Hyun Lee , Snehal Singh Tomar , Rajarshi Ray , Klaus Mueller , Sachin Chaudhary , Surya Vashisth , Akshay Dudhane , Praful Hambarde , Satya Naryan Tazi , Prashant Patil , Santosh Kumar Vipparthi , Subrahmanyam Murala , Bilel Benjdira , Anas M. Ali , Wadii Boulila , Zahra Moammeri , Ahmad Mahmoudi-Aznaveh , Ali Karbasi , Hossein Motamednia , Liangyan Li , Guanhua Zhao , Kevin Le , Yimo Ning , Haoxuan Huang , Jun Chen

This paper reviews the AIM 2025 Efficient Real-World Deblurring using Single Images Challenge, which aims to advance in efficient real-blur restoration. The challenge is based on a new test set based on the well known RSBlur dataset. Pairs…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Daniel Feijoo , Paula Garrido-Mellado , Marcos V. Conde , Jaesung Rim , Alvaro Garcia , Sunghyun Cho , Radu Timofte

Applying diffusion models to image-to-image translation (I2I) has recently received increasing attention due to its practical applications. Previous attempts inject information from the source image into each denoising step for an iterative…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Mengfei Xia , Yu Zhou , Ran Yi , Yong-Jin Liu , Wenping Wang

Image Difference Captioning (IDC) aims to generate natural language descriptions of subtle differences between image pairs, requiring both precise visual change localization and coherent semantic expression. Despite recent advancements,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Yuan Liu , Saihui Hou , Saijie Hou , Jiabao Du , Shibei Meng , Yongzhen Huang

In this paper, we present the Global Multimedia Deepfake Detection held concurrently with the Inclusion 2024. Our Multimedia Deepfake Detection aims to detect automatic image and audio-video manipulations including but not limited to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Yi Zhang , Weize Gao , Changtao Miao , Man Luo , Jianshu Li , Wenzhong Deng , Zhe Li , Bingyu Hu , Weibin Yao , Yunfeng Diao , Wenbo Zhou , Tao Gong , Qi Chu

This paper presents the computational challenge on differential geometry and topology that happened within the ICLR 2021 workshop "Geometric and Topological Representation Learning". The competition asked participants to provide creative…

This draft is a working document, having a summary of nighty-four (94) papers with additional sections on Traceability of Software Requirements (Section 4), Formal Methods and Its Tools (Section 5), Unifying Theories of Programming (UTP)…

Software Engineering · Computer Science 2025-06-24 Arshad Beg , Diarmuid O'Donoghue , Rosemary Monahan

This paper presents our solution for ICDAR 2021 competition on scientific literature parsing taskB: table recognition to HTML. In our method, we divide the table content recognition task into foursub-tasks: table structure recognition, text…

Computer Vision and Pattern Recognition · Computer Science 2021-05-06 Jiaquan Ye , Xianbiao Qi , Yelin He , Yihao Chen , Dengyi Gu , Peng Gao , Rong Xiao

Simultaneous machine translation (SiMT) aims to translate a continuous input text stream into another language with the lowest latency and highest quality possible. The translation thus has to start with an incomplete source text, which is…

Computation and Language · Computer Science 2020-10-14 Ozan Caglayan , Julia Ive , Veneta Haralampieva , Pranava Madhyastha , Loïc Barrault , Lucia Specia