中文
相关论文

相关论文: ICDAR 2025 Competition on End-to-End Document Imag…

200 篇论文

In-Image Machine Translation (IIMT) aims to translate texts within images from one language to another. Previous research on IIMT was primarily conducted on simplified scenarios such as images of one-line text with black font in white…

计算与语言 · 计算机科学 2025-05-22 Yanzhi Tian , Zeming Liu , Zhengyang Liu , Yuhang Guo

Motion blur is one of the most common degradation artifacts in dynamic scene photography. This paper reviews the NTIRE 2020 Challenge on Image and Video Deblurring. In this challenge, we present the evaluation results from 3 competition…

计算机视觉与模式识别 · 计算机科学 2020-05-12 Seungjun Nah , Sanghyun Son , Radu Timofte , Kyoung Mu Lee

This work presents a comparative evaluation of machine translation systems applied to images containing textual information, a task that lies at the intersection of computer vision and natural language processing. The study compares three…

计算与语言 · 计算机科学 2026-05-29 Blai Puchol , Sergio Gómez González , Miguel Domingo , Francisco Casacuberta

We introduce ABot-OCR, an end-to-end vision-language model that transcribes a page image directly into clean Markdown in a single forward pass. By doing so, our approach completely eliminates the need for brittle modular orchestration. To…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Kaitao Jiang , Ruiyan Gong , Xiaolong Cheng , Kangning Niu , Tianlun Li , Mu Xu

There has been a growing interest in developing multimodal machine translation (MMT) systems that enhance neural machine translation (NMT) with visual knowledge. This problem setup involves using images as auxiliary information during…

计算机视觉与模式识别 · 计算机科学 2023-08-30 Devaansh Gupta , Siddhant Kharbanda , Jiawei Zhou , Wanhua Li , Hanspeter Pfister , Donglai Wei

Recent studies on machine reading comprehension have focused on text-level understanding but have not yet reached the level of human understanding of the visual layout and content of real-world documents. In this study, we introduce a new…

计算与语言 · 计算机科学 2021-05-11 Ryota Tanaka , Kyosuke Nishida , Sen Yoshida

Chinese scene text reading is one of the most challenging problems in computer vision and has attracted great interest. Different from English text, Chinese has more than 6000 commonly used characters and Chinesecharacters can be arranged…

This paper reviews the NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images. This challenge received a wide range of impressive solutions, which are developed and evaluated using our collected real-world Raindrop…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Xin Li , Yeying Jin , Xin Jin , Zongwei Wu , Bingchen Li , Yufei Wang , Wenhan Yang , Yu Li , Zhibo Chen , Bihan Wen , Robby T. Tan , Radu Timofte , Qiyu Rong , Hongyuan Jing , Mengmeng Zhang , Jinglong Li , Xiangyu Lu , Yi Ren , Yuting Liu , Meng Zhang , Xiang Chen , Qiyuan Guan , Jiangxin Dong , Jinshan Pan , Conglin Gou , Qirui Yang , Fangpu Zhang , Yunlong Lin , Sixiang Chen , Guoxi Huang , Ruirui Lin , Yan Zhang , Jingyu Yang , Huanjing Yue , Jiyuan Chen , Qiaosi Yi , Hongjun Wang , Chenxi Xie , Shuai Li , Yuhui Wu , Kaiyi Ma , Jiakui Hu , Juncheng Li , Liwen Pan , Guangwei Gao , Wenjie Li , Zhenyu Jin , Heng Guo , Zhanyu Ma , Yubo Wang , Jinghua Wang , Wangzhi Xing , Anjusree Karnavar , Diqi Chen , Mohammad Aminul Islam , Hao Yang , Ruikun Zhang , Liyuan Pan , Qianhao Luo , XinCao , Han Zhou , Yan Min , Wei Dong , Jun Chen , Taoyi Wu , Weijia Dou , Yu Wang , Shengjie Zhao , Yongcheng Huang , Xingyu Han , Anyan Huang , Hongtao Wu , Hong Wang , Yefeng Zheng , Abhijeet Kumar , Aman Kumar , Marcos V. Conde , Paula Garrido , Daniel Feijoo , Juan C. Benito , Guanglu Dong , Xin Lin , Siyuan Liu , Tianheng Zheng , Jiayu Zhong , Shouyi Wang , Xiangtai Li , Lanqing Guo , Lu Qi , Chao Ren , Shuaibo Wang , Shilong Zhang , Wanyu Zhou , Yunze Wu , Qinzhong Tan , Jieyuan Pei , Zhuoxuan Li , Jiayu Wang , Haoyu Bian , Haoran Sun , Subhajit Paul , Ni Tang , Junhao Huang , Zihan Cheng , Hongyun Zhu , Yuehan Wu , Kaixin Deng , Hang Ouyang , Tianxin Xiao , Fan Yang , Zhizun Luo , Zeyu Xiao , Zhuoyuan Li , Nguyen Pham Hoang Le , An Dinh Thien , Son T. Luu , Kiet Van Nguyen , Ronghua Xu , Xianmin Tian , Weijian Zhou , Jiacheng Zhang , Yuqian Chen , Yihang Duan , Yujie Wu , Suresh Raikwar , Arsh Garg , Kritika , Jianhua Zheng , Xiaoshan Ma , Ruolin Zhao , Yongyu Yang , Yongsheng Liang , Guiming Huang , Qiang Li , Hongbin Zhang , Xiangyu Zheng , A. N. Rajagopalan

In this paper, we design and train a Generative Image-to-text Transformer, GIT, to unify vision-language tasks such as image/video captioning and question answering. While generative models provide a consistent network architecture between…

计算机视觉与模式识别 · 计算机科学 2022-12-19 Jianfeng Wang , Zhengyuan Yang , Xiaowei Hu , Linjie Li , Kevin Lin , Zhe Gan , Zicheng Liu , Ce Liu , Lijuan Wang

This competition investigates the performance of large-scale retrieval of historical document images based on writing style. Based on large image data sets provided by cultural heritage institutions and digital libraries, providing a total…

计算机视觉与模式识别 · 计算机科学 2019-12-10 Vincent Christlein , Anguelos Nicolaou , Mathias Seuret , Dominique Stutzmann , Andreas Maier

This paper presents the NTIRE 2025 image super-resolution ($\times$4) challenge, one of the associated competitions of the 10th NTIRE Workshop at CVPR 2025. The challenge aims to recover high-resolution (HR) images from low-resolution (LR)…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Zheng Chen , Kai Liu , Jue Gong , Jingkai Wang , Lei Sun , Zongwei Wu , Radu Timofte , Yulun Zhang , Xiangyu Kong , Xiaoxuan Yu , Hyunhee Park , Suejin Han , Hakjae Jeon , Dafeng Zhang , Hyung-Ju Chun , Donghun Ryou , Inju Ha , Bohyung Han , Lu Zhao , Yuyi Zhang , Pengyu Yan , Jiawei Hu , Pengwei Liu , Fengjun Guo , Hongyuan Yu , Pufan Xu , Zhijuan Huang , Shuyuan Cui , Peng Guo , Jiahui Liu , Dongkai Zhang , Heng Zhang , Huiyuan Fu , Huadong Ma , Yanhui Guo , Sisi Tian , Xin Liu , Jinwen Liang , Jie Liu , Jie Tang , Gangshan Wu , Zeyu Xiao , Zhuoyuan Li , Yinxiang Zhang , Wenxuan Cai , Vijayalaxmi Ashok Aralikatti , Nikhil Akalwadi , G Gyaneshwar Rao , Chaitra Desai , Ramesh Ashok Tabib , Uma Mudenagudi , Marcos V. Conde , Alejandro Merino , Bruno Longarela , Javier Abad , Weijun Yuan , Zhan Li , Zhanglu Chen , Boyang Yao , Aagam Jain , Milan Kumar Singh , Ankit Kumar , Shubh Kawa , Divyavardhan Singh , Anjali Sarvaiya , Kishor Upla , Raghavendra Ramachandra , Chia-Ming Lee , Yu-Fan Lin , Chih-Chung Hsu , Risheek V Hiremath , Yashaswini Palani , Yuxuan Jiang , Qiang Zhu , Siyue Teng , Fan Zhang , Shuyuan Zhu , Bing Zeng , David Bull , Jingwei Liao , Yuqing Yang , Wenda Shao , Junyi Zhao , Qisheng Xu , Kele Xu , Sunder Ali Khowaja , Ik Hyun Lee , Snehal Singh Tomar , Rajarshi Ray , Klaus Mueller , Sachin Chaudhary , Surya Vashisth , Akshay Dudhane , Praful Hambarde , Satya Naryan Tazi , Prashant Patil , Santosh Kumar Vipparthi , Subrahmanyam Murala , Bilel Benjdira , Anas M. Ali , Wadii Boulila , Zahra Moammeri , Ahmad Mahmoudi-Aznaveh , Ali Karbasi , Hossein Motamednia , Liangyan Li , Guanhua Zhao , Kevin Le , Yimo Ning , Haoxuan Huang , Jun Chen

This paper reviews the AIM 2025 Efficient Real-World Deblurring using Single Images Challenge, which aims to advance in efficient real-blur restoration. The challenge is based on a new test set based on the well known RSBlur dataset. Pairs…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Daniel Feijoo , Paula Garrido-Mellado , Marcos V. Conde , Jaesung Rim , Alvaro Garcia , Sunghyun Cho , Radu Timofte

Applying diffusion models to image-to-image translation (I2I) has recently received increasing attention due to its practical applications. Previous attempts inject information from the source image into each denoising step for an iterative…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Mengfei Xia , Yu Zhou , Ran Yi , Yong-Jin Liu , Wenping Wang

Image Difference Captioning (IDC) aims to generate natural language descriptions of subtle differences between image pairs, requiring both precise visual change localization and coherent semantic expression. Despite recent advancements,…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Yuan Liu , Saihui Hou , Saijie Hou , Jiabao Du , Shibei Meng , Yongzhen Huang

In this paper, we present the Global Multimedia Deepfake Detection held concurrently with the Inclusion 2024. Our Multimedia Deepfake Detection aims to detect automatic image and audio-video manipulations including but not limited to…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Yi Zhang , Weize Gao , Changtao Miao , Man Luo , Jianshu Li , Wenzhong Deng , Zhe Li , Bingyu Hu , Weibin Yao , Yunfeng Diao , Wenbo Zhou , Tao Gong , Qi Chu

This paper presents the computational challenge on differential geometry and topology that happened within the ICLR 2021 workshop "Geometric and Topological Representation Learning". The competition asked participants to provide creative…

This draft is a working document, having a summary of nighty-four (94) papers with additional sections on Traceability of Software Requirements (Section 4), Formal Methods and Its Tools (Section 5), Unifying Theories of Programming (UTP)…

软件工程 · 计算机科学 2025-06-24 Arshad Beg , Diarmuid O'Donoghue , Rosemary Monahan

This paper presents our solution for ICDAR 2021 competition on scientific literature parsing taskB: table recognition to HTML. In our method, we divide the table content recognition task into foursub-tasks: table structure recognition, text…

计算机视觉与模式识别 · 计算机科学 2021-05-06 Jiaquan Ye , Xianbiao Qi , Yelin He , Yihao Chen , Dengyi Gu , Peng Gao , Rong Xiao

Simultaneous machine translation (SiMT) aims to translate a continuous input text stream into another language with the lowest latency and highest quality possible. The translation thus has to start with an incomplete source text, which is…

计算与语言 · 计算机科学 2020-10-14 Ozan Caglayan , Julia Ive , Veneta Haralampieva , Pranava Madhyastha , Loïc Barrault , Lucia Specia