中文
相关论文

相关论文: ICDAR 2025 Competition on End-to-End Document Imag…

200 篇论文

We introduce DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages. Using a strictly controlled setup, 18 professional translators were instructed to…

计算与语言 · 计算机科学 2023-09-08 Gabriele Sarti , Arianna Bisazza , Ana Guerberof Arenas , Antonio Toral

This paper reports on the NTIRE 2022 challenge on perceptual image quality assessment (IQA), held in conjunction with the New Trends in Image Restoration and Enhancement workshop (NTIRE) workshop at CVPR 2022. This challenge is held to…

计算机视觉与模式识别 · 计算机科学 2022-06-24 Jinjin Gu , Haoming Cai , Chao Dong , Jimmy S. Ren , Radu Timofte

Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capability in document literacy. However, their effectiveness in real-world applications…

Chinese is the most widely used language in the world. Algorithms that read Chinese text in natural images facilitate applications of various kinds. Despite the large potential value, datasets and competitions in the past primarily focus on…

计算机视觉与模式识别 · 计算机科学 2018-09-27 Baoguang Shi , Cong Yao , Minghui Liao , Mingkun Yang , Pei Xu , Linyan Cui , Serge Belongie , Shijian Lu , Xiang Bai

This paper reports on the NTIRE 2025 challenge on HR Depth From images of Specular and Transparent surfaces, held in conjunction with the New Trends in Image Restoration and Enhancement (NTIRE) workshop at CVPR 2025. This challenge aims to…

This paper reviews the NTIRE 2022 challenge on efficient single image super-resolution with focus on the proposed solutions and results. The task of the challenge was to super-resolve an input image with a magnification factor of $\times$4…

计算机视觉与模式识别 · 计算机科学 2022-05-12 Yawei Li , Kai Zhang , Radu Timofte , Luc Van Gool , Fangyuan Kong , Mingxi Li , Songwei Liu , Zongcai Du , Ding Liu , Chenhui Zhou , Jingyi Chen , Qingrui Han , Zheyuan Li , Yingqi Liu , Xiangyu Chen , Haoming Cai , Yu Qiao , Chao Dong , Long Sun , Jinshan Pan , Yi Zhu , Zhikai Zong , Xiaoxiao Liu , Zheng Hui , Tao Yang , Peiran Ren , Xuansong Xie , Xian-Sheng Hua , Yanbo Wang , Xiaozhong Ji , Chuming Lin , Donghao Luo , Ying Tai , Chengjie Wang , Zhizhong Zhang , Yuan Xie , Shen Cheng , Ziwei Luo , Lei Yu , Zhihong Wen , Qi Wu1 , Youwei Li , Haoqiang Fan , Jian Sun , Shuaicheng Liu , Yuanfei Huang , Meiguang Jin , Hua Huang , Jing Liu , Xinjian Zhang , Yan Wang , Lingshun Long , Gen Li , Yuanfan Zhang , Zuowei Cao , Lei Sun , Panaetov Alexander , Yucong Wang , Minjie Cai , Li Wang , Lu Tian , Zheyuan Wang , Hongbing Ma , Jie Liu , Chao Chen , Yidong Cai , Jie Tang , Gangshan Wu , Weiran Wang , Shirui Huang , Honglei Lu , Huan Liu , Keyan Wang , Jun Chen , Shi Chen , Yuchun Miao , Zimo Huang , Lefei Zhang , Mustafa Ayazoğlu , Wei Xiong , Chengyi Xiong , Fei Wang , Hao Li , Ruimian Wen , Zhijing Yang , Wenbin Zou , Weixin Zheng , Tian Ye , Yuncheng Zhang , Xiangzhen Kong , Aditya Arora , Syed Waqas Zamir , Salman Khan , Munawar Hayat , Fahad Shahbaz Khan , Dandan Gaoand Dengwen Zhouand Qian Ning , Jingzhu Tang , Han Huang , Yufei Wang , Zhangheng Peng , Haobo Li , Wenxue Guan , Shenghua Gong , Xin Li , Jun Liu , Wanjun Wang , Dengwen Zhou , Kun Zeng , Hanjiang Lin , Xinyu Chen , Jinsheng Fang

This paper reviews the first-ever image demoireing challenge that was part of the Advances in Image Manipulation (AIM) workshop, held in conjunction with ICCV 2019. This paper describes the challenge, and focuses on the proposed solutions…

This paper presents final results of the Out-Of-Vocabulary 2022 (OOV) challenge. The OOV contest introduces an important aspect that is not commonly studied by Optical Character Recognition (OCR) models, namely, the recognition of unseen…

计算机视觉与模式识别 · 计算机科学 2022-09-15 Sergi Garcia-Bordils , Andrés Mafla , Ali Furkan Biten , Oren Nuriel , Aviad Aberdam , Shai Mazor , Ron Litman , Dimosthenis Karatzas

This paper presents a comprehensive review of the NTIRE 2026 Low Light Image Enhancement Challenge, highlighting the proposed solutions and final results. The objective of this challenge is to identify effective networks capable of…

In this paper, we present a comprehensive overview of the NTIRE 2025 challenge on the 2nd Restore Any Image Model (RAIM) in the Wild. This challenge established a new benchmark for real-world image restoration, featuring diverse scenarios…

图像与视频处理 · 电气工程与系统科学 2025-06-03 Jie Liang , Radu Timofte , Qiaosi Yi , Zhengqiang Zhang , Shuaizheng Liu , Lingchen Sun , Rongyuan Wu , Xindong Zhang , Hui Zeng , Lei Zhang

This report presents Team PA-VGG's solution for the ICDAR'25 Competition on Understanding Chinese College Entrance Exam Papers. In addition to leveraging high-resolution image processing and a multi-image end-to-end input strategy to…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Wei Wu , Wenjie Wang , Yang Tan , Ying Liu , Liang Diao , Lin Huang , Kaihe Xu , Wenfeng Xie , Ziling Lin

We introduce the AIM 2025 Real-World RAW Image Denoising Challenge, aiming to advance efficient and effective denoising techniques grounded in data synthesis. The competition is built upon a newly established evaluation benchmark featuring…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Feiran Li , Jiacheng Li , Marcos V. Conde , Beril Besbinar , Vlad Hosu , Daisuke Iso , Radu Timofte

Text Image Machine Translation (TIMT) aims to translate text embedded in images in the source-language into target-language, requiring synergistic integration of visual perception and linguistic understanding. Existing TIMT methods, whether…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Junxin Lu , Tengfei Song , Zhanglin Wu , Pengfei Li , Xiaowei Liang , Hui Yang , Kun Chen , Ning Xie , Yunfei Lu , Jing Zhao , Shiliang Sun , Daimeng Wei

The rise of multi-modal search requests from users has highlighted the importance of multi-modal retrieval (i.e. image-to-text or text-to-image retrieval), yet the more complex task of image-to-multi-modal retrieval, crucial for many…

信息检索 · 计算机科学 2024-06-11 Zida Cheng , Chen Ju , Shuai Xiao , Xu Chen , Zhonghua Zhai , Xiaoyi Zeng , Weilin Huang , Junchi Yan

We present the results from the second shared task on multimodal machine translation and multilingual image description. Nine teams submitted 19 systems to two tasks. The multimodal translation task, in which the source sentence is…

计算与语言 · 计算机科学 2017-10-20 Desmond Elliott , Stella Frank , Loïc Barrault , Fethi Bougares , Lucia Specia

Text-rich document understanding (TDU) requires comprehensive analysis of documents containing substantial textual content and complex layouts. While Multimodal Large Language Models (MLLMs) have achieved fast progress in this domain,…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Wenhui Liao , Jiapeng Wang , Hongliang Li , Chengyu Wang , Jun Huang , Lianwen Jin

We present Multimodal OCR (MOCR), a document parsing paradigm that jointly parses text and graphics into unified textual representations. Unlike conventional OCR systems that focus on text recognition and leave graphical regions as cropped…

Omnidirectional images (ODIs) provide full 360x180 view which are widely adopted in VR, AR and embodied intelligence applications. While multi-modal large language models (MLLMs) have demonstrated remarkable performance on conventional 2D…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Liu Yang , Huiyu Duan , Ran Tao , Juntao Cheng , Sijing Wu , Yunhao Li , Jing Liu , Xiongkuo Min , Guangtao Zhai

Recognition of identity documents using mobile devices has become a topic of a wide range of computer vision research. The portfolio of methods and algorithms for solving such tasks as face detection, document detection and rectification,…

计算机视觉与模式识别 · 计算机科学 2020-02-12 Konstantin Bulatov , Daniil Matalov , Vladimir V. Arlazarov