中文
相关论文

相关论文: ICDAR 2025 Competition on End-to-End Document Imag…

200 篇论文

Image relighting is attracting increasing interest due to its various applications. From a research perspective, image relighting can be exploited to conduct both image normalization for domain adaptation, and also for data augmentation. It…

计算机视觉与模式识别 · 计算机科学 2021-04-28 Majed El Helou , Ruofan Zhou , Sabine Susstrunk , Radu Timofte

The performance of optical character recognition (OCR) heavily relies on document image quality, which is crucial for automatic document processing and document intelligence. However, most existing document enhancement methods require…

计算机视觉与模式识别 · 计算机科学 2023-11-17 Jiaxin Zhang , Joy Rimchala , Lalla Mouatadid , Kamalika Das , Sricharan Kumar

In this report we present results of the ICDAR 2021 edition of the Document Visual Question Challenges. This edition complements the previous tasks on Single Document VQA and Document Collection VQA with a newly introduced on Infographics…

计算机视觉与模式识别 · 计算机科学 2021-11-11 Rubèn Tito , Minesh Mathew , C. V. Jawahar , Ernest Valveny , Dimosthenis Karatzas

Image Translation (IT) holds immense potential across diverse domains, enabling the translation of textual content within images into various languages. However, existing datasets often suffer from limitations in scale, diversity, and…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Bo Li , Shaolin Zhu , Lijie Wen

Scene video text spotting (SVTS) is a very important research topic because of many real-life applications. However, only a little effort has put to spotting scene video text, in contrast to massive studies of scene text spotting in static…

计算机视觉与模式识别 · 计算机科学 2021-07-27 Zhanzhan Cheng , Jing Lu , Baorui Zou , Shuigeng Zhou , Fei Wu

Document Layout Parsing serves as a critical gateway for Artificial Intelligence (AI) to access and interpret the world's vast stores of structured knowledge. This process,which encompasses layout detection, text recognition, and relational…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Yumeng Li , Guang Yang , Hao Liu , Bowen Wang , Colin Zhang

End-to-end text image translation (TIT), which aims at translating the source language embedded in images to the target language, has attracted intensive attention in recent research. However, data sparsity limits the performance of…

计算与语言 · 计算机科学 2022-10-11 Cong Ma , Yaping Zhang , Mei Tu , Xu Han , Linghui Wu , Yang Zhao , Yu Zhou

This paper presents results of Document Visual Question Answering Challenge organized as part of "Text and Documents in the Deep Learning Era" workshop, in CVPR 2020. The challenge introduces a new problem - Visual Question Answering on…

计算机视觉与模式识别 · 计算机科学 2021-07-20 Minesh Mathew , Ruben Tito , Dimosthenis Karatzas , R. Manmatha , C. V. Jawahar

Motion blur is a common photography artifact in dynamic environments that typically comes jointly with the other types of degradation. This paper reviews the NTIRE 2021 Challenge on Image Deblurring. In this challenge report, we describe…

计算机视觉与模式识别 · 计算机科学 2021-05-03 Seungjun Nah , Sanghyun Son , Suyoung Lee , Radu Timofte , Kyoung Mu Lee

Real-world infrared imagery presents unique challenges for vision-language models due to the scarcity of aligned text data and domain-specific characteristics. Although existing methods have advanced the field, their reliance on synthetic…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Zhe Cao , Jin Zhang , Ruiheng Zhang

This paper presents a comprehensive review of the NTIRE 2025 Low-Light Image Enhancement (LLIE) Challenge, highlighting the proposed solutions and final outcomes. The objective of the challenge is to identify effective networks capable of…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Xiaoning Liu , Zongwei Wu , Florin-Alexandru Vasluianu , Hailong Yan , Bin Ren , Yulun Zhang , Shuhang Gu , Le Zhang , Ce Zhu , Radu Timofte , Kangbiao Shi , Yixu Feng , Tao Hu , Yu Cao , Peng Wu , Yijin Liang , Yanning Zhang , Qingsen Yan , Han Zhou , Wei Dong , Yan Min , Mohab Kishawy , Jun Chen , Pengpeng Yu , Anjin Park , Seung-Soo Lee , Young-Joon Park , Zixiao Hu , Junyv Liu , Huilin Zhang , Jun Zhang , Fei Wan , Bingxin Xu , Hongzhe Liu , Cheng Xu , Weiguo Pan , Songyin Dai , Xunpeng Yi , Qinglong Yan , Yibing Zhang , Jiayi Ma , Changhui Hu , Kerui Hu , Donghang Jing , Tiesheng Chen , Zhi Jin , Hongjun Wu , Biao Huang , Haitao Ling , Jiahao Wu , Dandan Zhan , G Gyaneshwar Rao , Vijayalaxmi Ashok Aralikatti , Nikhil Akalwadi , Ramesh Ashok Tabib , Uma Mudenagudi , Ruirui Lin , Guoxi Huang , Nantheera Anantrasirichai , Qirui Yang , Alexandru Brateanu , Ciprian Orhei , Cosmin Ancuti , Daniel Feijoo , Juan C. Benito , Álvaro García , Marcos V. Conde , Yang Qin , Raul Balmez , Anas M. Ali , Bilel Benjdira , Wadii Boulila , Tianyi Mao , Huan Zheng , Yanyan Wei , Shengeng Tang , Dan Guo , Zhao Zhang , Sabari Nathan , K Uma , A Sasithradevi , B Sathya Bama , S. Mohamed Mansoor Roomi , Ao Li , Xiangtao Zhang , Zhe Liu , Yijie Tang , Jialong Tang , Zhicheng Fu , Gong Chen , Joe Nasti , John Nicholson , Zeyu Xiao , Zhuoyuan Li , Ashutosh Kulkarni , Prashant W. Patil , Santosh Kumar Vipparthi , Subrahmanyam Murala , Duan Liu , Weile Li , Hangyuan Lu , Rixian Liu , Tengfeng Wang , Jinxing Liang , Chenxin Yu

Text image super-resolution is a challenging yet open research problem in the computer vision community. In particular, low-resolution images hamper the performance of typical optical character recognition (OCR) systems. In this article, we…

计算机视觉与模式识别 · 计算机科学 2015-06-09 Chao Dong , Ximei Zhu , Yubin Deng , Chen Change Loy , Yu Qiao

Image Transformer has recently achieved significant progress for natural image understanding, either using supervised (ViT, DeiT, etc.) or self-supervised (BEiT, MAE, etc.) pre-training techniques. In this paper, we propose \textbf{DiT}, a…

计算机视觉与模式识别 · 计算机科学 2022-07-20 Junlong Li , Yiheng Xu , Tengchao Lv , Lei Cui , Cha Zhang , Furu Wei

This paper presents Centre for Development of Advanced Computing Mumbai's (CDACM) submission to NLP Tools Contest on Statistical Machine Translation in Indian Languages (ILSMT) 2015 (collocated with ICON 2015). The aim of the contest was to…

计算与语言 · 计算机科学 2016-10-26 Raj Nath Patel , Prakash B. Pimpale

Large language models (LLMs) have significantly advanced various natural language processing (NLP) tasks. Recent research indicates that moderately-sized LLMs often outperform larger ones after task-specific fine-tuning. This study focuses…

计算与语言 · 计算机科学 2024-10-14 Minghao Wu , Thuy-Trang Vu , Lizhen Qu , George Foster , Gholamreza Haffari

For many computer vision applications such as image captioning, visual question answering, and person search, learning discriminative feature representations at both image and text level is an essential yet challenging problem. Its…

计算机视觉与模式识别 · 计算机科学 2019-08-29 Nikolaos Sarafianos , Xiang Xu , Ioannis A. Kakadiaris

Image-to-image translation aims to learn a mapping between a source and a target domain, enabling tasks such as style transfer, appearance transformation, and domain adaptation. In this work, we explore a diffusion-based framework for…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Qiang Zhu , Kuan Lu , Menghao Huo , Yuxiao Li

Vision-Language Translation (VLT) is a challenging task that requires accurately recognizing multilingual text embedded in images and translating it into the target language with the support of visual context. While recent Large…

计算机视觉与模式识别 · 计算机科学 2025-06-16 Xintong Wang , Jingheng Pan , Yixiao Liu , Xiaohu Zhao , Chenyang Lyu , Minghao Wu , Chris Biemann , Longyue Wang , Linlong Xu , Weihua Luo , Kaifu Zhang

Developing and integrating advanced image sensors with novel algorithms in camera systems are prevalent with the increasing demand for computational photography and imaging on mobile platforms. However, the lack of high-quality data for…

图像与视频处理 · 电气工程与系统科学 2022-10-25 Ruicheng Feng , Chongyi Li , Shangchen Zhou , Wenxiu Sun , Qingpeng Zhu , Jun Jiang , Qingyu Yang , Chen Change Loy , Jinwei Gu