English
Related papers

Related papers: ICDAR 2023 Competition on Structured Text Extracti…

200 papers

Since real-world ubiquitous documents (e.g., invoices, tickets, resumes and leaflets) contain rich information, automatic document image understanding has become a hot topic. Most existing works decouple the problem into two separate tasks,…

Computer Vision and Pattern Recognition · Computer Science 2021-10-26 Peng Zhang , Yunlu Xu , Zhanzhan Cheng , Shiliang Pu , Jing Lu , Liang Qiao , Yi Niu , Fei Wu

This paper reports on the NTIRE 2024 Quality Assessment of AI-Generated Content Challenge, which will be held in conjunction with the New Trends in Image Restoration and Enhancement Workshop (NTIRE) at CVPR 2024. This challenge is to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Xiaohong Liu , Xiongkuo Min , Guangtao Zhai , Chunyi Li , Tengchuan Kou , Wei Sun , Haoning Wu , Yixuan Gao , Yuqin Cao , Zicheng Zhang , Xiele Wu , Radu Timofte , Fei Peng , Huiyuan Fu , Anlong Ming , Chuanming Wang , Huadong Ma , Shuai He , Zifei Dou , Shu Chen , Huacong Zhang , Haiyi Xie , Chengwei Wang , Baoying Chen , Jishen Zeng , Jianquan Yang , Weigang Wang , Xi Fang , Xiaoxin Lv , Jun Yan , Tianwu Zhi , Yabin Zhang , Yaohui Li , Yang Li , Jingwen Xu , Jianzhao Liu , Yiting Liao , Junlin Li , Zihao Yu , Yiting Lu , Xin Li , Hossein Motamednia , S. Farhad Hosseini-Benvidi , Fengbin Guan , Ahmad Mahmoudi-Aznaveh , Azadeh Mansouri , Ganzorig Gankhuyag , Kihwan Yoon , Yifang Xu , Haotian Fan , Fangyuan Kong , Shiling Zhao , Weifeng Dong , Haibing Yin , Li Zhu , Zhiling Wang , Bingchen Huang , Avinab Saha , Sandeep Mishra , Shashank Gupta , Rajesh Sureddi , Oindrila Saha , Luigi Celona , Simone Bianco , Paolo Napoletano , Raimondo Schettini , Junfeng Yang , Jing Fu , Wei Zhang , Wenzhi Cao , Limei Liu , Han Peng , Weijun Yuan , Zhan Li , Yihang Cheng , Yifan Deng , Haohui Li , Bowen Qu , Yao Li , Shuqing Luo , Shunzhou Wang , Wei Gao , Zihao Lu , Marcos V. Conde , Xinrui Wang , Zhibo Chen , Ruling Liao , Yan Ye , Qiulin Wang , Bing Li , Zhaokun Zhou , Miao Geng , Rui Chen , Xin Tao , Xiaoyu Liang , Shangkun Sun , Xingyuan Ma , Jiaze Li , Mengduo Yang , Haoran Xu , Jie Zhou , Shiding Zhu , Bohan Yu , Pengfei Chen , Xinrui Xu , Jiabin Shen , Zhichao Duan , Erfan Asadi , Jiahe Liu , Qi Yan , Youran Qu , Xiaohui Zeng , Lele Wang , Renjie Liao

Visually Rich Documents (VRDs) play a vital role in domains such as academia, finance, healthcare, and marketing, as they convey information through a combination of text, layout, and visual elements. Traditional approaches to extracting…

Computation and Language · Computer Science 2025-06-23 Yihao Ding , Soyeon Caren Han , Jean Lee , Eduard Hovy

Humans watch more than a billion hours of video per day. Most of this video was edited manually, which is a tedious process. However, AI-enabled video-generation and video-editing is on the rise. Building on text-to-image models like Stable…

This paper presents the NTIRE 2025 image super-resolution ($\times$4) challenge, one of the associated competitions of the 10th NTIRE Workshop at CVPR 2025. The challenge aims to recover high-resolution (HR) images from low-resolution (LR)…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Zheng Chen , Kai Liu , Jue Gong , Jingkai Wang , Lei Sun , Zongwei Wu , Radu Timofte , Yulun Zhang , Xiangyu Kong , Xiaoxuan Yu , Hyunhee Park , Suejin Han , Hakjae Jeon , Dafeng Zhang , Hyung-Ju Chun , Donghun Ryou , Inju Ha , Bohyung Han , Lu Zhao , Yuyi Zhang , Pengyu Yan , Jiawei Hu , Pengwei Liu , Fengjun Guo , Hongyuan Yu , Pufan Xu , Zhijuan Huang , Shuyuan Cui , Peng Guo , Jiahui Liu , Dongkai Zhang , Heng Zhang , Huiyuan Fu , Huadong Ma , Yanhui Guo , Sisi Tian , Xin Liu , Jinwen Liang , Jie Liu , Jie Tang , Gangshan Wu , Zeyu Xiao , Zhuoyuan Li , Yinxiang Zhang , Wenxuan Cai , Vijayalaxmi Ashok Aralikatti , Nikhil Akalwadi , G Gyaneshwar Rao , Chaitra Desai , Ramesh Ashok Tabib , Uma Mudenagudi , Marcos V. Conde , Alejandro Merino , Bruno Longarela , Javier Abad , Weijun Yuan , Zhan Li , Zhanglu Chen , Boyang Yao , Aagam Jain , Milan Kumar Singh , Ankit Kumar , Shubh Kawa , Divyavardhan Singh , Anjali Sarvaiya , Kishor Upla , Raghavendra Ramachandra , Chia-Ming Lee , Yu-Fan Lin , Chih-Chung Hsu , Risheek V Hiremath , Yashaswini Palani , Yuxuan Jiang , Qiang Zhu , Siyue Teng , Fan Zhang , Shuyuan Zhu , Bing Zeng , David Bull , Jingwei Liao , Yuqing Yang , Wenda Shao , Junyi Zhao , Qisheng Xu , Kele Xu , Sunder Ali Khowaja , Ik Hyun Lee , Snehal Singh Tomar , Rajarshi Ray , Klaus Mueller , Sachin Chaudhary , Surya Vashisth , Akshay Dudhane , Praful Hambarde , Satya Naryan Tazi , Prashant Patil , Santosh Kumar Vipparthi , Subrahmanyam Murala , Bilel Benjdira , Anas M. Ali , Wadii Boulila , Zahra Moammeri , Ahmad Mahmoudi-Aznaveh , Ali Karbasi , Hossein Motamednia , Liangyan Li , Guanhua Zhao , Kevin Le , Yimo Ning , Haoxuan Huang , Jun Chen

Recent deep learning models have demonstrated strong capabilities for classifying text and non-text components in natural images. They extract a high-level feature computed globally from a whole image component (patch), where the cluttered…

Computer Vision and Pattern Recognition · Computer Science 2016-05-04 Tong He , Weilin Huang , Yu Qiao , Jian Yao

Super-Resolution (SR) is a fundamental computer vision task that aims to obtain a high-resolution clean image from the given low-resolution counterpart. This paper reviews the NTIRE 2021 Challenge on Video Super-Resolution. We present…

Computer Vision and Pattern Recognition · Computer Science 2021-05-11 Sanghyun Son , Suyoung Lee , Seungjun Nah , Radu Timofte , Kyoung Mu Lee

Detecting scene text of arbitrary shapes has been a challenging task over the past years. In this paper, we propose a novel segmentation-based text detector, namely SAST, which employs a context attended multi-task learning framework based…

Computer Vision and Pattern Recognition · Computer Science 2019-08-16 Pengfei Wang , Chengquan Zhang , Fei Qi , Zuming Huang , Mengyi En , Junyu Han , Jingtuo Liu , Errui Ding , Guangming Shi

Reading text from natural images is challenging due to the great variety in text font, color, size, complex background and etc.. The perspective distortion and non-linear spatial arrangement of characters make it further difficult. While…

Computer Vision and Pattern Recognition · Computer Science 2019-11-12 Shangbang Long , Yushuo Guan , Bingxuan Wang , Kaigui Bian , Cong Yao

3D object retrieval is an important yet challenging task that has drawn more and more attention in recent years. While existing approaches have made strides in addressing this issue, they are often limited to restricted settings such as…

Document-level Relation Extraction (DocRE) involves identifying relations between entities across multiple sentences in a document. Evidence sentences, crucial for precise entity pair relationships identification, enhance focus on essential…

Computation and Language · Computer Science 2025-04-10 Khai Phan Tran , Xue Li

We introduce the structured scene-text spotting task, which requires a scene-text OCR system to spot text in the wild according to a query regular expression. Contrary to generic scene text OCR, structured scene-text spotting seeks to…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Sergi Garcia-Bordils , Dimosthenis Karatzas , Marçal Rusiñol

The goal of visual word sense disambiguation is to find the image that best matches the provided description of the word's meaning. It is a challenging problem, requiring approaches that combine language and image understanding. In this…

Computation and Language · Computer Science 2023-04-17 Sławomir Dadas

There has been a significant effort by the research community to address the problem of providing methods to organize documentation with the help of information Retrieval methods. In this report paper, we present several experiments with…

Information Retrieval · Computer Science 2022-06-07 Rui Portocarrero Sarmento , Douglas O. Cardoso , João Gama , Pavel Brazdil

We call on the Document AI (DocAI) community to reevaluate current methodologies and embrace the challenge of creating more practically-oriented benchmarks. Document Understanding Dataset and Evaluation (DUDE) seeks to remediate the halted…

This competition focus on Urban-Sense Segmentation based on the vehicle camera view. Class highly unbalanced Urban-Sense images dataset challenge the existing solutions and further studies. Deep Conventional neural network-based semantic…

Computer Vision and Pattern Recognition · Computer Science 2022-07-05 Akide Liu , Zihan Wang

In this paper, we present StrucTexTv2, an effective document image pre-training framework, by performing masked visual-textual prediction. It consists of two self-supervised pre-training tasks: masked image modeling and masked language…

Computer Vision and Pattern Recognition · Computer Science 2023-03-02 Yuechen Yu , Yulin Li , Chengquan Zhang , Xiaoqiang Zhang , Zengyuan Guo , Xiameng Qin , Kun Yao , Junyu Han , Errui Ding , Jingdong Wang

The field of visually rich document understanding (VRDU) aims to solve a multitude of well-researched NLP tasks in a multi-modal domain. Several datasets exist for research on specific tasks of VRDU such as document classification (DC), key…

The use of visually-rich documents (VRDs) in various fields has created a demand for Document AI models that can read and comprehend documents like humans, which requires the overcoming of technical, linguistic, and cognitive barriers.…

Human-Computer Interaction · Computer Science 2023-10-24 Hao Wang , Qingxuan Wang , Yue Li , Changqing Wang , Chenhui Chu , Rui Wang

Image-to-text tasks, such as open-ended image captioning and controllable image description, have received extensive attention for decades. Here, we further advance this line of work by presenting Visual Spatial Description (VSD), a new…

Computer Vision and Pattern Recognition · Computer Science 2022-10-27 Yu Zhao , Jianguo Wei , Zhichao Lin , Yueheng Sun , Meishan Zhang , Min Zhang
‹ Prev 1 3 4 5 6 7 10 Next ›