English
Related papers

Related papers: ICDAR 2021 Competition on Document VisualQuestion …

200 papers

This paper presents the final results of the ICDAR 2021 Competition on Historical Map Segmentation (MapSeg), encouraging research on a series of historical atlases of Paris, France, drawn at 1/5000 scale between 1894 and 1937. The…

Computer Vision and Pattern Recognition · Computer Science 2021-05-28 Joseph Chazalon , Edwin Carlinet , Yizi Chen , Julien Perret , Bertrand Duménieu , Clément Mallet , Thierry Géraud , Vincent Nguyen , Nam Nguyen , Josef Baloun , Ladislav Lenc , Pavel Král

This paper introduces a novel dataset for video enhancement and studies the state-of-the-art methods of the NTIRE 2021 challenge on quality enhancement of compressed video. The challenge is the first NTIRE challenge in this direction, with…

Image and Video Processing · Electrical Eng. & Systems 2021-05-04 Ren Yang , Radu Timofte

Aesthetic assessment of images can be categorized into two main forms: numerical assessment and language assessment. Aesthetics caption of photographs is the only task of aesthetic language assessment that has been addressed. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2022-08-12 Xin Jin , Wu Zhou , Xinghui Zhou , Shuai Cui , Le Zhang , Jianwen Lv , Shu Zhao

Medical Visual Question Answering~(VQA) is a combination of medical artificial intelligence and popular VQA challenges. Given a medical image and a clinically relevant question in natural language, the medical VQA system is expected to…

Computer Vision and Pattern Recognition · Computer Science 2023-06-12 Zhihong Lin , Donghao Zhang , Qingyi Tao , Danli Shi , Gholamreza Haffari , Qi Wu , Mingguang He , Zongyuan Ge

We present the 2017 WebVision Challenge, a public image recognition challenge designed for deep learning based on web images without instance-level human annotation. Following the spirit of previous vision challenges, such as ILSVRC,…

Computer Vision and Pattern Recognition · Computer Science 2017-05-17 Wen Li , Limin Wang , Wei Li , Eirikur Agustsson , Jesse Berent , Abhinav Gupta , Rahul Sukthankar , Luc Van Gool

Audio-Visual Question Answering (AVQA) is a complex multi-modal reasoning task, demanding intelligent systems to accurately respond to natural language queries based on audio-video input pairs. Nevertheless, prevalent AVQA approaches are…

Computer Vision and Pattern Recognition · Computer Science 2025-03-06 Jie Ma , Min Hu , Pinghui Wang , Wangchun Sun , Lingyun Song , Hongbin Pei , Jun Liu , Youtian Du

Current visual question answering (VQA) models tend to be trained and evaluated on image-question pairs in isolation. However, the questions people ask are dependent on their informational needs and prior knowledge about the image content.…

Computation and Language · Computer Science 2024-10-07 Nandita Shankar Naik , Christopher Potts , Elisa Kreiss

Visual Question Answering (VQA) has attracted attention from both computer vision and natural language processing communities. Most existing approaches adopt the pipeline of representing an image via pre-trained CNNs, and then using the…

Computer Vision and Pattern Recognition · Computer Science 2018-01-30 Qing Li , Jianlong Fu , Dongfei Yu , Tao Mei , Jiebo Luo

The VALUE (Video-And-Language Understanding Evaluation) benchmark is newly introduced to evaluate and analyze multi-modal representation learning algorithms on three video-and-language tasks: Retrieval, QA, and Captioning. The main…

Computer Vision and Pattern Recognition · Computer Science 2021-10-14 Minchul Shin , Jonghwan Mun , Kyoung-Woon On , Woo-Young Kang , Gunsoo Han , Eun-Sol Kim

Existing Multimodal Large Language Models (MLLMs) and Visual Language Pretrained Models (VLPMs) have shown remarkable performances in the general Visual Question Answering (VQA). However, these models struggle with VQA questions that…

Computation and Language · Computer Science 2024-11-06 Shuo Yang , Siwen Luo , Soyeon Caren Han

This report presents Team PA-VGG's solution for the ICDAR'25 Competition on Understanding Chinese College Entrance Exam Papers. In addition to leveraging high-resolution image processing and a multi-image end-to-end input strategy to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Wei Wu , Wenjie Wang , Yang Tan , Ying Liu , Liang Diao , Lin Huang , Kaihe Xu , Wenfeng Xie , Ziling Lin

In this paper, we present our solution to a Multi-modal Algorithmic Reasoning Task: SMART-101 Challenge. Different from the traditional visual question-answering datasets, this challenge evaluates the abstraction, deduction, and…

Computer Vision and Pattern Recognition · Computer Science 2023-10-11 Xiangyu Wu , Yang Yang , Shengdong Xu , Yifeng Wu , Qingguo Chen , Jianfeng Lu

Bar charts are an effective way to convey numeric information, but today's algorithms cannot parse them. Existing methods fail when faced with even minor variations in appearance. Here, we present DVQA, a dataset that tests many aspects of…

Computer Vision and Pattern Recognition · Computer Science 2018-03-30 Kushal Kafle , Brian Price , Scott Cohen , Christopher Kanan

This paper reports on the NTIRE 2024 Quality Assessment of AI-Generated Content Challenge, which will be held in conjunction with the New Trends in Image Restoration and Enhancement Workshop (NTIRE) at CVPR 2024. This challenge is to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Xiaohong Liu , Xiongkuo Min , Guangtao Zhai , Chunyi Li , Tengchuan Kou , Wei Sun , Haoning Wu , Yixuan Gao , Yuqin Cao , Zicheng Zhang , Xiele Wu , Radu Timofte , Fei Peng , Huiyuan Fu , Anlong Ming , Chuanming Wang , Huadong Ma , Shuai He , Zifei Dou , Shu Chen , Huacong Zhang , Haiyi Xie , Chengwei Wang , Baoying Chen , Jishen Zeng , Jianquan Yang , Weigang Wang , Xi Fang , Xiaoxin Lv , Jun Yan , Tianwu Zhi , Yabin Zhang , Yaohui Li , Yang Li , Jingwen Xu , Jianzhao Liu , Yiting Liao , Junlin Li , Zihao Yu , Yiting Lu , Xin Li , Hossein Motamednia , S. Farhad Hosseini-Benvidi , Fengbin Guan , Ahmad Mahmoudi-Aznaveh , Azadeh Mansouri , Ganzorig Gankhuyag , Kihwan Yoon , Yifang Xu , Haotian Fan , Fangyuan Kong , Shiling Zhao , Weifeng Dong , Haibing Yin , Li Zhu , Zhiling Wang , Bingchen Huang , Avinab Saha , Sandeep Mishra , Shashank Gupta , Rajesh Sureddi , Oindrila Saha , Luigi Celona , Simone Bianco , Paolo Napoletano , Raimondo Schettini , Junfeng Yang , Jing Fu , Wei Zhang , Wenzhi Cao , Limei Liu , Han Peng , Weijun Yuan , Zhan Li , Yihang Cheng , Yifan Deng , Haohui Li , Bowen Qu , Yao Li , Shuqing Luo , Shunzhou Wang , Wei Gao , Zihao Lu , Marcos V. Conde , Xinrui Wang , Zhibo Chen , Ruling Liao , Yan Ye , Qiulin Wang , Bing Li , Zhaokun Zhou , Miao Geng , Rui Chen , Xin Tao , Xiaoyu Liang , Shangkun Sun , Xingyuan Ma , Jiaze Li , Mengduo Yang , Haoran Xu , Jie Zhou , Shiding Zhu , Bohan Yu , Pengfei Chen , Xinrui Xu , Jiabin Shen , Zhichao Duan , Erfan Asadi , Jiahe Liu , Qi Yan , Youran Qu , Xiaohui Zeng , Lele Wang , Renjie Liao

Information fusion is used widely to improve document classification by the integration of multiple data sources (multimodal) or representations (multiview). However, the field lacks a unified framework, a quantitative synthesis of its…

Computation and Language · Computer Science 2026-05-27 Marcin Michał Mirończuk

Modern VLMs have achieved near-saturation accuracy in English document visual question-answering (VQA). However, this task remains challenging in lower resource languages due to a dearth of suitable training and evaluation data. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Jonathan Li , Zoltan Csaki , Nidhi Hiremath , Etash Guha , Fenglu Hong , Edward Ma , Urmish Thakker

The advent and proliferation of large multi-modal models (LMMs) have introduced new paradigms to computer vision, transforming various tasks into a unified visual question answering framework. Video Quality Assessment (VQA), a classic field…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Ziheng Jia , Zicheng Zhang , Jiaying Qian , Haoning Wu , Wei Sun , Chunyi Li , Xiaohong Liu , Weisi Lin , Guangtao Zhai , Xiongkuo Min

Visual Question Answering (VQA) task has showcased a new stage of interaction between language and vision, two of the most pivotal components of artificial intelligence. However, it has mostly focused on generating short and repetitive…

Computer Vision and Pattern Recognition · Computer Science 2016-09-22 Andrew Shin , Yoshitaka Ushiku , Tatsuya Harada

This paper reviews the Challenge on Super-Resolution of Compressed Image and Video at AIM 2022. This challenge includes two tracks. Track 1 aims at the super-resolution of compressed image, and Track~2 targets the super-resolution of…

Visual document understanding is a complex task that involves analyzing both the text and the visual elements in document images. Existing models often rely on manual feature engineering or domain-specific pipelines, which limit their…