English
Related papers

Related papers: StarVQA: Space-Time Attention for Video Quality As…

200 papers

Deep learning-based video quality assessment (deep VQA) has demonstrated significant potential in surpassing conventional metrics, with promising improvements in terms of correlation with human perception. However, the practical deployment…

Image and Video Processing · Electrical Eng. & Systems 2025-06-03 Chen Feng , Duolikun Danier , Haoran Wang , Fan Zhang , Benoit Vallade , Alex Mackin , David Bull

Visual Question Answering (VQA) is an increasingly popular topic in deep learning research, requiring coordination of natural language processing and computer vision modules into a single architecture. We build upon the model which placed…

Computation and Language · Computer Science 2018-03-22 Jasdeep Singh , Vincent Ying , Alex Nutkiewicz

Video-based Question Answering (Video QA) is a challenging task and becomes even more intricate when addressing Socially Intelligent Question Answering (SIQA). SIQA requires context understanding, temporal reasoning, and the integration of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Aviral Agrawal , Carlos Mateo Samudio Lezcano , Iqui Balam Heredia-Marin , Prabhdeep Singh Sethi

The data scaling law has been shown to significantly enhance the performance of large multi-modal models (LMMs) across various downstream tasks. However, in the domain of perceptual video quality assessment (VQA), the potential of scaling…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Ziheng Jia , Zicheng Zhang , Zeyu Zhang , Yingji Liang , Xiaorong Zhu , Chunyi Li , Jinliang Han , Haoning Wu , Bin Wang , Haoran Zhang , Guanyu Zhu , Qiyong Zhao , Xiaohong Liu , Guangtao Zhai , Xiongkuo Min

Recently, Space-Time Memory Network (STM) based methods have achieved state-of-the-art performance in semi-supervised video object segmentation (VOS). A crucial problem in this task is how to model the dependency both among different frames…

Computer Vision and Pattern Recognition · Computer Science 2021-09-21 Jianbiao Mei , Mengmeng Wang , Yeneng Lin , Yi Yuan , Yong Liu

We present a scalable, bottom-up and intrinsically diverse data collection scheme that can be used for high-level reasoning with long and medium horizons and that has 2.2x higher throughput compared to traditional narrow top-down…

The use of complex attention modules has improved the performance of the Visual Question Answering (VQA) task. This work aims to learn an improved multi-modal representation through dense interaction of visual and textual modalities. The…

Computer Vision and Pattern Recognition · Computer Science 2023-03-01 Aakansha Mishra , Ashish Anand , Prithwijit Guha

Short-form video poses new challenges to the quality assessment of user-generated content (UGC) due to its complex generation pipeline, rapid content variation, and mixed distortions. To address this challenge, we propose an end-to-end…

Image and Video Processing · Electrical Eng. & Systems 2026-05-20 Xinyi Wang , Angeliki Katsenou , Junxiao Shen , David Bull

In this paper, we aim to obtain improved attention for a visual question answering (VQA) task. It is challenging to provide supervision for attention. An observation we make is that visual explanations as obtained through class activation…

Computer Vision and Pattern Recognition · Computer Science 2019-11-21 Badri N. Patro , Anupriy , Vinay P. Namboodiri

Deep neural networks, especially transformer-based architectures, have achieved remarkable success in semantic segmentation for environmental perception. However, existing models process video frames independently, thus failing to leverage…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Serin Varghese , Kevin Ross , Fabian Hueger , Kira Maag

Rich and dense human labeled datasets are among the main enabling factors for the recent advance on vision-language understanding. Many seemingly distant annotations (e.g., semantic segmentation and visual question answering (VQA)) are…

Computer Vision and Pattern Recognition · Computer Science 2017-08-17 Chuang Gan , Yandong Li , Haoxiang Li , Chen Sun , Boqing Gong

Robust video scene classification models should capture the spatial (pixel-wise) and temporal (frame-wise) characteristics of a video effectively. Transformer models with self-attention which are designed to get contextualized…

Computer Vision and Pattern Recognition · Computer Science 2021-10-28 Saurabh Sahu , Palash Goyal

Long-term action quality assessment (AQA) focuses on evaluating the quality of human activities in videos lasting up to several minutes. This task plays an important role in the automated evaluation of artistic sports such as rhythmic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Xin Wang , Peng-Jie Li , Yuan-Yuan Shen

With the rapid growth of in-the-wild videos taken by non-specialists, blind video quality assessment (VQA) has become a challenging and demanding problem. Although lots of efforts have been made to solve this problem, it remains unclear how…

Computer Vision and Pattern Recognition · Computer Science 2022-07-11 Liang Liao , Kangmin Xu , Haoning Wu , Chaofeng Chen , Wenxiu Sun , Qiong Yan , Weisi Lin

In light of recent progress in video editing, deep learning models focusing on both spatial and temporal dependencies have emerged as the primary method. However, these models suffer from the quadratic computational complexity of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Abdelilah Aitrouga , Youssef Hmamouche , Amal El Fallah Seghrouchni

Transformer has attracted increasing interest in STVG, owing to its end-to-end pipeline and promising result. Existing Transformer-based STVG approaches often leverage a set of object queries, which are initialized simply using zeros and…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Xin Gu , Yaojie Shen , Chenxi Luo , Tiejian Luo , Yan Huang , Yuewei Lin , Heng Fan , Libo Zhang

Despite advances, video diffusion transformers still struggle to generalize beyond their training length, a challenge we term video length extrapolation. We identify two failure modes: model-specific periodic content repetition and a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Min Zhao , Hongzhou Zhu , Yingze Wang , Bokai Yan , Jintao Zhang , Guande He , Ling Yang , Chongxuan Li , Jun Zhu

This paper strives to solve complex video question answering (VideoQA) which features long video containing multiple objects and events at different time. To tackle the challenge, we highlight the importance of identifying question-critical…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Yicong Li , Junbin Xiao , Chun Feng , Xiang Wang , Tat-Seng Chua

In the domain of video question answering (VideoQA), the impact of question types on VQA systems, despite its critical importance, has been relatively under-explored to date. However, the richness of question types directly determines the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Zhixian He , Pengcheng Zhao , Fuwei Zhang , Shujin Lin

Recently, attention-based Visual Question Answering (VQA) has achieved great success by utilizing question to selectively target different visual areas that are related to the answer. Existing visual attention models are generally planar,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-07 Jingkuan Song , Pengpeng Zeng , Lianli Gao , Heng Tao Shen