English
Related papers

Related papers: KVQ: Boosting Video Quality Assessment via Salienc…

200 papers

Current state-of-the-art video quality models, such as VMAF, give excellent prediction results by comparing the degraded video with its reference video. However, they do not consider temporal distortions (e.g., frame freezes or skips) that…

Image and Video Processing · Electrical Eng. & Systems 2023-03-23 Gabriel Mittag , Babak Naderi , Vishak Gopal , Ross Cutler

Recent advances in deep learning have markedly improved the quality of visual-attention modelling. In this work we apply these advances to video compression. We propose a compression method that uses a saliency model to adaptively compress…

Computer Vision and Pattern Recognition · Computer Science 2019-07-25 Vitaliy Lyudvichenko , Mikhail Erofeev , Alexander Ploshkin , Dmitriy Vatolin

Point cloud quality assessment (PCQA) has become an appealing research field in recent days. Considering the importance of saliency detection in quality assessment, we propose an effective full-reference PCQA metric which makes the first…

Computer Vision and Pattern Recognition · Computer Science 2022-10-03 Zhengyu Wang , Yujie Zhang , Qi Yang , Yiling Xu , Jun Sun , Shan Liu

Visual query localization (VQL) aims to predict the spatio-temporal response of the most recent occurrence in a sequence given a query. Currently, most research focuses on visual query localization in 2D videos, while its counterpart in 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Liang Peng , Bohan Tan , Zhipeng Zhang , Haobo Li , Yifan Jiao , Xingping Dong , Libo Zhang

Video quality assessment (VQA) is a challenging research topic with broad applications. Traditional hand-crafted and discriminative learning-based VQA models mainly focus on pixel-level distortions and lack contextual understanding, while…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Wen Wen , Yaohong Wu , Yue Sheng , Neil Birkbeck , Balu Adsumilli , Yilin Wang

Scale-invariance is an open problem in many computer vision subfields. For example, object labels should remain constant across scales, yet model predictions diverge in many cases. This problem gets harder for tasks where the ground-truth…

Computer Vision and Pattern Recognition · Computer Science 2022-12-13 Oliver Wiedemann , Vlad Hosu , Shaolin Su , Dietmar Saupe

Video shakiness is an unpleasant distortion of User Generated Content (UGC) videos, which is usually caused by the unstable hold of cameras. In recent years, many video stabilization algorithms have been proposed, yet no specific and…

Computer Vision and Pattern Recognition · Computer Science 2023-10-30 Tengchuan Kou , Xiaohong Liu , Wei Sun , Jun Jia , Xiongkuo Min , Guangtao Zhai , Ning Liu

The rapid growth of user-generated (video) content (UGC) has driven increased demand for research on no-reference (NR) perceptual video quality assessment (VQA). NR-VQA is a key component for large-scale video quality monitoring in social…

Image and Video Processing · Electrical Eng. & Systems 2025-08-15 Xinyi Wang , Angeliki Katsenou , David Bull

Visual Question Answering (VQA) deep-learning systems tend to capture superficial statistical correlations in the training data because of strong language priors and fail to generalize to test data with a significantly different…

Computer Vision and Pattern Recognition · Computer Science 2020-01-01 Jialin Wu , Raymond J. Mooney

We introduce CausalVQA, a benchmark dataset for video question answering (VQA) composed of question-answer pairs that probe models' understanding of causality in the physical world. Existing VQA benchmarks either tend to focus on surface…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Aaron Foss , Chloe Evans , Sasha Mitts , Koustuv Sinha , Ammar Rizvi , Justine T. Kao

Dynamic Digital Humans (DDHs) are 3D digital models that are animated using predefined motions and are inevitably bothered by noise/shift during the generation process and compression distortion during the transmission process, which needs…

Computer Vision and Pattern Recognition · Computer Science 2023-10-25 Zicheng Zhang , Yingjie Zhou , Wei Sun , Xiongkuo Min , Guangtao Zhai

The great variations of videographic skills, camera designs, compression and processing protocols, and displays lead to an enormous variety of video impairments. Current no-reference (NR) video quality models are unable to handle this…

Image and Video Processing · Electrical Eng. & Systems 2018-11-06 Zeina Sinno , Alan C. Bovik

The data scaling law has been shown to significantly enhance the performance of large multi-modal models (LMMs) across various downstream tasks. However, in the domain of perceptual video quality assessment (VQA), the potential of scaling…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Ziheng Jia , Zicheng Zhang , Zeyu Zhang , Yingji Liang , Xiaorong Zhu , Chunyi Li , Jinliang Han , Haoning Wu , Bin Wang , Haoran Zhang , Guanyu Zhu , Qiyong Zhao , Xiaohong Liu , Guangtao Zhai , Xiongkuo Min

Visual saliency, which predicts regions in the field of view that draw the most visual attention, has attracted a lot of interest from researchers. It has already been used in several vision tasks, e.g., image classification, object…

Computer Vision and Pattern Recognition · Computer Science 2015-03-25 Qiang Zhang , Yilin Wang , Baoxin Li

Quality assessment for User Generated Content (UGC) videos plays an important role in ensuring the viewing experience of end-users. Previous UGC video quality assessment (VQA) studies either use the image recognition model or the image…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Wei Sun , Xiongkuo Min , Wei Lu , Guangtao Zhai

Recent advancements in language-model-based video understanding have been progressing at a remarkable pace, spurred by the introduction of Large Language Models (LLMs). However, the focus of prior research has been predominantly on devising…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Yizhou Wang , Ruiyi Zhang , Haoliang Wang , Uttaran Bhattacharya , Yun Fu , Gang Wu

We propose a new method for the visual quality assessment of 360-degree (omnidirectional) videos. The proposed method is based on computing multiple spatio-temporal objective quality features on viewports extracted from 360-degree videos. A…

Multimedia · Computer Science 2021-05-04 Roberto G. de A. Azevedo , Neil Birkbeck , Ivan Janatra , Balu Adsumilli , Pascal Frossard

The rapid advancement of large multimodal models (LMMs) has led to the rapid expansion of artificial intelligence generated videos (AIGVs), which highlights the pressing need for effective video quality assessment (VQA) models designed…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Jiarui Wang , Huiyu Duan , Guangtao Zhai , Juntong Wang , Xiongkuo Min

We consider the problem of capturing distortions arising from changes in frame rate as part of Video Quality Assessment (VQA). Variable frame rate (VFR) videos have become much more common, and streamed videos commonly range from 30 frames…

Image and Video Processing · Electrical Eng. & Systems 2022-05-24 Pavan C. Madhusudana , Neil Birkbeck , Yilin Wang , Balu Adsumilli , Alan C. Bovik

Omnidirectional images, aka 360 images, can deliver immersive and interactive visual experiences. As their popularity has increased dramatically in recent years, evaluating the quality of 360 images has become a problem of interest since it…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Nafiseh Jabbari Tofighi , Mohamed Hedi Elfkir , Nevrez Imamoglu , Cagri Ozcinar , Erkut Erdem , Aykut Erdem