中文
相关论文

相关论文: Learned Scanpaths Aid Blind Panoramic Video Qualit…

200 篇论文

Omnidirectional image quality assessment (OIQA) aims to predict the perceptual quality of omnidirectional images that cover the whole 180$\times$360$^{\circ}$ viewing range of the visual environment. Here we propose a blind/no-reference…

多媒体 · 计算机科学 2023-02-27 Wei Zhou , Zhou Wang

Visual question answering (VQA) is the task of answering questions about an image. The task assumes an understanding of both the image and the question to provide a natural language answer. VQA has gained popularity in recent years due to…

计算机视觉与模式识别 · 计算机科学 2023-11-01 Deepanway Ghosal , Navonil Majumder , Roy Ka-Wei Lee , Rada Mihalcea , Soujanya Poria

Subjective video quality assessment (VQA) is the gold standard for measuring end-user experience across communication, streaming, and UGC pipelines. Beyond high-validity lab studies, crowdsourcing offers accurate, reliable, faster, and…

图像与视频处理 · 电气工程与系统科学 2025-09-25 Babak Naderi , Ross Cutler

Generating high-quality 360{\deg} panoramic videos remains a significant challenge due to the fundamental differences between panoramic and traditional perspective-view projections. While perspective videos rely on a single viewpoint with a…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Zeyu Dong , Yuyang Yin , Yuqi Li , Eric Li , Hao-Xiang Guo , Yikai Wang

Video coding has traditionally been developed to support services such as video streaming, videoconferencing, digital TV, and so on. The main intent was to enable human viewing of the encoded content. However, with the advances in deep…

图像与视频处理 · 电气工程与系统科学 2024-11-19 Hadi Hadizadeh , Ivan V. Bajić

Human visual perception naturally evaluates image quality across multiple scales, a hierarchical process that existing blind image quality assessment (BIQA) algorithms struggle to replicate effectively. This limitation stems from a…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Runze Hu , Zihao Huang , Xudong Li , Bohan Fu , Yan Zhang , Sicheng Zhao

When observing objects, humans benefit from their spatial visualization and mental rotation ability to envision potential optimal viewpoints based on the current observation. This capability is crucial for enabling robots to achieve…

机器人学 · 计算机科学 2025-07-29 Jiayi Wu , Xiaomin Lin , Botao He , Cornelia Fermuller , Yiannis Aloimonos

This paper offers a comprehensive analysis of recent advancements in video inpainting techniques, a critical subset of computer vision and artificial intelligence. As a process that restores or fills in missing or corrupted portions of…

计算机视觉与模式识别 · 计算机科学 2024-02-01 Shreyank N Gowda , Yash Thakre , Shashank Narayana Gowda , Xiaobo Jin

Video quality assessment (VQA) methods focus on particular degradation types, usually artificially induced on a small set of reference videos. Hence, most traditional VQA methods under-perform in-the-wild. Deep learning approaches have had…

多媒体 · 计算机科学 2021-03-02 Franz Götz-Hahn , Vlad Hosu , Hanhe Lin , Dietmar Saupe

Deep Video Quality Assessment (VQA) methods have shown impressive high-performance capabilities. Notably, no-reference (NR) VQA methods play a vital role in situations where obtaining reference videos is restricted or not feasible.…

图像与视频处理 · 电气工程与系统科学 2024-07-31 Xiaoheng Tan , Jiabin Zhang , Yuhui Quan , Jing Li , Yajing Wu , Zilin Bian

In this work, we designed a completely blind video quality assessment algorithm using the deep video prior. This work mainly explores the utility of deep video prior in estimating the visual quality of the video. In our work, we have used a…

图像与视频处理 · 电气工程与系统科学 2024-11-06 Siddharath Narayan Shakya , Parimala Kancharla

The multimodal task of Visual Question Answering (VQA) encompassing elements of Computer Vision (CV) and Natural Language Processing (NLP), aims to generate answers to questions on any visual input. Over time, the scope of VQA has expanded…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Md Farhan Ishmam , Md Sakib Hossain Shovon , M. F. Mridha , Nilanjan Dey

Panoramic video generation has attracted growing attention due to its applications in virtual reality and immersive media. However, existing methods lack explicit motion control and struggle to generate scenes with large and complex…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Cheng Zhang , Hanwen Liang , Donny Y. Chen , Qianyi Wu , Konstantinos N. Plataniotis , Camilo Cruz Gambardella , Jianfei Cai

Quality assessment of fingerprints captured using digital cameras and smartphones, also called fingerphotos, is a challenging problem in biometric recognition systems. As contactless biometric modalities are gaining more attention, their…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Amol S. Joshi , Ali Dabouei , Jeremy Dawson , Nasser Nasrabadi

Embodied Question Answering (EQA) is a recently proposed task, where an agent is placed in a rich 3D environment and must act based solely on its egocentric input to answer a given question. The desired outcome is that the agent learns to…

计算机视觉与模式识别 · 计算机科学 2019-08-15 Cătălina Cangea , Eugene Belilovsky , Pietro Liò , Aaron Courville

Face video quality assessment (FVQA) deserves to be explored in addition to general video quality assessment (VQA), as face videos are the primary content on social media platforms and human visual system (HVS) is particularly sensitive to…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Sijing Wu , Yunhao Li , Ziwen Xu , Yixuan Gao , Huiyu Duan , Wei Sun , Guangtao Zhai

The data scaling law has been shown to significantly enhance the performance of large multi-modal models (LMMs) across various downstream tasks. However, in the domain of perceptual video quality assessment (VQA), the potential of scaling…

In this paper, we propose a novel virtual reality image quality assessment (VR IQA) with adversarial learning for omnidirectional images. To take into account the characteristics of the omnidirectional image, we devise deep networks…

计算机视觉与模式识别 · 计算机科学 2018-04-12 Heoun-taek Lim , Hak Gu Kim , Yong Man Ro

Most of existing blind omnidirectional image quality assessment (BOIQA) models rely on viewport generation by modeling user viewing behavior or transforming omnidirectional images (OIs) into varying formats; however, these methods are…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Jiebin Yan , Kangcheng Wu , Junjie Chen , Ziwen Tan , Yuming Fang

Video question answering (VideoQA) is a challenging task that requires integrating spatial, temporal, and semantic information to capture the complex dynamics of video sequences. Although recent advances have introduced various approaches…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Zhongyu Yang , Zuhao Yang , Shuo Zhan , Tan Yue , Wei Pang , Yingfang Yuan