中文
相关论文

相关论文: VDPVE: VQA Dataset for Perceptual Video Enhancemen…

200 篇论文

Video description is the automatic generation of natural language sentences that describe the contents of a given video. It has applications in human-robot interaction, helping the visually impaired and video subtitling. The past few years…

计算机视觉与模式识别 · 计算机科学 2020-03-04 Nayyer Aafaq , Ajmal Mian , Wei Liu , Syed Zulqarnain Gilani , Mubarak Shah

Learning-based video quality assessment (VQA) has advanced rapidly, yet progress is increasingly constrained by a disconnect between model design and dataset curation. Model-centric approaches often iterate on fixed benchmarks, while…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Jian Zou , Xiaoyu Xu , Zhihua Wang , Yilin Wang , Balu Adsumilli , Kede Ma

Owing to the proliferation of user-generated videos on the Internet, blind video quality assessment (BVQA) at the edge attracts growing attention. The usage of deep-learning-based methods is restricted to be applied at the edge due to their…

图像与视频处理 · 电气工程与系统科学 2023-10-31 Zhanxuan Mei , Yun-Cheng Wang , C. -C. Jay Kuo

Video conferencing, which includes both video and audio content, has contributed to dramatic increases in Internet traffic, as the COVID-19 pandemic forced millions of people to work and learn from home. Global Internet traffic of video…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Zhenqiang Ying , Deepti Ghadiyaram , Alan Bovik

Introduction: Video Quality Assessment (VQA) is one of the important areas of study in this modern era, where video is a crucial component of communication with applications in every field. Rapid technology developments in mobile technology…

图像与视频处理 · 电气工程与系统科学 2024-04-10 Anantha Prabhu , David Pratap , Narayana Darapeni , Anwesh P R

We introduce a new large-scale dataset that links the assessment of image quality issues to two practical vision tasks: image captioning and visual question answering. First, we identify for 39,181 images taken by people who are blind…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Tai-Yin Chiu , Yinan Zhao , Danna Gurari

The ever growing realism and quality of generated videos makes it increasingly harder for humans to spot deepfake content, who need to rely more and more on automatic deepfake detectors. However, deepfake detectors are also prone to errors,…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Vlad Hondru , Eduard Hogea , Darian Onchis , Radu Tudor Ionescu

This paper analyzes the joint assessment of quality, spatial and social presence, empathy, attitude, and attention in three conditions: (A)visualizing and rating the quality of contents in a Head-Mounted Display (HMD), (B)visualizing the…

多媒体 · 计算机科学 2022-02-10 Marta Orduna , Pablo Pérez , Jesús Gutiérrez , Narciso García

Video generation based on diffusion models presents a challenging multimodal task, with video editing emerging as a pivotal direction in this field. Recent video editing approaches primarily fall into two categories: training-required and…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Junhao Xia , Chaoyang Zhang , Yecheng Zhang , Chengyang Zhou , Zhichang Wang , Bochun Liu , Dongshuo Yin

Video captions play a crucial role in text-to-video generation tasks, as their quality directly influences the semantic coherence and visual fidelity of the generated videos. Although large vision-language models (VLMs) have demonstrated…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Shi-Xue Zhang , Hongfa Wang , Duojun Huang , Xin Li , Xiaobin Zhu , Xu-Cheng Yin

The commercialization of Virtual Reality (VR) headsets has made immersive and 360-degree video streaming the subject of intense interest in the industry and research communities. While the basic principles of video streaming are the same,…

多媒体 · 计算机科学 2021-02-17 Federico Chiariotti

Despite significant breakthroughs in video analysis driven by the rapid development of large multimodal models (LMMs), there remains a lack of a versatile evaluation benchmark to comprehensively assess these models' performance in video…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Yunxin Li , Xinyu Chen , Baotian Hu , Longyue Wang , Haoyuan Shi , Min Zhang

Video inpainting fills in corrupted video content with plausible replacements. While recent advances in endoscopic video inpainting have shown potential for enhancing the quality of endoscopic videos, they mainly repair 2D visual…

图像与视频处理 · 电气工程与系统科学 2024-07-04 Francis Xiatian Zhang , Shuang Chen , Xianghua Xie , Hubert P. H. Shum

We present the outcomes of a recent large-scale subjective study of Mobile Cloud Gaming Video Quality Assessment (MCG-VQA) on a diverse set of gaming videos. Rapid advancements in cloud services, faster video encoding technologies, and…

计算机视觉与模式识别 · 计算机科学 2024-06-25 Avinab Saha , Yu-Chih Chen , Chase Davis , Bo Qiu , Xiaoming Wang , Rahul Gowda , Ioannis Katsavounidis , Alan C. Bovik

Recent advancements in neural surface reconstruction have significantly enhanced 3D reconstruction. However, current real world datasets mainly focus on benchmarking multiview stereo (MVS) based on RGB inputs. Multiview photometric stereo…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Zaiyan Yang , Jieji Ren , Xiangyi Wang , zonglin li , Xu Cao , Heng Guo , Zhanyu Ma , Boxin Shi

Depth perception plays an essential role in the viewer experience for immersive virtual reality (VR) visual environments. However, previous research investigations in the depth quality of 3D/stereoscopic images are rather limited, and in…

计算机视觉与模式识别 · 计算机科学 2024-08-20 Wei Zhou , Zhou Wang

Although there have been notable advancements in video compression technologies in recent years, banding artifacts remain a serious issue affecting the quality of compressed videos, particularly on smooth regions of high-definition videos.…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Qi Zheng , Li-Heng Chen , Chenlong He , Neil Berkbeck , Yilin Wang , Balu Adsumilli , Alan C. Bovik , Yibo Fan , Zhengzhong Tu

Based on the Just-Noticeable-Difference (JND) criterion, a subjective video quality assessment (VQA) dataset, called the VideoSet, was constructed recently. In this work, we propose a JND-based VQA model using a probabilistic framework to…

多媒体 · 计算机科学 2018-07-04 Haiqiang Wang , Xinfeng Zhang , Chao Yang , C. -C. Jay Kuo

Laparoscopic videos can be affected by different distortions which may impact the performance of surgery and introduce surgical errors. In this work, we propose a framework for automatically detecting and identifying such distortions and…

We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeVAn (Dense Video Annotation). The dataset contains 8.5K…

计算机视觉与模式识别 · 计算机科学 2024-08-12 Tingkai Liu , Yunzhe Tao , Haogeng Liu , Qihang Fan , Ding Zhou , Huaibo Huang , Ran He , Hongxia Yang