中文
相关论文

相关论文: MDS-VQA: Model-Informed Data Selection for Video Q…

200 篇论文

Completely blind video quality assessment (VQA) refers to a class of quality assessment methods that do not use any reference videos, human opinion scores or training videos from the target database to learn a quality model. The design of…

图像与视频处理 · 电气工程与系统科学 2024-06-25 Shankhanil Mitra , Rajiv Soundararajan

We introduce HIDRO-VQA, a no-reference (NR) video quality assessment model designed to provide precise quality evaluations of High Dynamic Range (HDR) videos. HDR videos exhibit a broader spectrum of luminance, detail, and color than…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Shreshth Saini , Avinab Saha , Alan C. Bovik

No-reference video quality assessment (NR-VQA) for user generated content (UGC) is crucial for understanding and improving visual experience. Unlike video recognition tasks, VQA tasks are sensitive to changes in input resolution. Since…

计算机视觉与模式识别 · 计算机科学 2023-03-31 Junjie Ke , Tianhao Zhang , Yilin Wang , Peyman Milanfar , Feng Yang

Owing to the proliferation of user-generated videos on the Internet, blind video quality assessment (BVQA) at the edge attracts growing attention. The usage of deep-learning-based methods is restricted to be applied at the edge due to their…

图像与视频处理 · 电气工程与系统科学 2023-10-31 Zhanxuan Mei , Yun-Cheng Wang , C. -C. Jay Kuo

Reasoning about causal and temporal event relations in videos is a new destination of Video Question Answering (VideoQA).The major stumbling block to achieve this purpose is the semantic gap between language and video since they are at…

计算机视觉与模式识别 · 计算机科学 2022-11-03 Shaoning Xiao , Long Chen , Kaifeng Gao , Zhao Wang , Yi Yang , Zhimeng Zhang , Jun Xiao

With the rapid advancement of vision language models(VLM), their ability to assess visual content based on specific criteria and dimensions has become increasingly critical for applications such as video-theme consistency assessment and…

计算机视觉与模式识别 · 计算机科学 2025-08-18 Baihong Qian , Haotian Fan , Wenjie Liao , Yunqiu Wang , Tao Li , Junhui Cui

As multimedia services such as video streaming, video conferencing, virtual reality (VR), and online gaming continue to expand, ensuring high perceptual visual quality becomes a priority to maintain user satisfaction and competitiveness.…

多媒体 · 计算机科学 2025-03-04 Wei Zhou , Hadi Amirpour , Christian Timmerer , Guangtao Zhai , Patrick Le Callet , Alan C. Bovik

Perceptual quality assessment of user generated content (UGC) videos is challenging due to the requirement of large scale human annotated videos for training. In this work, we address this challenge by first designing a self-supervised…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Shankhanil Mitra , Rajiv Soundararajan

User-generated content (UGC) live videos are often bothered by various distortions during capture procedures and thus exhibit diverse visual qualities. Such source videos are further compressed and transcoded by media server providers…

计算机视觉与模式识别 · 计算机科学 2023-04-20 Zicheng Zhang , Wei Wu , Wei Sun , Dangyang Tu , Wei Lu , Xiongkuo Min , Ying Chen , Guangtao Zhai

Recently, User-Generated Content (UGC) videos have gained popularity in our daily lives. However, UGC videos often suffer from poor exposure due to the limitations of photographic equipment and techniques. Therefore, Video Exposure…

计算机视觉与模式识别 · 计算机科学 2024-05-15 Xunchu Zhou , Xiaohong Liu , Yunlong Dong , Tengchuan Kou , Yixuan Gao , Zicheng Zhang , Chunyi Li , Haoning Wu , Guangtao Zhai

The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Yuanhan Zhang , Jinming Wu , Wei Li , Bo Li , Zejun Ma , Ziwei Liu , Chunyuan Li

Visual Quality Assessment (QA) seeks to predict human perceptual judgments of visual fidelity. While recent multimodal large language models (MLLMs) show promise in reasoning about image and video quality, existing approaches mainly rely on…

计算机视觉与模式识别 · 计算机科学 2025-11-10 Zehui Feng , Tian Qiu , Tong Wu , Junxuan Li , Huayuan Xu , Ting Han

The study of video prediction models is believed to be a fundamental approach to representation learning for videos. While a plethora of generative models for predicting the future frame pixel values given the past few frames exist, the…

图像与视频处理 · 电气工程与系统科学 2023-05-09 Nagabhushan Somraj , Manoj Surya Kashi , S. P. Arun , Rajiv Soundararajan

We propose the LEHA-CVQAD (Large-scale Enriched Human-Annotated Compressed Video Quality Assessment) dataset, which comprises 6,240 clips for compression-oriented video quality assessment. 59 source videos are encoded with 186 codec-preset…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Aleksandr Gushchin , Maksim Smirnov , Dmitriy Vatolin , Anastasia Antsiferova

Through a simple multiple choice language prompt a VQA model can operate as a zero-shot image classifier, producing a classification label. Compared to typical image encoders, VQA models offer an advantage: VQA-produced image embeddings can…

机器学习 · 计算机科学 2024-07-24 Dipika Khullar , Emmett Goodman , Negin Sokhandan

Recently, Users Generated Content (UGC) videos becomes ubiquitous in our daily lives. However, due to the limitations of photographic equipments and techniques, UGC videos often contain various degradations, in which one of the most…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Yunlong Dong , Xiaohong Liu , Yixuan Gao , Xunchu Zhou , Tao Tan , Guangtao Zhai

With the rapid growth of Internet video data amounts and types, a unified Video Quality Assessment (VQA) is needed to inspire video communication with perceptual quality. To meet the real-time and universal requirements in providing such…

多媒体 · 计算机科学 2023-03-27 Xinhui Huang , Chunyi Li , Abdelhak Bentaleb , Roger Zimmermann , Guangtao Zhai

Humans apprehend the world through various sensory modalities, yet language is their predominant communication channel. Machine learning systems need to draw on the same multimodal richness to have informed discourses with humans in natural…

计算机视觉与模式识别 · 计算机科学 2022-08-25 Min Wang , Ata Mahjoubfar , Anupama Joshi

3D Visual Question Answering (3D VQA) is crucial for enabling models to perceive the physical world and perform spatial reasoning. In 3D VQA, the free-form nature of answers often leads to improper annotations that can confuse or mislead…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Shengli Zhou , Yang Liu , Feng Zheng

Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective quantitative metrics…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Shangkun Sun , Xiaoyu Liang , Songlin Fan , Wenxu Gao , Wei Gao