中文
相关论文

相关论文: Capturing Co-existing Distortions in User-Generate…

200 篇论文

Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feeding frame-level…

Recent advances in Multimodal Large Language Models (MLLMs) have introduced a paradigm shift for Image Quality Assessment (IQA) from unexplainable image quality scoring to explainable IQA, demonstrating practical applications like quality…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Wenjie Liao , Jieyu Yuan , Yifang Xu , Chunle Guo , Zilong Zhang , Jihong Li , Jiachen Fu , Haotian Fan , Tao Li , Junhui Cui , Chongyi Li

Despite rapid advancements in video generation models, aligning their outputs with complex user intent remains challenging. Existing test-time optimization methods are typically either computationally expensive or require white-box access…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Yiwen Song , Tomas Pfister , Yale Song

The main challenge in video question answering (VideoQA) is to capture and understand the complex spatial and temporal relations between objects based on given questions. Existing graph-based methods for VideoQA usually ignore keywords in…

计算机视觉与模式识别 · 计算机科学 2023-07-26 Yi Cheng , Hehe Fan , Dongyun Lin , Ying Sun , Mohan Kankanhalli , Joo-Hwee Lim

AI-driven video generation techniques have made significant progress in recent years. However, AI-generated videos (AGVs) involving human activities often exhibit substantial visual and semantic distortions, hindering the practical…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Zhichao Zhang , Wei Sun , Xinyue Li , Yunhao Li , Qihang Ge , Jun Jia , Zicheng Zhang , Zhongpeng Ji , Fengyu Sun , Shangling Jui , Xiongkuo Min , Guangtao Zhai

Point cloud is one of the most widely used digital representation formats for three-dimensional (3D) contents, the visual quality of which may suffer from noise and geometric shift distortions during the production procedure as well as…

计算机视觉与模式识别 · 计算机科学 2023-12-07 Zicheng Zhang , Wei Sun , Yucheng Zhu , Xiongkuo Min , Wei Wu , Ying Chen , Guangtao Zhai

Blind video quality assessment (BVQA) plays an indispensable role in monitoring and improving the end-users' viewing experience in various real-world video-enabled media applications. As an experimental field, the improvements of BVQA…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Wei Sun , Wen Wen , Xiongkuo Min , Long Lan , Guangtao Zhai , Kede Ma

Recent advancements in video-language understanding have been established on the foundation of image-text models, resulting in promising outcomes due to the shared knowledge between images and videos. However, video-language understanding…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Xiao Wang , Yaoyu Li , Tian Gan , Zheng Zhang , Jingjing Lv , Liqiang Nie

We propose a new model for no-reference video quality assessment (VQA). Our approach uses a new idea of highly-localized space-time (ST) slices called Space-Time Chips (ST Chips). ST Chips are localized cuts of video data along directions…

图像与视频处理 · 电气工程与系统科学 2021-10-04 Joshua P. Ebenezer , Zaixi Shang , Yongjun Wu , Hai Wei , Sriram Sethuraman , Alan C. Bovik

Intermediate features of a pre-trained model have been shown informative for making accurate predictions on downstream tasks, even if the model backbone is kept frozen. The key challenge is how to utilize these intermediate features given…

机器学习 · 计算机科学 2023-04-28 Cheng-Hao Tu , Zheda Mai , Wei-Lun Chao

High Dynamic Range (HDR) videos enhance visual experiences with superior brightness, contrast, and color depth. The surge of User-Generated Content (UGC) on platforms like YouTube and TikTok introduces unique challenges for HDR video…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Shreshth Saini , Alan C. Bovik , Neil Birkbeck , Yilin Wang , Balu Adsumilli

Accurate and efficient Video Quality Assessment (VQA) has long been a key research challenge. Current mainstream VQA methods typically improve performance by pretraining on large-scale classification datasets (e.g., ImageNet, Kinetics-400),…

计算机视觉与模式识别 · 计算机科学 2025-10-10 Yachun Mi , Yu Li , Yanting Li , Chen Hui , Tong Zhang , Zhixuan Li , Chenyue Song , Wei Yang Bryan Lim , Shaohui Liu

Millions of cameras at edge are being deployed to power a variety of different deep learning applications. However, the frames captured by these cameras are not always pristine - they can be distorted due to lighting issues, sensor noise,…

图像与视频处理 · 电气工程与系统科学 2021-10-27 Sibendu Paul , Utsav Drolia , Y. Charlie Hu , Srimat T. Chakradhar

Current state-of-the-art video quality models, such as VMAF, give excellent prediction results by comparing the degraded video with its reference video. However, they do not consider temporal distortions (e.g., frame freezes or skips) that…

图像与视频处理 · 电气工程与系统科学 2023-03-23 Gabriel Mittag , Babak Naderi , Vishak Gopal , Ross Cutler

Recent advances in video reward models and post-training strategies have improved text-to-video (T2V) generation. While these models typically assess visual quality, motion quality, and text alignment, they often overlook key structural…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Yuan Wang , Borui Liao , Huijuan Huang , Jinda Lu , Ouxiang Li , Kuien Liu , Meng Wang , Xiang Wang

In networked video applications, the frame rate (FR) and quantization stepsize (QS) of a compressed video are often adapted in response to the changes of the available bandwidth. It is important to understand how do the variation of FR and…

多媒体 · 计算机科学 2014-06-10 Yen-Fu Ou , Wenzhi Lin , Huiqi Zeng , Yao Wang

In this paper, we address the problem of enhancing perceptual quality in video super-resolution (VSR) using Diffusion Models (DMs) while ensuring temporal consistency among frames. We present StableVSR, a VSR method based on DMs that can…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Claudio Rota , Marco Buzzelli , Joost van de Weijer

In recent years, with the vigorous development of the video game industry, the proportion of gaming videos on major video websites like YouTube has dramatically increased. However, relatively little research has been done on the automatic…

图像与视频处理 · 电气工程与系统科学 2022-04-15 Xiangxu Yu , Zhengzhong Tu , Neil Birkbeck , Yilin Wang , Balu Adsumilli , Alan C. Bovik

The rising popularity of online User-Generated-Content (UGC) in the form of streamed and shared videos, has hastened the development of perceptual Video Quality Assessment (VQA) models, which can be used to help optimize their delivery.…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Xiangxu Yu , Zhenqiang Ying , Neil Birkbeck , Yilin Wang , Balu Adsumilli , Alan C. Bovik

Video conferencing, which includes both video and audio content, has contributed to dramatic increases in Internet traffic, as the COVID-19 pandemic forced millions of people to work and learn from home. Global Internet traffic of video…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Zhenqiang Ying , Deepti Ghadiyaram , Alan Bovik