English

Learning Perceptual Representations for Gaming NR-VQA with Multi-Task FR Signals

Image and Video Processing 2026-02-20 v2 Computer Vision and Pattern Recognition Multimedia

Abstract

No-reference video quality assessment (NR-VQA) for gaming videos is challenging due to limited human-rated datasets and unique content characteristics including fast motion, stylized graphics, and compression artifacts. We present MTL-VQA, a multi-task learning framework that uses full-reference metrics as supervisory signals to learn perceptually meaningful features without human labels for pretraining. By jointly optimizing multiple full-reference (FR) objectives with adaptive task weighting, our approach learns shared representations that transfer effectively to NR-VQA. Experiments on gaming video datasets show MTL-VQA achieves performance competitive with state-of-the-art NR-VQA methods across both MOS-supervised and label-efficient/self-supervised settings.

Keywords

Cite

@article{arxiv.2602.11903,
  title  = {Learning Perceptual Representations for Gaming NR-VQA with Multi-Task FR Signals},
  author = {Yu-Chih Chen and Michael Wang and Chieh-Dun Wen and Kai-Siang Ma and Avinab Saha and Li-Heng Chen and Alan Bovik},
  journal= {arXiv preprint arXiv:2602.11903},
  year   = {2026}
}

Comments

6 pages, 2 figures

R2 v1 2026-07-01T10:33:35.669Z