用于下采样色彩空间中基于深度学习的端到端图像/视频编码的变换网络架构
图像与视频处理
2021-08-30 v2 人工智能
计算机视觉与模式识别
机器学习
多媒体
摘要
现有大多数基于深度学习的端到端图像/视频编码(DLEC)架构是为非下采样RGB色彩格式设计的。然而,为获得更优编码性能,许多最先进的基于块的压缩标准如高效视频编码(HEVC/H.265)和通用视频编码(VVC/H.266)主要面向YUV 4:2:0格式设计,其中考虑人类视觉系统对U和V分量进行下采样。本文研究了多种支持YUV 4:2:0格式的DLEC设计,并在通用评估框架下将其性能与HEVC和VVC标准主档对比。此外,提出了一种新的变换网络架构以提升YUV 4:2:0数据编码效率。在YUV 4:2:0数据集上的实验结果表明,所提架构显著优于为RGB格式设计的现有架构的朴素扩展,并相较HEVC帧内编码取得约10%的平均BD-rate提升。
引用
@article{arxiv.2103.01760,
title = {Transform Network Architectures for Deep Learning based End-to-End Image/Video Coding in Subsampled Color Spaces},
author = {Hilmi E. Egilmez and Ankitesh K. Singh and Muhammed Coban and Marta Karczewicz and Yinhao Zhu and Yang Yang and Amir Said and Taco S. Cohen},
journal= {arXiv preprint arXiv:2103.01760},
year = {2021}
}
备注
10 pages, accepted in IEEE Open Journal of Signal Processing (Special issue on Applied Artificial Intelligence and Machine Learning for Video Coding and Streaming)