English

Preserving Semantic and Temporal Consistency for Unpaired Video-to-Video Translation

Computer Vision and Pattern Recognition 2019-08-22 v1 Multimedia

Abstract

In this paper, we investigate the problem of unpaired video-to-video translation. Given a video in the source domain, we aim to learn the conditional distribution of the corresponding video in the target domain, without seeing any pairs of corresponding videos. While significant progress has been made in the unpaired translation of images, directly applying these methods to an input video leads to low visual quality due to the additional time dimension. In particular, previous methods suffer from semantic inconsistency (i.e., semantic label flipping) and temporal flickering artifacts. To alleviate these issues, we propose a new framework that is composed of carefully-designed generators and discriminators, coupled with two core objective functions: 1) content preserving loss and 2) temporal consistency loss. Extensive qualitative and quantitative evaluations demonstrate the superior performance of the proposed method against previous approaches. We further apply our framework to a domain adaptation task and achieve favorable results.

Keywords

Cite

@article{arxiv.1908.07683,
  title  = {Preserving Semantic and Temporal Consistency for Unpaired Video-to-Video Translation},
  author = {Kwanyong Park and Sanghyun Woo and Dahun Kim and Donghyeon Cho and In So Kweon},
  journal= {arXiv preprint arXiv:1908.07683},
  year   = {2019}
}

Comments

Accepted by ACM Multimedia(ACM MM) 2019