English

Omnimatte: Associating Objects and Their Effects in Video

Computer Vision and Pattern Recognition 2021-10-04 v2

Abstract

Computer vision is increasingly effective at segmenting objects in images and videos; however, scene effects related to the objects -- shadows, reflections, generated smoke, etc -- are typically overlooked. Identifying such scene effects and associating them with the objects producing them is important for improving our fundamental understanding of visual scenes, and can also assist a variety of applications such as removing, duplicating, or enhancing objects in video. In this work, we take a step towards solving this novel problem of automatically associating objects with their effects in video. Given an ordinary video and a rough segmentation mask over time of one or more subjects of interest, we estimate an omnimatte for each subject -- an alpha matte and color image that includes the subject along with all its related time-varying scene elements. Our model is trained only on the input video in a self-supervised manner, without any manual labels, and is generic -- it produces omnimattes automatically for arbitrary objects and a variety of effects. We show results on real-world videos containing interactions between different types of subjects (cars, animals, people) and complex effects, ranging from semi-transparent elements such as smoke and reflections, to fully opaque effects such as objects attached to the subject.

Keywords

Cite

@article{arxiv.2105.06993,
  title  = {Omnimatte: Associating Objects and Their Effects in Video},
  author = {Erika Lu and Forrester Cole and Tali Dekel and Andrew Zisserman and William T. Freeman and Michael Rubinstein},
  journal= {arXiv preprint arXiv:2105.06993},
  year   = {2021}
}

Comments

CVPR 2021 Oral. Project webpage: https://omnimatte.github.io/. Added references

R2 v1 2026-06-24T02:07:33.731Z