English

M&M Mix: A Multimodal Multiview Transformer Ensemble

Computer Vision and Pattern Recognition 2022-06-22 v1

Abstract

This report describes the approach behind our winning solution to the 2022 Epic-Kitchens Action Recognition Challenge. Our approach builds upon our recent work, Multiview Transformer for Video Recognition (MTV), and adapts it to multimodal inputs. Our final submission consists of an ensemble of Multimodal MTV (M&M) models varying backbone sizes and input modalities. Our approach achieved 52.8% Top-1 accuracy on the test set in action classes, which is 4.1% higher than last year's winning entry.

Keywords

Cite

@article{arxiv.2206.09852,
  title  = {M&M Mix: A Multimodal Multiview Transformer Ensemble},
  author = {Xuehan Xiong and Anurag Arnab and Arsha Nagrani and Cordelia Schmid},
  journal= {arXiv preprint arXiv:2206.09852},
  year   = {2022}
}

Comments

Technical report for Epic-Kitchens challenge 2022

R2 v1 2026-06-24T11:57:26.094Z