English

Generating Fit Check Videos with a Handheld Camera

Computer Vision and Pattern Recognition 2025-12-02 v2

Abstract

Self-captured full-body videos are popular, but most deployments require mounted cameras, carefully-framed shots, and repeated practice. We propose a more convenient solution that enables full-body video capture using handheld mobile devices. Our approach takes as input two static photos (front and back) of you in a mirror, along with an IMU motion reference that you perform while holding your mobile phone, and synthesizes a realistic video of you performing a similar target motion. We enable rendering into a new scene, with consistent illumination and shadows. We propose a novel video diffusion-based model to achieve this. Specifically, we propose a parameter-free frame generation strategy and a multi-reference attention mechanism to effectively integrate appearance information from both the front and back selfies into the video diffusion model. Further, we introduce an image-based fine-tuning strategy to enhance frame sharpness and improve shadows and reflections generation for more realistic human-scene composition.

Keywords

Cite

@article{arxiv.2505.23886,
  title  = {Generating Fit Check Videos with a Handheld Camera},
  author = {Bowei Chen and Brian Curless and Ira Kemelmacher-Shlizerman and Steven M. Seitz},
  journal= {arXiv preprint arXiv:2505.23886},
  year   = {2025}
}
R2 v1 2026-07-01T02:49:13.990Z