English

E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control

Computer Vision and Pattern Recognition 2026-05-27 v1 Artificial Intelligence

Abstract

Controllable and physically grounded egocentric video generation is essential for embodied agents to reason about how their own and others' actions manifest and change the world. Compared to generic video synthesis, egocentric generation is especially challenging: the camera is tightly coupled to the actor, leading to rapid viewpoint changes and frequent self-occlusions; the underlying actions are subtle, articulated, and often only partially visible; and both the people and the scene state must evolve consistently with the specified controls. We present E3^3C, a controllable video diffusion framework for egocentric generation that builds structured and compact conditions disentangling persistent scene structure from human-driven dynamics. From context frames, E3^3C constructs a semi-dense point cloud-based 3D memory and augments each point with appearance descriptors from video-VAE features. Rendering this memory into target viewpoints produces conditioning aligned with the target frames. Human dynamics are modeled separately. The observed people in the scene are controlled by skeleton renderings (exo human control), while the camera wearer is specified by their 3D body joints and 6DoF wrist motion (ego human control). To preserve ego human control when the wearer's body parts are invisible, we introduce an ego motion encoder that produces persistent cross-attention tokens. Experiments on Nymeria show that E3^3C improves visual fidelity, camera-motion accuracy, object consistency, and ego & exo human control over strong baselines, while also enabling intuitive scene editing.

Keywords

Cite

@article{arxiv.2605.26316,
  title  = {E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control},
  author = {Qiao Gu and Lingni Ma and Adam W Harley and Richard Newcombe and Florian Shkurti and Julian Straub},
  journal= {arXiv preprint arXiv:2605.26316},
  year   = {2026}
}

Comments

Preprint. Project Page: https://e3c-videogen.github.io/

R2 v1 2026-07-22T07:33:21.954Z