中文

PortraitDirector:用于可控且实时的面部重作的层次化解缝框架

计算机视觉与模式识别 2026-04-22 v1

摘要

现有的面部重作方法在表达性和细粒度可控性之间存在权衡。整体性面部重作模型往往为表达性牺牲细粒度控制,而专为控制设计的方法可能在保真度和鲁棒的解缝方面受限。我们不将面部运动视为单一信号,而是探索另一种组合性视角。本文提出 PortraitDirector,一种新型框架,将面部重作形式化为层次化组合任务,实现高保真度和可控性。我们采用层次化运动解缝与组合策略,将面部运动分解为用于物理运动的空间层和用于情感内容的语义层。空间层包括:(i) 通过专用表示和注入路径管理的全局头部姿态;(ii) 从裁剪后的面部区域中提炼并通过情感滤波模块(利用信息瓶颈)剔除情感线索的空间分离局部面部表情。语义层包含派生的全局情感。解缝后的组件随后被组合成具有表达性的运动潜在向量。 Furthermore, we engineer the framework for real-time performance through a suite of optimizations, including diffusion distillation, causal attention and VAE acceleration. PortraitDirector achieves streaming, high-fidelity, controllable 512 x 512 face reenactment at 20 FPS with a end-to-end 800 ms latency on a single 5090 GPU. 我们通过一系列优化措施(包括扩散蒸馏、因果注意力和 VAE 加速)为实现实时性而进行工程化设计。PortraitDirector 在单张 5090 GPU 上以 20 FPS 的速度实现 512x512 的流式、高保真、可控面部重作,端到端延迟为 800 毫秒。

关键词

引用

@article{arxiv.2604.19129,
  title  = {PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment},
  author = {Chaonan Ji and Jinwei Qi and Sheng Xu and Peng Zhang and Bang Zhang},
  journal= {arXiv preprint arXiv:2604.19129},
  year   = {2026}
}

备注

accepted by CVPR2026