English

Livatar-1: Real-Time Talking Heads Generation with Tailored Flow Matching

Computer Vision and Pattern Recognition 2025-07-28 v1

Abstract

We present Livatar, a real-time audio-driven talking heads videos generation framework. Existing baselines suffer from limited lip-sync accuracy and long-term pose drift. We address these limitations with a flow matching based framework. Coupled with system optimizations, Livatar achieves competitive lip-sync quality with a 8.50 LipSync Confidence on the HDTF dataset, and reaches a throughput of 141 FPS with an end-to-end latency of 0.17s on a single A10 GPU. This makes high-fidelity avatars accessible to broader applications. Our project is available at https://www.hedra.com/ with with examples at https://h-liu1997.github.io/Livatar-1/

Keywords

Cite

@article{arxiv.2507.18649,
  title  = {Livatar-1: Real-Time Talking Heads Generation with Tailored Flow Matching},
  author = {Haiyang Liu and Xiaolin Hong and Xuancheng Yang and Yudi Ruan and Xiang Lian and Michael Lingelbach and Hongwei Yi and Wei Li},
  journal= {arXiv preprint arXiv:2507.18649},
  year   = {2025}
}

Comments

Technical Report