English
Related papers

Related papers: StableAvatar: Infinite-Length Audio-Driven Avatar …

200 papers

Although powerful for image generation, consistent and controllable video is a longstanding problem for diffusion models. Video models require extensive training and computational resources, leading to high costs and large environmental…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Muhammad Haaris Khan , Hadrien Reynaud , Bernhard Kainz

Avatar video generation models have achieved remarkable progress in recent years. However, prior work exhibits limited efficiency in generating long-duration high-resolution videos, suffering from temporal drifting, quality degradation, and…

The rising demand for creating lifelike avatars in the digital realm has led to an increased need for generating high-quality human videos guided by textual descriptions and poses. We propose Dancing Avatar, designed to fabricate human…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Bosheng Qin , Wentao Ye , Qifan Yu , Siliang Tang , Yueting Zhuang

Despite advances in audio-driven video generation, achieving commercial-grade stability remains challenging. We present LongCat-Video-Avatar 1.5, an upgraded open-source framework prioritizing systematic engineering and production-readiness…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Meituan LongCat Team , Xunliang Cai , Meng Cheng , Feng Gao , Zhe Kong , Jiamu Li , Le Li , Weiheng Li , Hongyu Liu , Shuai Tan , Xiaoming Wei , Tianyu Yang , Yong Zhang

Prevailing Video-to-Audio (V2A) generation models operate offline, assuming an entire video sequence or chunks of frames are available beforehand. This critically limits their use in interactive applications such as live content creation…

We propose VLOGGER, a method for audio-driven human video generation from a single input image of a person, which builds on the success of recent generative diffusion models. Our method consists of 1) a stochastic human-to-3d-motion…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Enric Corona , Andrei Zanfir , Eduard Gabriel Bazavan , Nikos Kolotouros , Thiemo Alldieck , Cristian Sminchisescu

This paper presents InfiniteAudio, a simple yet effective strategy for generating infinite-length audio using diffusion-based text-to-audio methods. Current approaches face memory constraints because the output size increases with input…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-04 Chaeyoung Jung , Hojoon Ki , Ji-Hoon Kim , Junmo Kim , Joon Son Chung

Recent advances in audio-driven avatar video generation have significantly enhanced audio-visual realism. However, existing methods treat instruction conditioning merely as low-level tracking driven by acoustic or visual cues, without…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Yikang Ding , Jiwen Liu , Wenyuan Zhang , Zekun Wang , Wentao Hu , Liyuan Cui , Mingming Lao , Yingchao Shao , Hui Liu , Xiaohan Li , Ming Chen , Xiaoqiang Liu , Yu-Shen Liu , Pengfei Wan

Stable Audio 3 is a family of fast latent diffusion models (small, medium, large) for variable-length audio generation and editing. Since our models can generate several minutes of audio, variable-length generations are key to avoid the…

Sound · Computer Science 2026-05-19 Zach Evans , Julian D. Parker , Matthew Rice , CJ Carr , Zack Zukowski , Josiah Taylor , Jordi Pons

SmartAvatar is a vision-language-agent-driven framework for generating fully rigged, animation-ready 3D human avatars from a single photo or textual prompt. While diffusion-based methods have made progress in general 3D object generation,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Alexander Huang-Menders , Xinhang Liu , Andy Xu , Yuyao Zhang , Chi-Keung Tang , Yu-Wing Tai

Video inverse problems are fundamental to streaming, telepresence, and AR/VR, where high perceptual quality must coexist with tight latency constraints. Diffusion-based priors currently deliver state-of-the-art reconstructions, but existing…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Weimin Bai , Suzhe Xu , Yiwei Ren , Jinhua Hao , Ming Sun , Wenzheng Chen , He Sun

Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Zhen Ye , Xu Tan , Aoxiong Yin , Hongzhan Lin , Guangyan Zhang , Peiwen Sun , Yiming Li , Chi-Min Chan , Wei Ye , Shikun Zhang , Wei Xue

Recent breakthroughs in video AIGC have ushered in a transformative era for audio-driven human animation. However, conventional video dubbing techniques remain constrained to mouth region editing, resulting in discordant facial expressions…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Shaoshu Yang , Zhe Kong , Feng Gao , Meng Cheng , Xiangyu Liu , Yong Zhang , Zhuoliang Kang , Wenhan Luo , Xunliang Cai , Ran He , Xiaoming Wei

Text-based diffusion models have exhibited remarkable success in generation and editing, showing great promise for enhancing visual content with their generative prior. However, applying these models to video super-resolution remains…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Shangchen Zhou , Peiqing Yang , Jianyi Wang , Yihang Luo , Chen Change Loy

Diffusion-based audio-driven talking avatar methods have recently gained attention for their high-fidelity, vivid, and expressive results. However, their slow inference speed limits practical applications. Despite the development of various…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Tianyun Zhong , Chao Liang , Jianwen Jiang , Gaojie Lin , Jiaqi Yang , Zhou Zhao

This paper focuses on the task of speech-driven 3D facial animation, which aims to generate realistic and synchronized facial motions driven by speech inputs. Recent methods have employed audio-conditioned diffusion models for 3D facial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yifan Yang , Zhi Cen , Sida Peng , Xiangwei Chen , Yifu Deng , Xinyu Zhu , Fan Jia , Xiaowei Zhou , Hujun Bao

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader…

Sound · Computer Science 2025-03-25 Yong Ren , Chenxing Li , Manjie Xu , Wei Liang , Yu Gu , Rilin Chen , Dong Yu

Diffusion-based video generation techniques have significantly improved zero-shot talking-head avatar generation, enhancing the naturalness of both head motion and facial expressions. However, existing methods suffer from poor…

Graphics · Computer Science 2025-04-24 Lingzhou Mu , Baiji Liu , Ruonan Zhang , Guiming Mo , Jiawei Jin , Kai Zhang , Haozhi Huang

Video generation has drawn significant interest recently, pushing the development of large-scale models capable of producing realistic videos with coherent motion. Due to memory constraints, these models typically generate short video…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Idan Kligvasser , Regev Cohen , George Leifman , Ehud Rivlin , Michael Elad

In this paper, we explore the overlooked challenge of stability and temporal consistency in interactive video generation, which synthesizes dynamic and controllable video worlds through interactive behaviors such as camera movements and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Ying Yang , Zhengyao Lv , Tianlin Pan , Haofan Wang , Binxin Yang , Hubery Yin , Chen Li , Ziwei Liu , Chenyang Si