中文
相关论文

相关论文: HunyuanVideo-Avatar: High-Fidelity Audio-Driven Hu…

200 篇论文

To address key limitations in human-object interaction (HOI) video generation -- specifically the reliance on curated motion data, limited generalization to novel objects/scenarios, and restricted accessibility -- we introduce…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Ziyao Huang , Zixiang Zhou , Juan Cao , Yifeng Ma , Yi Chen , Zejing Rao , Zhiyong Xu , Hongmei Wang , Qin Lin , Yuan Zhou , Qinglin Lu , Fan Tang

Customized video generation aims to produce videos featuring specific subjects under flexible user-defined conditions, yet existing methods often struggle with identity consistency and limited input modalities. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2025-05-09 Teng Hu , Zhentao Yu , Zhengguang Zhou , Sen Liang , Yuan Zhou , Qin Lin , Qinglin Lu

Recent advances in video generation produce visually realistic content, yet the absence of synchronized audio severely compromises immersion. To address key challenges in video-to-audio generation, including multimodal data scarcity,…

音频与语音处理 · 电气工程与系统科学 2025-08-26 Sizhe Shan , Qiulin Li , Yutao Cui , Miles Yang , Yuehai Wang , Qun Yang , Jin Zhou , Zhao Zhong

We introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance reference and video clips…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Zunnan Xu , Zhentao Yu , Zixiang Zhou , Jun Zhou , Xiaoyu Jin , Fa-Ting Hong , Xiaozhong Ji , Junwei Zhu , Chengfei Cai , Shiyu Tang , Qin Lin , Xiu Li , Qinglin Lu

Recent advances in diffusion-based and controllable video generation have enabled high-quality and temporally coherent video synthesis, laying the groundwork for immersive interactive gaming experiences. However, current methods face…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Jiaqi Li , Junshu Tang , Zhiyong Xu , Longhuang Wu , Yuan Zhou , Shuai Shao , Tianbao Yu , Zhiguo Cao , Qinglin Lu

Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and fluidity. They also…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Qijun Gan , Ruizi Yang , Jianke Zhu , Shaofei Xue , Steven Hoi

Head avatars animated by visual signals have gained popularity, particularly in cross-driving synthesis where the driver differs from the animated character, a challenging but highly practical approach. The recently presented MegaPortraits…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Nikita Drobyshev , Antoni Bigata Casademunt , Konstantinos Vougioukas , Zoe Landgraf , Stavros Petridis , Maja Pantic

We present Hunyuan-DiT, a text-to-image diffusion transformer with fine-grained understanding of both English and Chinese. To construct Hunyuan-DiT, we carefully design the transformer structure, text encoder, and positional encoding. We…

Intelligent game creation represents a transformative advancement in game development, utilizing generative artificial intelligence to dynamically generate and enhance game content. Despite notable progress in generative models, the…

Recent years have witnessed remarkable advances in audio-driven talking head generation. However, existing approaches predominantly focus on single-character scenarios. While some methods can create separate conversation videos between two…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Yubo Huang , Weiqiang Wang , Sirui Zhao , Tong Xu , Lin Liu , Enhong Chen

Producing expressive facial animations from static images is a challenging task. Prior methods relying on explicit geometric priors (e.g., facial landmarks or 3DMM) often suffer from artifacts in cross reenactment and struggle to capture…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Qiang Wang , Mengchao Wang , Fan Jiang , Yaqi Fan , Yonggang Qi , Mu Xu

Video-driven 3D facial animation transfer aims to drive avatars to reproduce the expressions of actors. Existing methods have achieved remarkable results by constraining both geometric and perceptual consistency. However, geometric…

图形学 · 计算机科学 2024-10-10 Feng Qiu , Wei Zhang , Chen Liu , Rudong An , Lincheng Li , Yu Ding , Changjie Fan , Zhipeng Hu , Xin Yu

Existing video avatar models can produce fluid human animations, yet they struggle to move beyond mere physical likeness to capture a character's authentic essence. Their motions typically synchronize with low-level cues like audio rhythm,…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Jianwen Jiang , Weihong Zeng , Zerong Zheng , Jiaqi Yang , Chao Liang , Wang Liao , Han Liang , Yuan Zhang , Mingyuan Gao

Whole-body audio-driven avatar pose and expression generation is a critical task for creating lifelike digital humans and enhancing the capabilities of interactive virtual agents, with wide-ranging applications in virtual reality, digital…

声音 · 计算机科学 2025-10-15 Tianbao Zhang , Jian Zhao , Yuer Li , Zheng Zhu , Ping Hu , Zhaoxin Fan , Wenjun Wu , Xuelong Li

Existing methodologies for animating portrait images face significant challenges, particularly in handling non-frontal perspectives, rendering dynamic objects around the portrait, and generating immersive, realistic backgrounds. In this…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Jiahao Cui , Hui Li , Yun Zhan , Hanlin Shang , Kaihui Cheng , Yuqi Ma , Shan Mu , Hang Zhou , Jingdong Wang , Siyu Zhu

Current state-of-the-art (SOTA) methods for audio-driven character animation demonstrate promising performance for scenarios primarily involving speech and singing. However, they often fall short in more complex film and television…

We present an audio-driven real-time system for animating photorealistic 3D facial avatars with minimal latency, designed for social interactions in virtual reality for anyone. Central to our approach is an encoder model that transforms…

图形学 · 计算机科学 2025-11-04 Jiye Lee , Chenghui Li , Linh Tran , Shih-En Wei , Jason Saragih , Alexander Richard , Hanbyul Joo , Shaojie Bai

While recent image-based human animation methods achieve realistic body and facial motion synthesis, critical gaps remain in fine-grained holistic controllability, multi-scale adaptability, and long-term temporal coherence, which leads to…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Yuxuan Luo , Zhengkun Rong , Lizhen Wang , Longhao Zhang , Tianshu Hu , Yongming Zhu

The generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial expressions. Current…

计算机视觉与模式识别 · 计算机科学 2025-01-10 Huaize Liu , Wenzhang Sun , Donglin Di , Shibo Sun , Jiahui Yang , Changqing Zou , Hujun Bao

Existing video avatar models have demonstrated impressive capabilities in scenarios such as talking, public speaking, and singing. However, the majority of these methods exhibit limited alignment with respect to text instructions,…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Ruikui Wang , Jinheng Feng , Lang Tian , Huaishao Luo , Chaochao Li , Liangbo Zhou , Huan Zhang , Youzheng Wu , Xiaodong He
‹ 上一页 1 2 3 10 下一页 ›