中文
相关论文

相关论文: Dimitra: Audio-driven Diffusion model for Expressi…

200 篇论文

In this paper, we consider a novel and practical case for talking face video generation. Specifically, we focus on the scenarios involving multi-people interactions, where the talking context, such as audience or surroundings, is present.…

计算机视觉与模式识别 · 计算机科学 2024-02-29 Meidai Xuanyuan , Yuwang Wang , Honglei Guo , Qionghai Dai

We propose a novel 3D-aware diffusion-based method for generating photorealistic talking head videos directly from a single identity image and explicit control signals (e.g., expressions). Our method generates Multiplane Images (MPIs) that…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Yuan Li , Ziqian Bai , Feitong Tan , Zhaopeng Cui , Sean Fanello , Yinda Zhang

Recent advances in audio-driven talking head generation have achieved impressive results in lip synchronization and emotional expression. However, they largely overlook the crucial task of facial attribute editing. This capability is…

计算机视觉与模式识别 · 计算机科学 2025-08-28 Guanwen Feng , Zhiyuan Ma , Yunan Li , Jiahao Yang , Junwei Jing , Qiguang Miao

We describe our novel deep learning approach for driving animated faces using both acoustic and visual information. In particular, speech-related facial movements are generated using audiovisual information, and non-speech facial movements…

音频与语音处理 · 电气工程与系统科学 2020-05-29 Ahmed Hussen Abdelaziz , Barry-John Theobald , Paul Dixon , Reinhard Knothe , Nicholas Apostoloff , Sachin Kajareker

Talking face generation has gained significant attention as a core application of generative models. To enhance the expressiveness and realism of synthesized videos, emotion editing in talking face video plays a crucial role. However,…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Chanhyuk Choi , Taesoo Kim , Donggyu Lee , Siyeol Jung , Taehwan Kim

Singing-driven 3D head animation is a challenging yet promising task with applications in virtual avatars, entertainment, and education. Unlike speech, singing involves richer emotional nuance, dynamic prosody, and lyric-based semantics,…

图形学 · 计算机科学 2025-09-03 Zikai Huang , Yihan Zhou , Xuemiao Xu , Cheng Xu , Xiaofen Xing , Jing Qin , Shengfeng He

Speech-driven 3D talking head generation aims to produce lifelike facial animations precisely synchronized with speech. While considerable progress has been made in achieving high lip-synchronization accuracy, existing methods largely…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Bin Wang , Yang Xu , Huan Zhao , Hao Zhang , Zixing Zhang

This paper presents a novel approach for generating 3D talking heads from raw audio inputs. Our method grounds on the idea that speech related movements can be comprehensively and efficiently described by the motion of a few control points…

计算机视觉与模式识别 · 计算机科学 2023-07-27 Federico Nocentini , Claudio Ferrari , Stefano Berretti

While increasing attention has been paid to co-speech gesture synthesis, most previous works neglect to investigate hand gestures with explicit and essential semantics. In this paper, we study co-speech gesture generation with an emphasis…

Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper introduces JoyGen, a…

计算机视觉与模式识别 · 计算机科学 2025-01-06 Qili Wang , Dajiang Wu , Zihang Xu , Junshi Huang , Jun Lv

Audio-driven 3D talking avatar generation is increasingly important in virtual communication, digital humans, and interactive media, where avatars must preserve identity, synchronize lip motion with speech, express emotion, and exhibit…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Zhongju Wang , Zhenhong Sun , Beier Wang , Yifu Wang , Daoyi Dong , Huadong Mo , Hongdong Li

Speech-driven 3D facial animation with accurate lip synchronization has been widely studied. However, synthesizing realistic motions for the entire face during speech has rarely been explored. In this work, we present a joint audio-text…

计算机视觉与模式识别 · 计算机科学 2021-12-08 Yingruo Fan , Zhaojiang Lin , Jun Saito , Wenping Wang , Taku Komura

Recent advancements in deep learning and computer vision have led to a surge of interest in generating realistic talking heads. This paper presents a comprehensive survey of state-of-the-art methods for talking head generation. We…

计算机视觉与模式识别 · 计算机科学 2023-08-31 Shreyank N Gowda , Dheeraj Pandey , Shashank Narayana Gowda

This paper presents FaceXHuBERT, a text-less speech-driven 3D facial animation generation method that allows to capture personalized and subtle cues in speech (e.g. identity, emotion and hesitation). It is also very robust to background…

计算机视觉与模式识别 · 计算机科学 2023-03-10 Kazi Injamamul Haque , Zerrin Yumak

Talking head synthesis, an advanced method for generating portrait videos from a still image driven by specific content, has garnered widespread attention in virtual reality, augmented reality and game production. Recently, significant…

计算机视觉与模式识别 · 计算机科学 2024-06-19 Ming Meng , Yufei Zhao , Bo Zhang , Yonggui Zhu , Weimin Shi , Maxwell Wen , Zhaoxin Fan

Sora-like video generation models have achieved remarkable progress with a Multi-Modal Diffusion Transformer MM-DiT architecture. However, the current video generation models predominantly focus on single-prompt, struggling to generate…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Minghong Cai , Xiaodong Cun , Xiaoyu Li , Wenze Liu , Zhaoyang Zhang , Yong Zhang , Ying Shan , Xiangyu Yue

We present a framework for modeling interactional communication in dyadic conversations: given multimodal inputs of a speaker, we autoregressively output multiple possibilities of corresponding listener motion. We combine the motion and…

计算机视觉与模式识别 · 计算机科学 2022-04-19 Evonne Ng , Hanbyul Joo , Liwen Hu , Hao Li , Trevor Darrell , Angjoo Kanazawa , Shiry Ginosar

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Jiaxin Ye , Hongming Shan

We present DiffusionBERT, a new generative masked language model based on discrete diffusion models. Diffusion models and many pre-trained language models have a shared training objective, i.e., denoising, making it possible to combine the…

计算与语言 · 计算机科学 2022-12-02 Zhengfu He , Tianxiang Sun , Kuanning Wang , Xuanjing Huang , Xipeng Qiu

In this work, we tackle the challenge of enhancing the realism and expressiveness in talking head video generation by focusing on the dynamic and nuanced relationship between audio cues and facial movements. We identify the limitations of…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Linrui Tian , Qi Wang , Bang Zhang , Liefeng Bo
‹ 上一页 1 8 9 10 下一页 ›