中文
相关论文

相关论文: G4G:A Generic Framework for High Fidelity Talking …

200 篇论文

We present TANGO, a framework for generating co-speech body-gesture videos. Given a few-minute, single-speaker reference video and target speech audio, TANGO produces high-fidelity videos with synchronized body gestures. TANGO builds on…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Haiyang Liu , Xingchao Yang , Tomoya Akiyama , Yuantian Huang , Qiaoge Li , Shigeru Kuriyama , Takafumi Taketomi

Deepfakes are AI-generated media in which the original content is digitally altered to create convincing but manipulated images, videos, or audio. Among the various types of deepfakes, lip-syncing deepfakes are one of the most challenging…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Soumyya Kanti Datta , Shan Jia , Siwei Lyu

Articulatory-to-acoustic (forward) mapping is a technique to predict speech using various articulatory acquisition techniques as input (e.g. ultrasound tongue imaging, MRI, lip video). The advantage of lip video is that it is easily…

计算机视觉与模式识别 · 计算机科学 2021-04-30 Frigyes Viktor Arthur , Tamás Gábor Csapó

Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head…

多媒体 · 计算机科学 2023-09-21 Songlin Yang , Wei Wang , Jun Ling , Bo Peng , Xu Tan , Jing Dong

Speech-driven facial video generation has been a complex problem due to its multi-modal aspects namely audio and video domain. The audio comprises lots of underlying features such as expression, pitch, loudness, prosody(speaking style) and…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Neeraj Kumar , Srishti Goel , Ankur Narang , Brejesh Lall

Animating human face images aims to synthesize a desired source identity in a natural-looking way mimicking a driving video's facial movements. In this context, Generative Adversarial Networks have demonstrated remarkable potential in…

计算机视觉与模式识别 · 计算机科学 2024-08-26 Alireza Javanmardi , Alain Pagani , Didier Stricker

Audio-driven lip sync has recently drawn significant attention due to its widespread application in the multimedia domain. Individuals exhibit distinct lip shapes when speaking the same utterance, attributed to the unique speaking styles of…

计算机视觉与模式识别 · 计算机科学 2025-06-19 Weizhi Zhong , Jichang Li , Yinqi Cai , Ming Li , Feng Gao , Liang Lin , Guanbin Li

Full-stack multimodal interaction in real-time is a central goal in building intelligent embodied agents capable of natural, dynamic communication. However, existing systems are either limited to unimodal generation or suffer from degraded…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Xiang Deng , Feng Gao , Yong Zhang , Youxin Pang , Xu Xiaoming , Zhuoliang Kang , Xiaoming Wei , Yebin Liu

Despite exhibiting impressive performance in synthesizing lifelike personalized 3D talking heads, prevailing methods based on radiance fields suffer from high demands for training data and time for each new identity. This paper introduces…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Jiahe Li , Jiawei Zhang , Xiao Bai , Jin Zheng , Jun Zhou , Lin Gu

While recent advances in deep neural networks have made it possible to render high-quality images, generating photo-realistic and personalized talking head remains challenging. With given audio, the key to tackling this task is…

计算机视觉与模式识别 · 计算机科学 2022-01-04 Shunyu Yao , RuiZhe Zhong , Yichao Yan , Guangtao Zhai , Xiaokang Yang

Generative adversarial networks have seen rapid development in recent years and have led to remarkable improvements in generative modelling of images. However, their application in the audio domain has received limited attention, and…

Despite progress in speech-to-video synthesis, existing methods often struggle to capture cross-individual dependencies and provide fine-grained control over reactive behaviors in dyadic settings. To address these challenges, we propose…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Dongwei Pan , Longwei Guo , Jiazhi Guan , Luying Huang , Yiding Li , Haojie Liu , Haocheng Feng , Wei He , Kaisiyuan Wang , Hang Zhou

Audio-driven talking face generation has received growing interest, particularly for applications requiring expressive and natural human-avatar interaction. However, most existing emotion-aware methods rely on a single modality (either…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Phyo Thet Yee , Dimitrios Kollias , Sudeepta Mishra , Abhinav Dhall

In recent years, the increasing demand for dynamic 3D assets in design and gaming applications has given rise to powerful generative pipelines capable of synthesizing high-quality 4D objects. Previous methods generally rely on score…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Qi Sun , Zhiyang Guo , Ziyu Wan , Jing Nathan Yan , Shengming Yin , Wengang Zhou , Jing Liao , Houqiang Li

The synthesis of 3D facial animations from speech has garnered considerable attention. Due to the scarcity of high-quality 4D facial data and well-annotated abundant multi-modality labels, previous methods often suffer from limited realism…

计算机视觉与模式识别 · 计算机科学 2024-01-31 Qingcheng Zhao , Pengyu Long , Qixuan Zhang , Dafei Qin , Han Liang , Longwen Zhang , Yingliang Zhang , Jingyi Yu , Lan Xu

In the domain of photorealistic avatar generation, the fidelity of audio-driven lip motion synthesis is essential for realistic virtual interactions. Existing methods face two key challenges: a lack of vivacity due to limited diversity in…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Deng Junli , Luo Yihao , Yang Xueting , Li Siyou , Wang Wei , Guo Jinyang , Shi Ping

Recent advances in audio-driven talking head generation have achieved impressive results in lip synchronization and emotional expression. However, they largely overlook the crucial task of facial attribute editing. This capability is…

计算机视觉与模式识别 · 计算机科学 2025-08-28 Guanwen Feng , Zhiyuan Ma , Yunan Li , Jiahao Yang , Junwei Jing , Qiguang Miao

The recent GPT-4 has demonstrated extraordinary multi-modal abilities, such as directly generating websites from handwritten text and identifying humorous elements within images. These features are rarely observed in previous…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Deyao Zhu , Jun Chen , Xiaoqian Shen , Xiang Li , Mohamed Elhoseiny

We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and eye blinks.…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Dongchan Min , Minyoung Song , Eunji Ko , Sung Ju Hwang

Over the last few decades, many aspects of human life have been enhanced with virtual domains, from the advent of digital assistants such as Amazon's Alexa and Apple's Siri to the latest metaverse efforts of the rebranded Meta. These trends…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Siddarth Ravichandran , Ondřej Texler , Dimitar Dinev , Hyun Jae Kang
‹ 上一页 1 8 9 10 下一页 ›