中文
相关论文

相关论文: SwapTalk: Audio-Driven Talking Face Generation wit…

200 篇论文

Talking head synthesis is a practical technique with wide applications. Current Neural Radiance Field (NeRF) based approaches have shown their superiority on driving one-shot talking heads with videos or signals regressed from audio.…

计算机视觉与模式识别 · 计算机科学 2024-08-20 Dongze Li , Kang Zhao , Wei Wang , Yifeng Ma , Bo Peng , Yingya Zhang , Jing Dong

Speech-driven 3D facial animation aims to generate realistic lip movements and facial expressions for 3D head models from arbitrary audio clips. Although existing diffusion-based methods are capable of producing natural motions, their slow…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Xuangeng Chu , Nabarun Goswami , Ziteng Cui , Hanqin Wang , Tatsuya Harada

Recent advances have demonstrated compelling capabilities in synthesizing real individuals into generated videos, reflecting the growing demand for identity-aware content creation. Nevertheless, an openly accessible framework enabling…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Yingjie Chen , Shilun Lin , Cai Xing , Binxin Yang , Long Zhou , Qixin Yan , Wenjing Wang , Dingming Liu , Hao Liu , Chen Li , Jing Lyu

While recent advances in deep neural networks have made it possible to render high-quality images, generating photo-realistic and personalized talking head remains challenging. With given audio, the key to tackling this task is…

计算机视觉与模式识别 · 计算机科学 2022-01-04 Shunyu Yao , RuiZhe Zhong , Yichao Yan , Guangtao Zhai , Xiaokang Yang

In this paper, we propose a neural end-to-end system for voice preserving, lip-synchronous translation of videos. The system is designed to combine multiple component models and produces a video of the original speaker speaking in the…

In this work, we propose a joint system combining a talking face generation system with a text-to-speech system that can generate multilingual talking face videos from only the text input. Our system can synthesize natural multilingual…

计算机视觉与模式识别 · 计算机科学 2022-05-16 Hyoung-Kyu Song , Sang Hoon Woo , Junhyeok Lee , Seungmin Yang , Hyunjae Cho , Youseong Lee , Dongho Choi , Kang-wook Kim

Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in situations where audio signals are unavailable or degraded.…

声音 · 计算机科学 2026-02-03 Jaejun Lee , Yoori Oh , Kyogu Lee

Speech-to-face generation is an intriguing area of research that focuses on generating realistic facial images based on a speaker's audio speech. However, state-of-the-art methods employing GAN-based architectures lack stability and cannot…

计算机视觉与模式识别 · 计算机科学 2023-10-06 Jinting Wang , Li Liu , Jun Wang , Hei Victor Cheng

Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consistency. To mitigate these drawbacks, we propose UniAVGen, a…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Guozhen Zhang , Zixiang Zhou , Teng Hu , Ziqiao Peng , Youliang Zhang , Yi Chen , Yuan Zhou , Qinglin Lu , Limin Wang

We introduce a novel method for joint expression and audio-guided talking face generation. Recent approaches either struggle to preserve the speaker identity or fail to produce faithful facial expressions. To address these challenges, we…

计算机视觉与模式识别 · 计算机科学 2024-09-19 Sai Tanmay Reddy Chakkera , Aggelina Chatziagapi , Dimitris Samaras

We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied avatar must achieve real-time generation, unified listening-speaking behavior, and high…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Yu Zhang , Kaiyuan Shen , Yang Li

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a…

计算机视觉与模式识别 · 计算机科学 2020-03-06 Ran Yi , Zipeng Ye , Juyong Zhang , Hujun Bao , Yong-Jin Liu

The task of video-to-speech aims to translate silent video of lip movement to its corresponding audio signal. Previous approaches to this task are generally limited to the case of a single speaker, but a method that accounts for multiple…

音频与语音处理 · 电气工程与系统科学 2021-05-21 Dan Oneata , Adriana Stan , Horia Cucu

Creating high-quality, generalizable speech-driven 3D talking heads remains a persistent challenge. Previous methods achieve satisfactory results for fixed viewpoints and small-scale audio variations, but they struggle with large head…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Wentao Hu , Shunkai Li , Ziqiao Peng , Haoxian Zhang , Fan Shi , Xiaoqiang Liu , Pengfei Wan , Di Zhang , Hui Tian

Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production. While existing methods accurately synchronize lip movements with input audio, they typically ignore…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Jiadong Liang , Feng Lu

We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often produce disembodied "talking heads" akin to a video conference…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Mengyi Shan , Shouchieh Chang , Ziqian Bai , Shichen Liu , Yinda Zhang , Luchuan Song , Rohit Pandey , Sean Fanello , Zeng Huang

Talking head synthesis has emerged as a prominent research topic in computer graphics and multimedia, yet most existing methods often struggle to strike a balance between generation quality and computational efficiency, particularly under…

图形学 · 计算机科学 2025-06-30 Shuai Shen , Wanhua Li , Yunpeng Zhang , Yap-Peng Tan , Jiwen Lu

Audio-driven talking head generation faces a fundamental trade-off between personalization and generalization, limiting its practical application. Implicit models often achieve generalization at the cost of structural incoherence, resulting…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Shiyu Liu , Kui Jiang , Junjun Jiang , Xianming Liu , Xiaocheng Feng , Hongxun Yao , Qi Tian

Video generation has achieved remarkable progress, with generated videos increasingly resembling real ones. However, the rapid advance in generation has outpaced the development of adequate evaluation metrics. Currently, the assessment of…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Nabyl Quignon , Baptiste Chopin , Yaohui Wang , Antitza Dantcheva

The goal of this work is to reconstruct speech from a silent talking face video. Recent studies have shown impressive performance on synthesizing speech from silent talking face videos. However, they have not explicitly considered on…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Joanna Hong , Minsu Kim , Yong Man Ro