中文
相关论文

相关论文: LatentSync: Taming Audio-Conditioned Latent Diffus…

200 篇论文

Generating consecutive images of lip movements that align with a given speech in audio-driven lip synthesis is a challenging task. While previous studies have made strides in synchronization and visual quality, lip intelligibility and video…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Shiyan Liu , Rui Qu , Yan Jin

Instead of performing text-conditioned denoising in the image domain, latent diffusion models (LDMs) operate in latent space of a variational autoencoder (VAE), enabling more efficient processing at reduced computational costs. However,…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Jason Becker , Chris Wendler , Peter Baylies , Robert West , Christian Wressnegger

Lip sync is a fundamental audio-visual task. However, existing lip sync methods fall short of being robust in the wild. One important cause could be distracting factors on the visual input side, making extracting lip motion information…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Chun Wang

Many speech segments in movies are re-recorded in a studio during postproduction, to compensate for poor sound quality as recorded on location. Manual alignment of the newly-recorded speech with the original lip movements is a tedious task.…

计算机视觉与模式识别 · 计算机科学 2018-08-21 Tavi Halperin , Ariel Ephrat , Shmuel Peleg

Lip sync has emerged as a promising technique for generating mouth movements from audio signals. However, synthesizing a high-resolution and photorealistic virtual news anchor is still challenging. Lack of natural appearance, visual…

计算机视觉与模式识别 · 计算机科学 2021-05-06 Ruobing Zheng , Zhou Zhu , Bo Song , Changjiang Ji

Generating synchronized and natural lip movement with speech is one of the most important tasks in creating realistic virtual characters. In this paper, we present a combined deep neural network of one-dimensional convolutions and LSTM to…

声音 · 计算机科学 2022-05-03 Xiaohong Li , Xiang Wang , Kai Wang , Shiguo Lian

Emotionally talking head video generation aims to generate expressive portrait videos with accurate lip synchronization and emotional facial expressions. Current methods rely on simple emotional labels, leading to insufficient semantic…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Yahui Li , Yinfeng Yu , Liejun Wang , Shengjie Shen

Speech-to-face generation is an intriguing area of research that focuses on generating realistic facial images based on a speaker's audio speech. However, state-of-the-art methods employing GAN-based architectures lack stability and cannot…

计算机视觉与模式识别 · 计算机科学 2023-10-06 Jinting Wang , Li Liu , Jun Wang , Hei Victor Cheng

Researchers have shown a growing interest in Audio-driven Talking Head Generation. The primary challenge in talking head generation is achieving audio-visual coherence between the lips and the audio, known as lip synchronization. This paper…

声音 · 计算机科学 2026-02-03 Zhipeng Chen , Xinheng Wang , Lun Xie , Haijie Yuan , Hang Pan

The goal of this paper is to develop state-of-the-art models for lip reading -- visual speech recognition. We develop three architectures and compare their accuracy and training times: (i) a recurrent model using LSTMs; (ii) a fully…

计算机视觉与模式识别 · 计算机科学 2018-06-18 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Contrastively pretrained audio-language models (e.g., CLAP) excel at clip-level understanding but struggle with frame-level tasks. Existing extensions fail to exploit the varying granularity of real-world audio-text data, where massive…

声音 · 计算机科学 2026-04-02 Xiquan Li , Xuenan Xu , Ziyang Ma , Wenxi Chen , Haolin He , Qiuqiang Kong , Xie Chen

Movie dubbing aims to synthesize speech that preserves the vocal identity of a reference audio while synchronizing with the lip movements in a target video. Existing methods fail to achieve precise lip-sync and lack naturalness due to…

声音 · 计算机科学 2026-04-15 Gaoxiang Cong , Liang Li , Jiaxin Ye , Zhedong Zhang , Hongming Shan , Yuankai Qi , Qingming Huang

The goal of this paper is to learn strong lip reading models that can recognise speech in silent videos. Most prior works deal with the open-set visual speech recognition problem by adapting existing automatic speech recognition techniques…

计算机视觉与模式识别 · 计算机科学 2021-12-06 K R Prajwal , Triantafyllos Afouras , Andrew Zisserman

Synthesizing realistic videos according to a given speech is still an open challenge. Previous works have been plagued by issues such as inaccurate lip shape generation and poor image quality. The key reason is that only motions and…

计算机视觉与模式识别 · 计算机科学 2023-09-12 Xiuzhe Wu , Pengfei Hu , Yang Wu , Xiaoyang Lyu , Yan-Pei Cao , Ying Shan , Wenming Yang , Zhongqian Sun , Xiaojuan Qi

The Video-to-Audio (V2A) model has recently gained attention for its practical application in generating audio directly from silent videos, particularly in video/film production. However, previous methods in V2A have limited generation…

声音 · 计算机科学 2023-07-03 Simian Luo , Chuanhao Yan , Chenxu Hu , Hang Zhao

Combining face swapping with lip synchronization technology offers a cost-effective solution for customized talking face generation. However, directly cascading existing models together tends to introduce significant interference between…

计算机视觉与模式识别 · 计算机科学 2024-05-10 Zeren Zhang , Haibo Qin , Jiayu Huang , Yixin Li , Hui Lin , Yitao Duan , Jinwen Ma

Current text-to-speech (TTS) models face a persistent limitation: autoregressive (AR) models suffer from low generation efficiency, while modern non-autoregressive (NAR) models experience high latency due to their unordered temporal nature.…

声音 · 计算机科学 2026-03-17 Zhengyan Sheng , Zhihao Du , Shiliang Zhang , Zhijie Yan , Liping Chen

The word-level lipreading approach typically employs a two-stage framework with separate frontend and backend architectures to model dynamic lip movements. Each component has been extensively studied, and in the backend architecture,…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Byung Hoon Lee , Wooseok Shin , Sung Won Han

Latent diffusion models offer an attractive alternative to discrete diffusion for non-autoregressive text generation by operating on continuous text representations and denoising entire sequences in parallel. The major challenge in latent…

Animating still face images with deep generative models using a speech input signal is an active research topic and has seen important recent progress.However, much of the effort has been put into lip syncing and rendering quality while the…

图形学 · 计算机科学 2024-12-13 Louis Airale , Dominique Vaufreydaz , Xavier Alameda-Pineda