中文
相关论文

相关论文: FluentLip: A Phonemes-Based Two-stage Approach for…

200 篇论文

Real-time video dubbing that preserves identity consistency while achieving accurate lip synchronization remains a critical challenge. Existing approaches face a trilemma: diffusion-based methods achieve high visual fidelity but suffer from…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Yue Zhang , Zhizhou Zhong , Minhao Liu , Zhaokang Chen , Bin Wu , Yubin Zeng , Chao Zhan , Yingjie He , Junxin Huang , Wenjiang Zhou

Most lip-to-speech (LTS) synthesis models are trained and evaluated under the assumption that the audio-video pairs in the dataset are perfectly synchronized. In this work, we show that the commonly used audio-visual datasets, such as GRID,…

声音 · 计算机科学 2023-03-02 Zhe Niu , Brian Mak

Video editing-based talking face generation aims to preserve video details such as pose, lighting, and gestures while modifying only lip motion, often using an identity reference image to maintain speaker consistency. However, this…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Dogucan Yaman , Fevziye Irem Eyiokur , Hazım Kemal Ekenel , Alexander Waibel

Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human communication, co-speech gestures are also temporally aligned…

音频与语音处理 · 电气工程与系统科学 2026-01-28 Zexu Pan , Xinyuan Qian , Shengkui Zhao , Kun Zhou , Bin Ma

Despite recent advances in syncing lip movements with any audio waves, current methods still struggle to balance generation quality and the model's generalization ability. Previous studies either require long-term data for training or…

计算机视觉与模式识别 · 计算机科学 2023-05-10 Jiazhi Guan , Zhanwang Zhang , Hang Zhou , Tianshu Hu , Kaisiyuan Wang , Dongliang He , Haocheng Feng , Jingtuo Liu , Errui Ding , Ziwei Liu , Jingdong Wang

The task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current approaches, they commonly require high-quality homologous data…

计算机视觉与模式识别 · 计算机科学 2022-01-19 Tianyi Xie , Liucheng Liao , Cheng Bi , Benlai Tang , Xiang Yin , Jianfei Yang , Mingjie Wang , Jiali Yao , Yang Zhang , Zejun Ma

Lip sync has emerged as a promising technique for generating mouth movements from audio signals. However, synthesizing a high-resolution and photorealistic virtual news anchor is still challenging. Lack of natural appearance, visual…

计算机视觉与模式识别 · 计算机科学 2021-05-06 Ruobing Zheng , Zhou Zhu , Bo Song , Changjiang Ji

As a key component of talking face generation, lip movements generation determines the naturalness and coherence of the generated talking face video. Prior literature mainly focuses on speech-to-lip generation while there is a paucity in…

多媒体 · 计算机科学 2021-12-21 Jinglin Liu , Zhiying Zhu , Yi Ren , Wencan Huang , Baoxing Huai , Nicholas Yuan , Zhou Zhao

In this work, we revisit the effectiveness of 3DMM for talking head synthesis by jointly learning a 3D face reconstruction model and a talking head synthesis model. This enables us to obtain a FACS-based blendshape representation of facial…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Sungjoon Park , Minsik Park , Haneol Lee , Jaesub Yun , Donggeon Lee

In this paper, we introduce a novel approach to address the task of synthesizing speech from silent videos of any in-the-wild speaker solely based on lip movements. The traditional approach of directly generating speech from lip videos…

多媒体 · 计算机科学 2024-03-05 Sindhu Hegde , Rudrabha Mukhopadhyay , C. V. Jawahar , Vinay Namboodiri

Lip reading, also known as visual speech recognition, aims to recognize the speech content from videos by analyzing the lip dynamics. There have been several appealing progress in recent years, benefiting much from the rapidly developed…

计算机视觉与模式识别 · 计算机科学 2020-11-17 Dalu Feng , Shuang Yang , Shiguang Shan , Xilin Chen

Language-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches, such as time-frequency masking, to separate target sounds…

声音 · 计算机科学 2025-01-10 Yi Yuan , Xubo Liu , Haohe Liu , Mark D. Plumbley , Wenwu Wang

Talking face generation with great practical significance has attracted more attention in recent audio-visual studies. How to achieve accurate lip synchronization is a long-standing challenge to be further investigated. Motivated by xxx, in…

计算机视觉与模式识别 · 计算机科学 2022-04-12 Ganglai Wang , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha

Video dubbing requires content accuracy, expressive prosody, high-quality acoustics, and precise lip synchronization, yet existing approaches struggle on all four fronts. To address these issues, we propose DiFlowDubber, the first video…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Ngoc-Son Nguyen , Thanh V. T. Tran , Jeongsoo Choi , Hieu-Nghia Huynh-Nguyen , Truong-Son Hy , Van Nguyen

In this work, we re-think the task of speech enhancement in unconstrained real-world environments. Current state-of-the-art methods use only the audio stream and are limited in their performance in a wide range of real-world noises. Recent…

计算机视觉与模式识别 · 计算机科学 2020-12-22 Sindhu B Hegde , K R Prajwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C. V. Jawahar

Visual speaker recognition based on lip motion offers a silent, hands-free, and behavior-driven biometric solution that remains effective even when acoustic cues are unavailable. Compared to traditional methods that rely heavily on…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Junguang Yao , Wenye Liu , Stjepan Picek , Yue Zheng

Given a piece of text, a video clip, and a reference audio, the movie dubbing task aims to generate speech that aligns with the video while cloning the desired voice. The existing methods have two primary deficiencies: (1) They struggle to…

Direct speech-to-speech translation achieves high-quality results through the introduction of discrete units obtained from self-supervised learning. This approach circumvents delays and cascading errors associated with model cascading.…

We aim to edit the lip movements in talking video according to the given speech while preserving the personal identity and visual details. The task can be decomposed into two sub-problems: (1) speech-driven lip motion generation and (2)…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Runyi Yu , Tianyu He , Ailing Zhang , Yuchi Wang , Junliang Guo , Xu Tan , Chang Liu , Jie Chen , Jiang Bian

Achieving high-fidelity lip-speech synchronization in audio-driven talking portrait synthesis remains challenging. While multi-stage pipelines or diffusion models yield high-quality results, they suffer from high computational costs. Some…

计算机视觉与模式识别 · 计算机科学 2025-04-24 Ziqi Ni , Ao Fu , Yi Zhou