English
Related papers

Related papers: VoiceCraft-Dub: Automated Video Dubbing with Neura…

200 papers

This work presents CLIPDraw, an algorithm that synthesizes novel drawings based on natural language input. CLIPDraw does not require any training; rather a pre-trained CLIP language-image encoder is used as a metric for maximizing…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Kevin Frans , L. B. Soros , Olaf Witkowski

Current movie dubbing technology can produce the desired speech using a reference voice and input video, maintaining perfect synchronization with the visuals while effectively conveying the intended emotions. However, crucial aspects of…

Multimedia · Computer Science 2025-05-23 Junjie Zheng , Zihao Chen , Chaofan Ding , Yunming Liang , Yihan Fan , Huan Yang , Lei Xie , Xinhan Di

Existing automated dubbing methods are usually designed for Professionally Generated Content (PGC) production, which requires massive training data and training time to learn a person-specific audio-video mapping. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2023-09-04 Linsen Song , Wayne Wu , Chaoyou Fu , Chen Change Loy , Ran He

We describe a system for large-scale audiovisual translation and dubbing, which translates videos from one language to another. The source language's speech content is transcribed to text, translated, and automatically synthesized into…

Computer Vision and Pattern Recognition · Computer Science 2020-11-09 Yi Yang , Brendan Shillingford , Yannis Assael , Miaosen Wang , Wendi Liu , Yutian Chen , Yu Zhang , Eren Sezener , Luis C. Cobo , Misha Denil , Yusuf Aytar , Nando de Freitas

Since facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate than their audio-only counterparts. Yet, state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2022-04-01 Karren Yang , Dejan Markovic , Steven Krenn , Vasu Agrawal , Alexander Richard

Dubbing is a type of audiovisual translation where dialogues are translated and enacted so that they give the impression that the media is in the target language. It requires a careful alignment of dubbed recordings with the lip movements…

Computation and Language · Computer Science 2019-08-21 Alp Öktem , Mireia Farrús , Antonio Bonafonte

The task of Visual Text-to-Speech (VisualTTS), also known as video dubbing, aims to generate speech synchronized with the lip movements in an input video, in additional to being consistent with the content of input text and cloning the…

Multimedia · Computer Science 2025-12-01 Yuyue Wang , Xin Cheng , Yihan Wu , Xihua Wang , Jinchuan Tian , Ruihua Song

The visual dubbing task aims to generate mouth movements synchronized with the driving audio, which has seen significant progress in recent years. However, two critical deficiencies hinder their wide application: (1) Audio-only driving…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Liyang Chen , Tianze Zhou , Xu He , Boshi Tang , Zhiyong Wu , Yang Huang , Yang Wu , Zhongqian Sun , Wei Yang , Helen Meng

The task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current approaches, they commonly require high-quality homologous data…

Computer Vision and Pattern Recognition · Computer Science 2022-01-19 Tianyi Xie , Liucheng Liao , Cheng Bi , Benlai Tang , Xiang Yin , Jianfei Yang , Mingjie Wang , Jiali Yao , Yang Zhang , Zejun Ma

We propose EnCLAP, a novel framework for automated audio captioning. EnCLAP employs two acoustic representation models, EnCodec and CLAP, along with a pretrained language model, BART. We also introduce a new training objective called masked…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-01 Jaeyeon Kim , Jaeyoon Jung , Jinjoo Lee , Sang Hoon Woo

Movie dubbing aims to synthesize speech that preserves the vocal identity of a reference audio while synchronizing with the lip movements in a target video. Existing methods fail to achieve precise lip-sync and lack naturalness due to…

Sound · Computer Science 2026-04-15 Gaoxiang Cong , Liang Li , Jiaxin Ye , Zhedong Zhang , Hongming Shan , Yuankai Qi , Qingming Huang

Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head…

Multimedia · Computer Science 2023-09-21 Songlin Yang , Wei Wang , Jun Ling , Bo Peng , Xu Tan , Jing Dong

Automatic speaker naming is the problem of localizing as well as identifying each speaking character in a TV/movie/live show video. This is a challenging problem mainly attributes to its multimodal nature, namely face cue alone is…

Computer Vision and Pattern Recognition · Computer Science 2015-07-20 Yongtao Hu , Jimmy Ren , Jingwen Dai , Chang Yuan , Li Xu , Wenping Wang

Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody expressiveness but overlooks two key issues: 1) Multiscale…

Multimedia · Computer Science 2025-01-03 Yuan Zhao , Rui Liu , Gaoxiang Cong

Large-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited research on learning video-text representations for general…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Xingjian He , Sihan Chen , Fan Ma , Zhicheng Huang , Xiaojie Jin , Zikang Liu , Dongmei Fu , Yi Yang , Jing Liu , Jiashi Feng

Recent breakthroughs in video AIGC have ushered in a transformative era for audio-driven human animation. However, conventional video dubbing techniques remain constrained to mouth region editing, resulting in discordant facial expressions…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Shaoshu Yang , Zhe Kong , Feng Gao , Meng Cheng , Xiangyu Liu , Yong Zhang , Zhuoliang Kang , Wenhan Luo , Xunliang Cai , Ran He , Xiaoming Wei

Movie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker's voice demonstrated in a short reference audio clip. This task demands…

Sound · Computer Science 2025-03-19 Zhedong Zhang , Liang Li , Chenggang Yan , Chunshan Liu , Anton van den Hengel , Yuankai Qi

The generation of audio-driven talking head videos is a key challenge in computer vision and graphics, with applications in virtual avatars and digital media. Traditional approaches often struggle with capturing the complex interaction…

Sound · Computer Science 2025-04-01 Ao Fu , Ziqi Ni , Yi Zhou

Video dubbing aims to translate the original speech in a film or television program into the speech in a target language, which can be achieved with a cascaded system consisting of speech recognition, machine translation and speech…

Computation and Language · Computer Science 2023-12-06 Yihan Wu , Junliang Guo , Xu Tan , Chen Zhang , Bohan Li , Ruihua Song , Lei He , Sheng Zhao , Arul Menezes , Jiang Bian

We present a method for generating a video of a talking face. The method takes as inputs: (i) still images of the target face, and (ii) an audio speech segment; and outputs a video of the target face lip synched with the audio. The method…

Computer Vision and Pattern Recognition · Computer Science 2017-07-19 Joon Son Chung , Amir Jamaludin , Andrew Zisserman