English
Related papers

Related papers: Dynamic Temporal Alignment of Speech to Lips

200 papers

Talking face generation aims to create realistic videos with accurate lip synchronization and high visual quality, using given audio and reference video while preserving identity and visual characteristics. In this paper, we start by…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazim Kemal Ekenel , Alexander Waibel

Generating consecutive images of lip movements that align with a given speech in audio-driven lip synthesis is a challenging task. While previous studies have made strides in synchronization and visual quality, lip intelligibility and video…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Shiyan Liu , Rui Qu , Yan Jin

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-24 Hang Zhou , Yu Liu , Ziwei Liu , Ping Luo , Xiaogang Wang

Direct speech-to-speech translation achieves high-quality results through the introduction of discrete units obtained from self-supervised learning. This approach circumvents delays and cascading errors associated with model cascading.…

This paper introduces a parallel and asynchronous Transformer framework designed for efficient and accurate multilingual lip synchronization in real-time video conferencing systems. The proposed architecture integrates translation, speech…

Multimedia · Computer Science 2025-12-23 Eren Caglar , Amirkia Rafiei Oskooei , Mehmet Kutanoglu , Mustafa Keles , Mehmet S. Aktas

As a key component of talking face generation, lip movements generation determines the naturalness and coherence of the generated talking face video. Prior literature mainly focuses on speech-to-lip generation while there is a paucity in…

Multimedia · Computer Science 2021-12-21 Jinglin Liu , Zhiying Zhu , Yi Ren , Wencan Huang , Baoxing Huai , Nicholas Yuan , Zhou Zhao

Audio-visual alignment after dubbing is a challenging research problem. To this end, we propose a novel method, DubWise Multi-modal Large Language Model (LLM)-based Text-to-Speech (TTS), which can control the speech duration of synthesized…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-14 Neha Sahipjohn , Ashishkumar Gudmalwar , Nirmesh Shah , Pankaj Wasnik , Rajiv Ratn Shah

Acoustic matching aims to re-synthesize an audio clip to sound as if it were recorded in a target acoustic environment. Existing methods assume access to paired training data, where the audio is observed in both source and target…

Multimedia · Computer Science 2023-11-27 Arjun Somayazulu , Changan Chen , Kristen Grauman

Understanding the relationship between vocal tract motion during speech and the resulting acoustic signal is crucial for aided clinical assessment and developing personalized treatment and rehabilitation strategies. Toward this goal, we…

Large text-to-video models hold immense potential for a wide range of downstream applications. However, they struggle to accurately depict dynamic object interactions, often resulting in unrealistic movements and frequent violations of…

Machine Learning · Computer Science 2026-04-21 Hiroki Furuta , Heiga Zen , Dale Schuurmans , Aleksandra Faust , Yutaka Matsuo , Percy Liang , Sherry Yang

Talking face generation is the challenging task of synthesizing a natural and realistic face that requires accurate synchronization with a given audio. Due to co-articulation, where an isolated phone is influenced by the preceding or…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Se Jin Park , Minsu Kim , Jeongsoo Choi , Yong Man Ro

Learning to associate audio with textual descriptions is valuable for a range of tasks, including pretraining, zero-shot classification, audio retrieval, audio captioning, and text-conditioned audio generation. Existing contrastive…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-13 Paul Primus , Florian Schmid , Gerhard Widmer

Realistic lip synchronization is essential for the natural human-robot non-verbal interaction of humanoid robots. Motivated by this need, this paper presents a lip motion generation framework based on 3D dynamic viseme and coarticulation…

Robotics · Computer Science 2026-04-03 Sheng Li , Jingcheng Huang , Min Li

Generating semantically and temporally aligned audio content in accordance with video input has become a focal point for researchers, particularly following the remarkable breakthrough in text-to-video generation. In this work, we aim to…

Sound · Computer Science 2025-03-12 Manjie Xu , Chenxing Li , Xinyi Tu , Yong Ren , Rilin Chen , Yu Gu , Wei Liang , Dong Yu

Talking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and visual quality. Despite…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Jiadong Wang , Xinyuan Qian , Malu Zhang , Robby T. Tan , Haizhou Li

Unconstrained lip-to-speech synthesis aims to generate corresponding speeches from silent videos of talking faces with no restriction on head poses or vocabulary. Current works mainly use sequence-to-sequence models to solve this problem,…

Sound · Computer Science 2022-07-14 Yongqi Wang , Zhou Zhao

Creating realistic, natural, and lip-readable talking face videos remains a formidable challenge. Previous research primarily concentrated on generating and aligning single-frame images while overlooking the smoothness of frame-to-frame…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Shuheng Ge , Haoyu Xing , Li Zhang , Xiangqian Wu

Audio-driven talking face video generation has attracted increasing attention due to its huge industrial potential. Some previous methods focus on learning a direct mapping from audio to visual content. Despite progress, they often struggle…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Weizhi Zhong , Junfan Lin , Peixin Chen , Liang Lin , Guanbin Li

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Jiaxin Ye , Hongming Shan

Speech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input…

Computer Vision and Pattern Recognition · Computer Science 2023-01-03 Balamurugan Thambiraja , Ikhsanul Habibie , Sadegh Aliakbarian , Darren Cosker , Christian Theobalt , Justus Thies