English
Related papers

Related papers: VisualTTS: TTS with Accurate Lip-Speech Synchroniz…

200 papers

In many applications, synchronizing audio with visuals is crucial, such as in creating graphic animations for films or games, translating movie audio into different languages, and developing metaverse applications. This review explores…

Talking face generation aims to create realistic videos with accurate lip synchronization and high visual quality, using given audio and reference video while preserving identity and visual characteristics. In this paper, we start by…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazim Kemal Ekenel , Alexander Waibel

In this work, we investigate the problem of lip-syncing a talking face video of an arbitrary identity to match a target speech segment. Current works excel at producing accurate lip movements on a static image or videos of specific people…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 K R Prajwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

Creating realistic and natural-sounding synthetic speech remains a big challenge for voice identities unseen during training. As there is growing interest in synthesizing voices of new speakers, here we investigate the ability of…

Text-to-Speech (TTS) has recently seen great progress in synthesizing high-quality speech owing to the rapid development of parallel TTS systems, but producing speech with naturalistic prosodic variations, speaking styles and emotional…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-21 Yinghao Aaron Li , Cong Han , Nima Mesgarani

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

Computation and Language · Computer Science 2025-04-11 Lakshmipathi Balaji , Karan Singla

Visual Speech Recognition (VSR) differs from the common perception tasks as it requires deeper reasoning over the video sequence, even by human experts. Despite the recent advances in VSR, current approaches rely on labeled data to fully…

Several recent end-to-end text-to-speech (TTS) models enabling single-stage training and parallel sampling have been proposed, but their sample quality does not match that of two-stage TTS systems. In this work, we present a parallel…

Sound · Computer Science 2021-06-14 Jaehyeon Kim , Jungil Kong , Juhee Son

Speech-to-text translation (ST), which translates source language speech into target language text, has attracted intensive attention in recent years. Compared to the traditional pipeline system, the end-to-end ST model has potential…

Computation and Language · Computer Science 2019-12-17 Yuchen Liu , Jiajun Zhang , Hao Xiong , Long Zhou , Zhongjun He , Hua Wu , Haifeng Wang , Chengqing Zong

Audio is the primary modality for human communication and has driven the success of Automatic Speech Recognition (ASR) technologies. However, such audio-centric systems inherently exclude individuals who are deaf or hard of hearing. Visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Jeong Hun Yeo , Hyeongseop Rha , Sungjune Park , Junil Won , Yong Man Ro

Text-to-speech synthesis (TTS) has witnessed rapid progress in recent years, where neural methods became capable of producing audios with high naturalness. However, these efforts still suffer from two types of latencies: (a) the {\em…

Computation and Language · Computer Science 2020-10-08 Mingbo Ma , Baigong Zheng , Kaibo Liu , Renjie Zheng , Hairong Liu , Kainan Peng , Kenneth Church , Liang Huang

Text-to-speech synthesis (TTS) is a task to convert texts into speech. Two of the factors that have been driving TTS are the advancements of probabilistic models and latent representation learning. We propose a TTS method based on latent…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-19 Yusuke Yasuda , Tomoki Toda

Direct speech-to-speech translation (S2ST) aims to convert speech from one language into another, and has demonstrated significant progress to date. Despite the recent success, current S2ST models still suffer from distinct degradation in…

Computation and Language · Computer Science 2023-05-25 Rongjie Huang , Huadai Liu , Xize Cheng , Yi Ren , Linjun Li , Zhenhui Ye , Jinzheng He , Lichao Zhang , Jinglin Liu , Xiang Yin , Zhou Zhao

In this work, we address the problem of generating speech from silent lip videos for any speaker in the wild. In stark contrast to previous works, our method (i) is not restricted to a fixed number of speakers, (ii) does not explicitly…

Computer Vision and Pattern Recognition · Computer Science 2022-09-02 Sindhu B Hegde , K R Prajwal , Rudrabha Mukhopadhyay , Vinay P Namboodiri , C. V. Jawahar

While accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Hang Zhou , Yasheng Sun , Wayne Wu , Chen Change Loy , Xiaogang Wang , Ziwei Liu

We present ObamaNet, the first architecture that generates both audio and synchronized photo-realistic lip-sync videos from any new text. Contrary to other published lip-sync approaches, ours is only composed of fully trainable neural…

Computer Vision and Pattern Recognition · Computer Science 2018-01-08 Rithesh Kumar , Jose Sotelo , Kundan Kumar , Alexandre de Brebisson , Yoshua Bengio

This paper introduces Stereo-Talker, a novel one-shot audio-driven human video synthesis system that generates 3D talking videos with precise lip synchronization, expressive body gestures, temporally consistent photo-realistic quality, and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Xiang Deng , Youxin Pang , Xiaochen Zhao , Chao Xu , Lizhen Wang , Hongjiang Xiao , Shi Yan , Hongwen Zhang , Yebin Liu

Neural Text-to-speech (TTS) synthesis is a powerful technology that can generate speech using neural networks. One of the most remarkable features of TTS synthesis is its capability to produce speech in the voice of different speakers. This…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-19 Vinotha R , Hepsiba D , L. D. Vijay Anand , Deepak John Reji

Lip sync has emerged as a promising technique for generating mouth movements from audio signals. However, synthesizing a high-resolution and photorealistic virtual news anchor is still challenging. Lack of natural appearance, visual…

Computer Vision and Pattern Recognition · Computer Science 2021-05-06 Ruobing Zheng , Zhou Zhu , Bo Song , Changjiang Ji

Text-to-audio (TTA) generation is a recent popular problem that aims to synthesize general audio given text descriptions. Previous methods utilized latent diffusion models to learn audio embedding in a latent space with text embedding as…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Shentong Mo , Jing Shi , Yapeng Tian
‹ Prev 1 3 4 5 6 7 10 Next ›