English
Related papers

Related papers: Audiovisual Speech Synthesis using Tacotron2

200 papers

Despite previous success in generating audio-driven talking heads, most of the previous studies focus on the correlation between speech content and the mouth shape. Facial emotion, which is one of the most important features on natural…

Computer Vision and Pattern Recognition · Computer Science 2021-05-21 Xinya Ji , Hang Zhou , Kaisiyuan Wang , Wayne Wu , Chen Change Loy , Xun Cao , Feng Xu

Recent advances in speech synthesis suggest that limitations such as the lossy nature of the amplitude spectrum with minimum phase approximation and the over-smoothing effect in acoustic modeling can be overcome by using advanced machine…

Audio and Speech Processing · Electrical Eng. & Systems 2018-04-10 Xin Wang , Jaime Lorenzo-Trueba , Shinji Takaki , Lauri Juvela , Junichi Yamagishi

In this paper, we propose a talking face generation method that takes an audio signal as input and a short target video clip as reference, and synthesizes a photo-realistic video of the target face with natural lip motions, head poses, and…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Chenxu Zhang , Yifan Zhao , Yifei Huang , Ming Zeng , Saifeng Ni , Madhukar Budagavi , Xiaohu Guo

In this work, we propose a joint system combining a talking face generation system with a text-to-speech system that can generate multilingual talking face videos from only the text input. Our system can synthesize natural multilingual…

Computer Vision and Pattern Recognition · Computer Science 2022-05-16 Hyoung-Kyu Song , Sang Hoon Woo , Junhyeok Lee , Seungmin Yang , Hyunjae Cho , Youseong Lee , Dongho Choi , Kang-wook Kim

People talk with diversified styles. For one piece of speech, different talking styles exhibit significant differences in the facial and head pose movements. For example, the "excited" style usually talks with the mouth wide open, while the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Haozhe Wu , Jia Jia , Haoyu Wang , Yishun Dou , Chao Duan , Qingshan Deng

Audiovisual active speaker detection (ASD) is conventionally performed by modelling the temporal synchronisation of acoustic and visual speech cues. In egocentric recordings, however, the efficacy of synchronisation-based methods is…

Multimedia · Computer Science 2025-06-24 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

One way of expressing an environmental sound is using vocal imitations, which involve the process of replicating or mimicking the rhythm and pitch of sounds by voice. We can effectively express the features of environmental sounds, such as…

Speech enhancement can potentially benefit from the visual information from the target speaker, such as lip movement and facial expressions, because the visual aspect of speech is essentially unaffected by acoustic environment. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-24 Xinmeng Xu , Jianjun Hao

Audio-visual speech recognition (AVSR) system is thought to be one of the most promising solutions for robust speech recognition, especially in noisy environment. In this paper, we propose a novel multimodal attention based method for…

Computation and Language · Computer Science 2019-04-24 Pan Zhou , Wenwen Yang , Wei Chen , Yanfeng Wang , Jia Jia

This paper investigates a novel task of talking face video generation solely from speeches. The speech-to-video generation technique can spark interesting applications in entertainment, customer service, and human-computer-interaction…

Sound · Computer Science 2021-07-15 Shijing Si , Jianzong Wang , Xiaoyang Qu , Ning Cheng , Wenqi Wei , Xinghua Zhu , Jing Xiao

Audio and visual modalities are inherently connected in speech signals: lip movements and facial expressions are correlated with speech sounds. This motivates studies that incorporate the visual modality to enhance an acoustic speech signal…

Sound · Computer Science 2023-06-02 Juan F. Montesinos , Daniel Michelsanti , Gloria Haro , Zheng-Hua Tan , Jesper Jensen

Synthesizing realistic videos according to a given speech is still an open challenge. Previous works have been plagued by issues such as inaccurate lip shape generation and poor image quality. The key reason is that only motions and…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Xiuzhe Wu , Pengfei Hu , Yang Wu , Xiaoyang Lyu , Yan-Pei Cao , Ying Shan , Wenming Yang , Zhongqian Sun , Xiaojuan Qi

In this work, we present an end-to-end binaural speech synthesis system that combines a low-bitrate audio codec with a powerful binaural decoder that is capable of accurate speech binauralization while faithfully reconstructing…

Sound · Computer Science 2022-07-11 Wen Chin Huang , Dejan Markovic , Alexander Richard , Israel Dejene Gebru , Anjali Menon

Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However, in a more realistic setting, when multiple faces are…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-12 Otavio Braga , Takaki Makino , Olivier Siohan , Hank Liao

State-of-the-art Text-to-Video (T2V) diffusion models can generate visually impressive results, yet they still frequently fail to compose complex scenes or follow logical temporal instructions. In this paper, we argue that many errors,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Mariam Hassan , Bastien Van Delft , Wuyang Li , Alexandre Alahi

Recent deep learning Text-to-Speech (TTS) systems have achieved impressive performance by generating speech close to human parity. However, they suffer from training stability issues as well as incorrect alignment of the intermediate…

Visual emotion expression plays an important role in audiovisual speech communication. In this work, we propose a novel approach to rendering visual emotion expression in speech-driven talking face generation. Specifically, we design an…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-23 Sefik Emre Eskimez , You Zhang , Zhiyao Duan

Due to the increasing demand in films and games, synthesizing 3D avatar animation has attracted much attention recently. In this work, we present a production-ready text/speech-driven full-body animation synthesis system. Given the text and…

Graphics · Computer Science 2022-06-01 Wenlin Zhuang , Jinwei Qi , Peng Zhang , Bang Zhang , Ping Tan

Person-generic audio-driven face generation is a challenging task in computer vision. Previous methods have achieved remarkable progress in audio-visual synchronization, but there is still a significant gap between current results and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Xiaozhong Ji , Chuming Lin , Zhonggan Ding , Ying Tai , Junwei Zhu , Xiaobin Hu , Donghao Luo , Yanhao Ge , Chengjie Wang

Speech technology is becoming ever more ubiquitous with the advance of speech enabled devices and services. The use of speech synthesis in Augmentative and Alternative Communication tools, has facilitated inclusion of individuals with…