English
Related papers

Related papers: Text-driven Talking Face Synthesis by Reprogrammin…

200 papers

This paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker embeddings to enhance the performance of the phrasing model. We…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Dong Yang , Yuki Saito , Takaaki Saeki , Tomoki Koriyama , Wataru Nakata , Detai Xin , Hiroshi Saruwatari

Recent work on speech representation models jointly pre-trained with text has demonstrated the potential of improving speech representations by encoding speech and text in a shared space. In this paper, we leverage such shared…

Computation and Language · Computer Science 2023-10-10 Chung-Ming Chien , Mingjiamei Zhang , Ju-Chieh Chou , Karen Livescu

Text-to-audio (TTA) system has recently gained attention for its ability to synthesize general audio based on text descriptions. However, previous studies in TTA have limited generation quality with high computational costs. In this study,…

Sound · Computer Science 2023-09-12 Haohe Liu , Zehua Chen , Yi Yuan , Xinhao Mei , Xubo Liu , Danilo Mandic , Wenwu Wang , Mark D. Plumbley

Editing talking-head video to change the speech content or to remove filler words is challenging. We propose a novel method to edit talking-head video based on its transcript to produce a realistic output video in which the dialogue of the…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Ohad Fried , Ayush Tewari , Michael Zollhöfer , Adam Finkelstein , Eli Shechtman , Dan B Goldman , Kyle Genova , Zeyu Jin , Christian Theobalt , Maneesh Agrawala

Some recent studies have demonstrated the feasibility of single-stage neural text-to-speech, which does not need to generate mel-spectrograms but generates the raw waveforms directly from the text. Single-stage text-to-speech often faces…

Sound · Computer Science 2022-07-14 Zhengxi Liu , Qiao Tian , Chenxu Hu , Xudong Liu , Menglin Wu , Yuping Wang , Hang Zhao , Yuxuan Wang

Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in situations where audio signals are unavailable or degraded.…

Sound · Computer Science 2026-02-03 Jaejun Lee , Yoori Oh , Kyogu Lee

We present Face0, a novel way to instantaneously condition a text-to-image generation model on a face, in sample time, without any optimization procedures such as fine-tuning or inversions. We augment a dataset of annotated images with…

Computer Vision and Pattern Recognition · Computer Science 2023-06-13 Dani Valevski , Danny Wasserman , Yossi Matias , Yaniv Leviathan

Speech-Preserving Facial Expression Manipulation (SPFEM) is an innovative technique aimed at altering facial expressions in images and videos while retaining the original mouth movements. Despite advancements, SPFEM still struggles with…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Zhenxuan Lu , Zhihua Xu , Zhijing Yang , Feng Gao , Yongyi Lu , Keze Wang , Tianshui Chen

With the popularity of deep neural network, speech synthesis task has achieved significant improvements based on the end-to-end encoder-decoder framework in the recent days. More and more applications relying on speech synthesis technology…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-23 Dongyang Dai , Li Chen , Yuping Wang , Mu Wang , Rui Xia , Xuchen Song , Zhiyong Wu , Yuxuan Wang

In this paper we present the first model for directly synthesizing fluent, natural-sounding spoken audio captions for images that does not require natural language text as an intermediate representation or source of supervision. Instead, we…

Computation and Language · Computer Science 2021-01-01 Wei-Ning Hsu , David Harwath , Christopher Song , James Glass

Target speech separation refers to isolating target speech from a multi-speaker mixture signal by conditioning on auxiliary information about the target speaker. Different from the mainstream audio-visual approaches which usually require…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Leyuan Qu , Cornelius Weber , Stefan Wermter

Speech synthesis has significantly advanced from statistical methods to deep neural network architectures, leading to various text-to-speech (TTS) models that closely mimic human speech patterns. However, capturing nuances such as emotion…

Sound · Computer Science 2025-01-14 Shaozuo Zhang , Ambuj Mehrish , Yingting Li , Soujanya Poria

Speech-driven 3D facial animation has been widely studied, yet there is still a gap to achieving realism and vividness due to the highly ill-posed nature and scarcity of audio-visual data. Existing works typically formulate the cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Jinbo Xing , Menghan Xia , Yuechen Zhang , Xiaodong Cun , Jue Wang , Tien-Tsin Wong

We present a novel audio-driven facial animation approach that can generate realistic lip-synchronized 3D facial animations from the input audio. Our approach learns viseme dynamics from speech videos, produces animator-friendly viseme…

Graphics · Computer Science 2023-01-18 Linchao Bao , Haoxian Zhang , Yue Qian , Tangli Xue , Changhai Chen , Xuefei Zhe , Di Kang

Text-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts describing the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Yuming Jiang , Shuai Yang , Tong Liang Koh , Wayne Wu , Chen Change Loy , Ziwei Liu

We introduce a text-to-speech(TTS) framework based on a neural transducer. We use discretized semantic tokens acquired from wav2vec2.0 embeddings, which makes it easy to adopt a neural transducer for the TTS framework enjoying its monotonic…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-09 Minchan Kim , Myeonghun Jeong , Byoung Jin Choi , Dongjune Lee , Nam Soo Kim

Combining face swapping with lip synchronization technology offers a cost-effective solution for customized talking face generation. However, directly cascading existing models together tends to introduce significant interference between…

Computer Vision and Pattern Recognition · Computer Science 2024-05-10 Zeren Zhang , Haibo Qin , Jiayu Huang , Yixin Li , Hui Lin , Yitao Duan , Jinwen Ma

Speech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input…

Computer Vision and Pattern Recognition · Computer Science 2023-01-03 Balamurugan Thambiraja , Ikhsanul Habibie , Sadegh Aliakbarian , Darren Cosker , Christian Theobalt , Justus Thies

Text-based talking-head video editing aims to efficiently insert, delete, and substitute segments of talking videos through a user-friendly text editing approach. It is challenging because of \textbf{1)} generalizable talking-face…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Bo Han , Heqing Zou , Haoyang Li , Guangcong Wang , Chng Eng Siong

Text-to-speech conversion has traditionally been performed either by concatenating short samples of speech or by using rule-based systems to convert a phonetic representation of speech into an acoustic representation, which is then…

Neural and Evolutionary Computing · Computer Science 2007-05-23 Orhan Karaali , Gerald Corrigan , Ira Gerson