English
Related papers

Related papers: Residual-guided Personalized Speech Synthesis base…

200 papers

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. In this paper, we present an automatic method to generate…

Computer Vision and Pattern Recognition · Computer Science 2021-08-29 Xinsheng Wang , Qicong Xie , Jihua Zhu , Lei Xie , Scharenborg

In this paper, we present a method for reprogramming pre-trained audio-driven talking face synthesis models to operate in a text-driven manner. Consequently, we can easily generate face videos that articulate the provided textual sentences,…

Graphics · Computer Science 2024-01-19 Jeongsoo Choi , Minsu Kim , Se Jin Park , Yong Man Ro

Speech-based image retrieval has been studied as a proxy for joint representation learning, usually without emphasis on retrieval itself. As such, it is unclear how well speech-based retrieval can work in practice -- both in an absolute…

Computation and Language · Computer Science 2021-06-16 Ramon Sanabria , Austin Waters , Jason Baldridge

With the advances in deep learning, speech enhancement systems benefited from large neural network architectures and achieved state-of-the-art quality. However, speaker-agnostic methods are not always desirable, both in terms of quality and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-15 Anastasia Kuznetsova , Aswin Sivaraman , Minje Kim

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a…

Computer Vision and Pattern Recognition · Computer Science 2020-03-06 Ran Yi , Zipeng Ye , Juyong Zhang , Hujun Bao , Yong-Jin Liu

Recent strides in neural speech synthesis technologies, while enjoying widespread applications, have nonetheless introduced a series of challenges, spurring interest in the defence against the threat of misuse and abuse. Notably, source…

Sound · Computer Science 2024-06-18 Chu Yuan Zhang , Jiangyan Yi , Jianhua Tao , Chenglong Wang , Xinrui Yan

Novel text-to-speech systems can generate entirely new voices that were not seen during training. However, it remains a difficult task to efficiently create personalized voices from a high-dimensional speaker space. In this work, we use…

Controllable speech synthesis aims to control the style of generated speech using reference input, which can be of various modalities. Existing face-based methods struggle with robustness and generalization due to data quality constraints,…

Sound · Computer Science 2025-06-27 Rui Niu , Weihao Wu , Jie Chen , Long Ma , Zhiyong Wu

The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Youngjoon Jang , Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Hong-Sun Yang , Yoon-Cheol Ju , Il-Hwan Kim , Byeong-Yeol Kim , Joon Son Chung

Visual perception of a person is easily influenced by many factors such as camera parameters, pose and viewpoint variations. These variations make person Re-Identification (ReID) a challenging problem. Nevertheless, human attributes usually…

Computer Vision and Pattern Recognition · Computer Science 2021-08-11 Fariborz Taherkhani , Ali Dabouei , Sobhan Soleymani , Jeremy Dawson , Nasser M. Nasrabadi

In this work, we focus on source tracing of synthetic speech generation systems (STSGS). Each source embeds distinctive paralinguistic features--such as pitch, tone, rhythm, and intonation--into their synthesized speech, reflecting the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Girish , Mohd Mujtaba Akhtar , Orchid Chetia Phukan , Drishti Singh , Swarup Ranjan Behera , Pailla Balakrishna Reddy , Arun Balaji Buduru , Rajesh Sharma

This paper describes a human-in-the-loop approach to personalized voice synthesis in the absence of reference speech data from the target speaker. It is intended to help vocally disabled individuals restore their lost voices without…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-27 Yusheng Tian , Junbin Liu , Tan Lee

While deep learning technologies are now capable of generating realistic images confusing humans, the research efforts are turning to the synthesis of images for more concrete and application-specific purposes. Facial image generation based…

Computer Vision and Pattern Recognition · Computer Science 2020-06-11 Yeqi Bai , Tao Ma , Lipo Wang , Zhenjie Zhang

A text-to-speech (TTS) model typically factorizes speech attributes such as content, speaker and prosody into disentangled representations.Recent works aim to additionally model the acoustic conditions explicitly, in order to disentangle…

In recent years, the role of image generative models in facial reenactment has been steadily increasing. Such models are usually subject-agnostic and trained on domain-wide datasets. The appearance of the reenacted individual is learned…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Ariel Elazary , Yotam Nitzan , Daniel Cohen-Or

Speech-driven facial animation involves using a speech signal to generate realistic videos of talking faces. Recent deep learning approaches to facial synthesis rely on extracting low-dimensional representations and concatenating them,…

This chapter presents a novel approach to brain-to-speech (BTS) synthesis from intracranial electroencephalography (iEEG) data, emphasizing prosody-aware feature engineering and advanced transformer-based models for high-fidelity speech…

Signal Processing · Electrical Eng. & Systems 2026-04-08 Mohammed Salah Al-Radhi , Géza Németh , Andon Tchechmedjiev , Binbin Xu

We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic information, and speaker…

Since facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate than their audio-only counterparts. Yet, state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2022-04-01 Karren Yang , Dejan Markovic , Steven Krenn , Vasu Agrawal , Alexander Richard

Personalized speech enhancement (PSE) models can improve the audio quality of teleconferencing systems by adapting to the characteristics of a speaker's voice. However, most existing methods require a separate speaker embedding model to…

Sound · Computer Science 2024-06-17 Tanel Pärnamaa , Ando Saabas