English
Related papers

Related papers: Seeing What You Say: Expressive Image Generation f…

200 papers

Generating speech from a face image is crucial for developing virtual humans capable of interacting using their unique voices, without relying on pre-recorded human speech. In this paper, we propose Face-StyleSpeech, a zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Minki Kang , Wooseok Han , Eunho Yang

Text-based speech editing allows users to edit speech by intuitively cutting, copying, and pasting text to speed up the process of editing speech. In the previous work, CampNet (context-aware mask prediction network) is proposed to realize…

Sound · Computer Science 2022-12-21 Tao Wang , Jiangyan Yi , Ruibo Fu , Jianhua Tao , Zhengqi Wen , Chu Yuan Zhang

Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-17 Leying Zhang , Yao Qian , Long Zhou , Shujie Liu , Dongmei Wang , Xiaofei Wang , Midia Yousefi , Yanmin Qian , Jinyu Li , Lei He , Sheng Zhao , Michael Zeng

Speech sounds of spoken language are obtained by varying configuration of the articulators surrounding the vocal tract. They contain abundant information that can be utilized to better understand the underlying mechanism of human speech…

Image and Video Processing · Electrical Eng. & Systems 2021-06-17 Laxmi Pandey , Ahmed Sabbir Arif

Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selected from a predefined set, average variation learned from…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-10 Sarath Sivaprasad , Saiteja Kosgi , Vineet Gandhi

In this paper, we propose a new methodology for emotional speech recognition using visual deep neural network models. We employ the transfer learning capabilities of the pre-trained computer vision deep models to have a mandate for the…

Computer Vision and Pattern Recognition · Computer Science 2022-04-08 Waleed Ragheb , Mehdi Mirzapour , Ali Delfardi , Hélène Jacquenet , Lawrence Carbon

State-of-the-art speech synthesis models try to get as close as possible to the human voice. Hence, modelling emotions is an essential part of Text-To-Speech (TTS) research. In our work, we selected FastSpeech2 as the starting point and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-04 Daria Diatlova , Vitaly Shutov

It is in high demand to generate facial animation with high realism, but it remains a challenging task. Existing approaches of speech-driven facial animation can produce satisfactory mouth movement and lip synchronization, but show weakness…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Yutong Chen , Junhong Zhao , Wei-Qiang Zhang

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Jiaxin Ye , Hongming Shan

Generating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessitates disentangled…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Xuli Shen , Hua Cai , Dingding Yu , Weilin Shen , Qing Xu , Xiangyang Xue

Text-to-image models have rapidly evolved from casual creative tools to professional-grade systems, achieving unprecedented levels of image quality and realism. Yet, most models are trained to map short prompts into detailed images,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Eyal Gutflaish , Eliran Kachlon , Hezi Zisman , Tal Hacham , Nimrod Sarid , Alexander Visheratin , Saar Huberman , Gal Davidi , Guy Bukchin , Kfir Goldberg , Ron Mokady

Recent advancements in generative speech models based on audio-text prompts have enabled remarkable innovations like high-quality zero-shot text-to-speech. However, existing models still face limitations in handling diverse audio-text…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-27 Xiaofei Wang , Manthan Thakker , Zhuo Chen , Naoyuki Kanda , Sefik Emre Eskimez , Sanyuan Chen , Min Tang , Shujie Liu , Jinyu Li , Takuya Yoshioka

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. However, existing studies…

Multimedia · Computer Science 2023-05-25 Zheng-Yan Sheng , Yang Ai , Zhen-Hua Ling

Spoken dialogue systems often rely on cascaded pipelines that transcribe, process, and resynthesize speech. While effective, this design discards paralinguistic cues and limits expressivity. Recent end-to-end methods reduce latency and…

In this paper, we introduce V2SFlow, a novel Video-to-Speech (V2S) framework designed to generate natural and intelligible speech directly from silent talking face videos. While recent V2S systems have shown promising results on constrained…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Jeongsoo Choi , Ji-Hoon Kim , Jinyu Li , Joon Son Chung , Shujie Liu

Speech emotion conversion is the task of modifying the perceived emotion of a speech utterance while preserving the lexical content and speaker identity. In this study, we cast the problem of emotion conversion as a spoken language…

Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating synthetic versions…

Machine Learning · Computer Science 2021-01-15 Simon Alexanderson , Éva Székely , Gustav Eje Henter , Taras Kucherenko , Jonas Beskow

In recent years, the field of image generation has been revolutionized by the application of autoregressive transformers and DDPMs. These approaches model the process of image generation as a step-wise probabilistic processes and leverage…

Sound · Computer Science 2023-05-25 James Betker

The goal of this work is to reconstruct speech from a silent talking face video. Recent studies have shown impressive performance on synthesizing speech from silent talking face videos. However, they have not explicitly considered on…

Computer Vision and Pattern Recognition · Computer Science 2022-07-21 Joanna Hong , Minsu Kim , Yong Man Ro

A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module. Building these components often requires extensive domain expertise and may contain…