English
Related papers

Related papers: Controlling your Attributes in Voice

200 papers

The speech enhancement task usually consists of removing additive noise or reverberation that partially mask spoken utterances, affecting their intelligibility. However, little attention is drawn to other, perhaps more aggressive signal…

Sound · Computer Science 2019-04-09 Santiago Pascual , Joan Serrà , Antonio Bonafonte

Embodied agents, in the form of virtual agents or social robots, are rapidly becoming more widespread. In human-human interactions, humans use nonverbal behaviours to convey their attitudes, feelings, and intentions. Therefore, this…

Artificial Intelligence · Computer Science 2026-04-30 Carson Yu Liu , Gelareh Mohammadi , Yang Song , Wafa Johal

We present a method that generates expressive talking heads from a single facial image with audio as the only input. In contrast to previous approaches that attempt to learn direct mappings from audio to raw pixels or points for creating…

Computer Vision and Pattern Recognition · Computer Science 2021-02-26 Yang Zhou , Xintong Han , Eli Shechtman , Jose Echevarria , Evangelos Kalogerakis , Dingzeyu Li

Human voice is the source of several important information. This is in the form of features. These Features help in interpreting various features associated with the speaker and speech. The speaker dependent work researchersare targeted…

Sound · Computer Science 2022-03-30 Shankhanil Ghosh , Chhanda Saha , Naagamani Molakathaala

Generative adversarial networks have been widely used in image synthesis in recent years and the quality of the generated image has been greatly improved. However, the flexibility to control and decouple facial attributes (e.g., eyes, nose,…

Computer Vision and Pattern Recognition · Computer Science 2021-08-26 Xiao Cui , Wengang Zhou , Yang Hu , Weilun Wang , Houqiang Li

Recent approaches have achieved great success in image generation from structured inputs, e.g., semantic segmentation, scene graph or layout. Although these methods allow specification of objects and their locations at image-level, they…

Computer Vision and Pattern Recognition · Computer Science 2020-08-28 Ke Ma , Bo Zhao , Leonid Sigal

The advance of Generative Adversarial Networks (GANs) enables realistic face image synthesis. However, synthesizing face images that preserve facial identity as well as have high diversity within each identity remains challenging. To…

Computer Vision and Pattern Recognition · Computer Science 2018-12-05 Yujun Shen , Bolei Zhou , Ping Luo , Xiaoou Tang

AI-driven image generation has improved significantly in recent years. Generative adversarial networks (GANs), like StyleGAN, are able to generate high-quality realistic data and have artistic control over the output, as well. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2022-04-19 Mohamed Shawky Sabae , Mohamed Ahmed Dardir , Remonda Talaat Eskarous , Mohamed Ramzy Ebbed

The field of text-to-audio generation has seen significant advancements, and yet the ability to finely control the acoustic characteristics of generated audio remains under-explored. In this paper, we introduce a novel yet simple approach…

Sound · Computer Science 2024-12-16 Sonal Kumar , Prem Seetharaman , Justin Salamon , Dinesh Manocha , Oriol Nieto

Unlike other data modalities such as text and vision, speech does not lend itself to easy interpretation. While lay people can understand how to describe an image or sentence via perception, non-expert descriptions of speech often end at…

Sound · Computer Science 2023-10-05 Robin Netzorg , Bohan Yu , Andrea Guzman , Peter Wu , Luna McNulty , Gopala Anumanchipalli

We devise a cascade GAN approach to generate talking face video, which is robust to different face shapes, view angles, facial characteristics, and noisy audio conditions. Instead of learning a direct mapping from audio to video frames, we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-13 Lele Chen , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

Achieving nuanced and accurate emulation of human voice has been a longstanding goal in artificial intelligence. Although significant progress has been made in recent years, the mainstream of speech synthesis models still relies on…

Sound · Computer Science 2024-03-04 Weiwei Lin , Chenhang He , Man-Wai Mak , Jiachen Lian , Kong Aik Lee

Speech-driven facial animation is the process that automatically synthesizes talking characters based on speech signals. The majority of work in this domain creates a mapping from audio features to visual features. This approach often…

Computer Vision and Pattern Recognition · Computer Science 2019-06-18 Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

Humans can perceive speakers' characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vision-driven Text-to-speech (TTS) scholars grounded their…

Sound · Computer Science 2025-04-17 Tian-Hao Zhang , Jiawei Zhang , Jun Wang , Xinyuan Qian , Xu-Cheng Yin

Controlled generation refers to the problem of creating text that contains stylistic or semantic attributes of interest. Many approaches reduce this problem to training a predictor of the desired attribute. For example, researchers hoping…

Computation and Language · Computer Science 2023-06-02 Carolina Zheng , Claudia Shi , Keyon Vafa , Amir Feder , David M. Blei

Text-to-speech models trained on large-scale datasets have demonstrated impressive in-context learning capabilities and naturalness. However, control of speaker identity and style in these models typically requires conditioning on reference…

Sound · Computer Science 2024-02-08 Dan Lyth , Simon King

Speech enhancement is an essential task of improving speech quality in noise scenario. Several state-of-the-art approaches have introduced visual information for speech enhancement,since the visual aspect of speech is essentially unaffected…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-21 Xinmeng Xu , Yang Wang , Dongxiang Xu , Yiyuan Peng , Cong Zhang , Jie Jia , Binbin Chen

Recent advancements in Text-to-Speech (TTS) systems have enabled the generation of natural and expressive speech from textual input. Accented TTS aims to enhance user experience by making the synthesized speech more relatable to minority…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-18 Jan Melechovsky , Ambuj Mehrish , Berrak Sisman , Dorien Herremans

Privacy-preserving voice conversion aims to remove only the attributes of speech audio that convey identity information, keeping other speech characteristics intact. This paper presents a mechanism for privacy-preserving voice conversion…

Sound · Computer Science 2024-09-24 Jacob J Webber , Oliver Watts , Gustav Eje Henter , Jennifer Williams , Simon King

We propose a novel method for generating high-resolution videos of talking-heads from speech audio and a single 'identity' image. Our method is based on a convolutional neural network model that incorporates a pre-trained StyleGAN…

Computer Vision and Pattern Recognition · Computer Science 2022-09-12 Mohammed M. Alghamdi , He Wang , Andrew J. Bulpitt , David C. Hogg