中文
相关论文

相关论文: Controlling your Attributes in Voice

200 篇论文

Audio-driven talking face generation has garnered significant interest within the domain of digital human research. Existing methods are encumbered by intricate model architectures that are intricately dependent on each other, complicating…

计算机视觉与模式识别 · 计算机科学 2024-11-08 Dong Zhao , Jiaying Shi , Wenjun Li , Shudong Wang , Shenghui Xu , Zhaoming Pan

Controllable generation using StyleGANs is usually achieved by training the model using labeled data. For audio textures, however, there is currently a lack of large semantically labeled datasets. Therefore, to control generation, we…

音频与语音处理 · 电气工程与系统科学 2024-10-08 Purnima Kamath , Chitralekha Gupta , Lonce Wyse , Suranga Nanayakkara

Voice conversion (VC) techniques aim to modify speaker identity of an utterance while preserving the underlying linguistic information. Most VC approaches ignore modeling of the speaking style (e.g. emotion and emphasis), which may contain…

音频与语音处理 · 电气工程与系统科学 2020-05-20 Songxiang Liu , Yuewen Cao , Shiyin Kang , Na Hu , Xunying Liu , Dan Su , Dong Yu , Helen Meng

Facial attribute editing aims to manipulate single or multiple attributes of a face image, i.e., to generate a new face with desired attributes while preserving other details. Recently, generative adversarial net (GAN) and encoder-decoder…

计算机视觉与模式识别 · 计算机科学 2018-07-26 Zhenliang He , Wangmeng Zuo , Meina Kan , Shiguang Shan , Xilin Chen

Recently, and under the umbrella of Responsible AI, efforts have been made to develop gender-ambiguous synthetic speech to represent with a single voice all individuals in the gender spectrum. However, research efforts have completely…

音频与语音处理 · 电气工程与系统科学 2024-03-18 Maria Koutsogiannaki , Shafel Mc Dowall , Ioannis Agiomyrgiannakis

Advancement in speech technology has brought convenience to our life. However, the concern is on the rise as speech signal contains multiple personal attributes, which would lead to either sensitive information leakage or bias toward…

音频与语音处理 · 电气工程与系统科学 2021-09-09 Yu-Lin Huang , Bo-Hao Su , Y. -W. Peter Hong , Chi-Chun Lee

We present LingGen, a controlled text generation model that allows fine-grained control over a large number of real-valued linguistic attributes. It encodes target attribute values with a dedicated linguistic attribute encoder and…

计算与语言 · 计算机科学 2026-01-27 Mohamed Elgaar , Hadi Amiri

Currently, many multi-speaker speech synthesis and voice conversion systems address speaker variations with an embedding vector. Modeling it directly allows new voices outside of training data to be synthesized. GMM based approaches such as…

声音 · 计算机科学 2023-09-26 Yao Shi , Ming Li

The goal of this paper is to synthesise talking faces with controllable facial motions. To achieve this goal, we propose two key ideas. The first is to establish a canonical space where every face has the same motion patterns but different…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Youngjoon Jang , Kyeongha Rho , Jong-Bin Woo , Hyeongkeun Lee , Jihwan Park , Youshin Lim , Byeong-Yeol Kim , Joon Son Chung

Besides its linguistic content, our speech is rich in biometric information that can be inferred by classifiers. Learning privacy-preserving representations for speech signals enables downstream tasks without sharing unnecessary, private…

声音 · 计算机科学 2021-06-18 Dimitrios Stoidis , Andrea Cavallaro

Controlling chatbot utterance generation with multiple attributes such as personalities, emotions and dialogue acts is a practically useful but under-studied problem. We propose a novel framework called DASC that possesses strong…

计算与语言 · 计算机科学 2023-11-07 Zhiling Zhang , Mengyue Wu , Kenny Q. Zhu

Generative models have been applied in the medical imaging domain for various image recognition and synthesis tasks. However, a more controllable and interpretable image synthesis model is still lacking yet necessary for important…

图像与视频处理 · 电气工程与系统科学 2021-11-15 Jiarong Ye , Yuan Xue , Peter Liu , Richard Zaino , Keith Cheng , Xiaolei Huang

Non-goal oriented dialog agents (i.e. chatbots) aim to produce varying and engaging conversations with a user; however, they typically exhibit either inconsistent personality across conversations or the average personality of all users.…

计算与语言 · 计算机科学 2020-05-14 Alex Boyd , Raul Puri , Mohammad Shoeybi , Mostofa Patwary , Bryan Catanzaro

Accent is an integral part of society, reflecting multiculturalism and shaping how individuals express identity. The majority of English speakers are non-native (L2) speakers, yet current Text-To-Speech (TTS) systems primarily model…

计算与语言 · 计算机科学 2026-03-10 Thanathai Lertpetchpun , Thanapat Trachu , Jihwan Lee , Tiantian Feng , Dani Byrd , Shrikanth Narayanan

Generative models are now capable of synthesizing images, speeches, and videos that are hardly distinguishable from authentic contents. Such capabilities cause concerns such as malicious impersonation and IP theft. This paper investigates a…

声音 · 计算机科学 2022-03-16 Yongbaek Cho , Changhoon Kim , Yezhou Yang , Yi Ren

State-of-the-art generative models (e.g. StyleGAN3 \cite{karras2021alias}) often generate photorealistic images based on vectors sampled from their latent space. However, the ability to control the output is limited. Here we present our…

计算机视觉与模式识别 · 计算机科学 2024-02-28 Róbert Belanec , Peter Lacko , Kristína Malinovská

Generating and manipulating human facial images using high-level attributal controls are important and interesting problems. The models proposed in previous work can solve one of these two problems (generation or manipulation), but not both…

计算机视觉与模式识别 · 计算机科学 2017-04-10 Weidong Yin , Yanwei Fu , Leonid Sigal , Xiangyang Xue

In recent years, image generation has made great strides in improving the quality of images, producing high-fidelity ones. Also, quite recently, there are architecture designs, which enable GAN to unsupervisedly learn the semantic…

计算机视觉与模式识别 · 计算机科学 2022-08-10 Xin Jin , Shu Zhao , Le Zhang , Xin Zhao , Qiang Deng , Chaoen Xiao

The face reenactment is a popular facial animation method where the person's identity is taken from the source image and the facial motion from the driving image. Recent works have demonstrated high quality results by combining the facial…

计算机视觉与模式识别 · 计算机科学 2020-11-10 Soumya Tripathy , Juho Kannala , Esa Rahtu

A Prompt-based Text-To-Speech model allows a user to control different aspects of speech, such as speaking rate and perceived gender, through natural language instruction. Although user-friendly, such approaches are on one hand constrained:…

计算与语言 · 计算机科学 2025-07-14 Atli Sigurgeirsson , Simon King