中文
相关论文

相关论文: Residual-guided Personalized Speech Synthesis base…

200 篇论文

We propose a framework based on Generative Adversarial Networks to disentangle the identity and attributes of faces, such that we can conveniently recombine different identities and attributes for identity preserving face synthesis in open…

计算机视觉与模式识别 · 计算机科学 2018-08-10 Jianmin Bao , Dong Chen , Fang Wen , Houqiang Li , Gang Hua

Sarcastic speech synthesis, which involves generating speech that effectively conveys sarcasm, is essential for enhancing natural interactions in applications such as entertainment and human-computer interaction. However, synthesizing…

计算与语言 · 计算机科学 2025-08-19 Zhu Li , Yuqing Zhang , Xiyuan Gao , Devraj Raghuvanshi , Nagendra Kumar , Shekhar Nayak , Matt Coler

Recent work has shown improved lesion detectability and flexibility to reconstruction hyperparameters (e.g. scanner geometry or dose level) when PET images are reconstructed by leveraging pre-trained diffusion models. Such methods train a…

医学物理 · 物理学 2025-08-28 George Webber , Alexander Hammers , Andrew P. King , Andrew J. Reader

This work presents self-supervised learning methods for developing monaural speaker-specific (i.e., personalized) speech enhancement models. While generalist models must broadly address many speakers, specialist models can adapt their…

音频与语音处理 · 电气工程与系统科学 2022-07-28 Aswin Sivaraman , Minje Kim

In this paper we investigate a technique to find out vocal source based features from the LP residual of speech signal for automatic speaker identification. Autocorrelation with some specific lag is computed for the residual signal to…

人机交互 · 计算机科学 2011-05-24 Md. Sahidullah , Goutam Saha

While large-scale pre-trained text-to-image models can synthesize diverse and high-quality human-centric images, an intractable problem is how to preserve the face identity for conditioned face images. Existing methods either require…

计算机视觉与模式识别 · 计算机科学 2023-07-04 Zhuowei Chen , Shancheng Fang , Wei Liu , Qian He , Mengqi Huang , Yongdong Zhang , Zhendong Mao

Face synthesis, including face aging, in particular, has been one of the major topics that witnessed a substantial improvement in image fidelity by using generative adversarial networks (GANs). Most existing face aging approaches divide the…

计算机视觉与模式识别 · 计算机科学 2021-05-04 Zeqi Li , Ruowei Jiang , Parham Aarabi

We present a generative model for controllable person image synthesis,as shown in Figure , which can be applied to pose-guided person image synthesis, $i.e.$, converting the pose of a source person image to the target pose while preserving…

计算机视觉与模式识别 · 计算机科学 2020-12-24 Shilong Shen

We propose the first approach to automatically and jointly synthesize both the synchronous 3D conversational body and hand gestures, as well as 3D face and head animations, of a virtual character from speech input. Our algorithm uses a CNN…

计算机视觉与模式识别 · 计算机科学 2021-02-16 Ikhsanul Habibie , Weipeng Xu , Dushyant Mehta , Lingjie Liu , Hans-Peter Seidel , Gerard Pons-Moll , Mohamed Elgharib , Christian Theobalt

Facial attributes can provide rich ancillary information which can be utilized for different applications such as targeted marketing, human computer interaction, and law enforcement. This research focuses on facial attribute prediction…

计算机视觉与模式识别 · 计算机科学 2018-03-21 Akshay Sethi , Maneet Singh , Richa Singh , Mayank Vatsa

In this work, we present an end-to-end binaural speech synthesis system that combines a low-bitrate audio codec with a powerful binaural decoder that is capable of accurate speech binauralization while faithfully reconstructing…

声音 · 计算机科学 2022-07-11 Wen Chin Huang , Dejan Markovic , Alexander Richard , Israel Dejene Gebru , Anjali Menon

Autoregressive neural vocoders have achieved outstanding performance in speech synthesis tasks such as text-to-speech and voice conversion. An autoregressive vocoder predicts a sample at some time step conditioned on those at previous time…

声音 · 计算机科学 2024-06-06 Po-chun Hsu , Da-rong Liu , Andy T. Liu , Hung-yi Lee

Recently end-to-end neural audio/speech coding has shown its great potential to outperform traditional signal analysis based audio codecs. This is mostly achieved by following the VQ-VAE paradigm where blind features are learned,…

声音 · 计算机科学 2023-02-28 Xue Jiang , Xiulian Peng , Yuan Zhang , Yan Lu

This paper introduces a multi-scale speech style modeling method for end-to-end expressive speech synthesis. The proposed method employs a multi-scale reference encoder to extract both the global-scale utterance-level and the local-scale…

声音 · 计算机科学 2021-04-09 Xiang Li , Changhe Song , Jingbei Li , Zhiyong Wu , Jia Jia , Helen Meng

This paper proposes a zero-shot text-to-speech (TTS) conditioned by a self-supervised speech-representation model acquired through self-supervised learning (SSL). Conventional methods with embedding vectors from x-vector or global style…

声音 · 计算机科学 2023-12-19 Kenichi Fujita , Takanori Ashihara , Hiroki Kanagawa , Takafumi Moriya , Yusuke Ijima

The task of synthetic speech generation is to generate language content from a given text, then simulating fake human voice.The key factors that determine the effect of synthetic speech generation mainly include speed of generation,…

声音 · 计算机科学 2023-07-04 Sheng Zhao , Qilong Yuan , Yibo Duan , Zhuoyue Chen

Personalized image synthesis has emerged as a pivotal application in text-to-image generation, enabling the creation of images featuring specific subjects in diverse contexts. While diffusion models have dominated this domain,…

计算机视觉与模式识别 · 计算机科学 2025-04-18 Kaiyue Sun , Xian Liu , Yao Teng , Xihui Liu

Recent face reenactment works are limited by the coarse reference landmarks, leading to unsatisfactory identity preserving performance due to the distribution gap between the manipulated landmarks and those sampled from a real person. To…

计算机视觉与模式识别 · 计算机科学 2021-10-13 Haichao Zhang , Youcheng Ben , Weixi Zhang , Tao Chen , Gang Yu , Bin Fu

The modeling of speech production often relies on a source-filter approach. Although methods parameterizing the filter have nowadays reached a certain maturity, there is still a lot to be gained for several speech processing applications in…

声音 · 计算机科学 2020-01-07 Thomas Drugman , Thierry Dutoit

In recent Text-to-Speech (TTS) systems, a neural vocoder often generates speech samples by solely conditioning on acoustic features predicted from an acoustic model. However, there are always distortions existing in the predicted acoustic…

声音 · 计算机科学 2023-04-25 Jianzong Wang , Xulong Zhang , Haobin Tang , Aolan Sun , Ning Cheng , Jing Xiao