中文
相关论文

相关论文: PromptSpeaker: Speaker Generation Based on Text De…

200 篇论文

Style voice conversion aims to transform the style of source speech to a desired style according to real-world application demands. However, the current style voice conversion approach relies on pre-defined labels or reference speech to…

音频与语音处理 · 电气工程与系统科学 2023-12-27 Jixun Yao , Yuguang Yang , Yi Lei , Ziqian Ning , Yanni Hu , Yu Pan , Jingjing Yin , Hongbin Zhou , Heng Lu , Lei Xie

We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and eye blinks.…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Dongchan Min , Minyoung Song , Eunji Ko , Sung Ju Hwang

Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content and style. Some studies try to improve the granularity by…

音频与语音处理 · 电气工程与系统科学 2025-04-01 Yuanyuan Wang , Hangting Chen , Dongchao Yang , Zhiyong Wu , Xixin Wu

We investigate the emergent abilities of the recently proposed web-scale speech model Whisper, by adapting it to unseen tasks with prompt engineering. We selected three tasks: audio-visual speech recognition (AVSR), code-switched speech…

音频与语音处理 · 电气工程与系统科学 2023-08-17 Puyuan Peng , Brian Yan , Shinji Watanabe , David Harwath

It is appealing to have a system that generates a story or scripts automatically from a story-line, even though this is still out of our reach. In dialogue systems, it would also be useful to drive dialogues by a dialogue plan. In this…

计算与语言 · 计算机科学 2020-11-17 Yutao Zhu , Ruihua Song , Zhicheng Dou , Jian-Yun Nie , Jin Zhou

This paper presents a novel framework for speech-driven gesture production, applicable to virtual agents to enhance human-computer interaction. Specifically, we extend recent deep-learning-based, data-driven methods for speech-driven…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Taras Kucherenko , Dai Hasegawa , Naoshi Kaneko , Gustav Eje Henter , Hedvig Kjellström

Current mainstream audio generation methods primarily rely on simple text prompts, often failing to capture the nuanced details necessary for multi-style audio generation. To address this limitation, the Sound Event Enhanced Prompt Adapter…

音频与语音处理 · 电气工程与系统科学 2024-09-17 Chenxu Xiong , Ruibo Fu , Shuchen Shi , Zhengqi Wen , Jianhua Tao , Tao Wang , Chenxing Li , Chunyu Qiang , Yuankun Xie , Xin Qi , Guanjun Li , Zizheng Yang

In this report we present a system that can generate political speeches for a desired political party. Furthermore, the system allows to specify whether a speech should hold a supportive or opposing opinion. The system relies on a…

计算与语言 · 计算机科学 2016-01-21 Valentin Kassarnig

Scaling text-to-speech (TTS) to large-scale, multi-speaker, and in-the-wild datasets is important to capture the diversity in human speech such as speaker identities, prosodies, and styles (e.g., singing). Current large TTS systems usually…

音频与语音处理 · 电气工程与系统科学 2023-05-31 Kai Shen , Zeqian Ju , Xu Tan , Yanqing Liu , Yichong Leng , Lei He , Tao Qin , Sheng Zhao , Jiang Bian

Recent advancements in zero-shot speech generation have enabled models to synthesize speech that mimics speaker identity and speaking style from speech prompts. However, these models' effectiveness is significantly limited in real-world…

音频与语音处理 · 电气工程与系统科学 2025-08-14 Boyu Zhu , Cheng Gong , Muyang Wu , Ruihao Jing , Fan Liu , Xiaolei Zhang , Chi Zhang , Xuelong Li

The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately. However, this…

Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to…

音频与语音处理 · 电气工程与系统科学 2024-12-17 Leying Zhang , Yao Qian , Long Zhou , Shujie Liu , Dongmei Wang , Xiaofei Wang , Midia Yousefi , Yanmin Qian , Jinyu Li , Lei He , Sheng Zhao , Michael Zeng

Spontaneous style speech synthesis, which aims to generate human-like speech, often encounters challenges due to the scarcity of high-quality data and limitations in model capabilities. Recent language model-based TTS systems can be trained…

声音 · 计算机科学 2024-07-19 Weiqin Li , Peiji Yang , Yicheng Zhong , Yixuan Zhou , Zhisheng Wang , Zhiyong Wu , Xixin Wu , Helen Meng

Speech representations learned from Self-supervised learning (SSL) models can benefit various speech processing tasks. However, utilizing SSL representations usually requires fine-tuning the pre-trained models or designing task-specific…

音频与语音处理 · 电气工程与系统科学 2022-07-12 Kai-Wei Chang , Wei-Cheng Tseng , Shang-Wen Li , Hung-yi Lee

Methods for modeling and controlling prosody with acoustic features have been proposed for neural text-to-speech (TTS) models. Prosodic speech can be generated by conditioning acoustic features. However, synthesized speech with a large…

音频与语音处理 · 电气工程与系统科学 2021-06-30 Taejun Bak , Jae-Sung Bae , Hanbin Bae , Young-Ik Kim , Hoon-Young Cho

Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing systems struggle to maintain speaker consistency, model…

The rapid development of the Internet has profoundly changed human life. Humans are increasingly expressing themselves and interacting with others on social media platforms. However, although artificial intelligence technology has been…

计算与语言 · 计算机科学 2024-07-11 Haochen Xue , Chong Zhang , Chengzhi Liu , Fangyu Wu , Xiaobo Jin

During speech, people spontaneously gesticulate, which plays a key role in conveying information. Similarly, realistic co-speech gestures are crucial to enable natural and smooth interactions with social agents. Current end-to-end co-speech…

Reference-based Text-to-Speech (TTS) models can generate multiple, prosodically-different renditions of the same target text. Such models jointly learn a latent acoustic space during training, which can be sampled from during inference.…

计算与语言 · 计算机科学 2023-09-20 Atli Thor Sigurgeirsson , Simon King

Controllable speech generation methods typically rely on single or fixed prompts, hindering creativity and flexibility. These limitations make it difficult to meet specific user needs in certain scenarios, such as adjusting the style while…

音频与语音处理 · 电气工程与系统科学 2025-05-01 Hanzhao Li , Yuke Li , Xinsheng Wang , Jingbin Hu , Qicong Xie , Shan Yang , Lei Xie