中文
相关论文

相关论文: JoyVoice: Long-Context Conditioning for Anthropomo…

200 篇论文

How much can we infer about an emotional voice solely from an expressive face? This intriguing question holds great potential for applications such as virtual character dubbing and aiding individuals with expressive language disorders.…

声音 · 计算机科学 2025-02-04 Jiaxin Ye , Boyuan Cao , Hongming Shan

The idea of using phonological features instead of phonemes as input to sequence-to-sequence TTS has been recently proposed for zero-shot multilingual speech synthesis. This approach is useful for code-switching, as it facilitates the…

Recent advancements in large language models (LLMs) and multimodal speech-text models have laid the groundwork for seamless voice interactions, enabling real-time, natural, and human-like conversations. Previous models for voice…

Spoken dialogue is essential for human-AI interactions, providing expressive capabilities beyond text. Developing effective spoken dialogue systems (SDSs) requires large-scale, high-quality, and diverse spoken dialogue corpora. However,…

计算与语言 · 计算机科学 2026-04-03 Wataru Nakata , Kentaro Seki , Hitomi Yanaka , Yuki Saito , Shinnosuke Takamichi , Hiroshi Saruwatari

We present MParrotTTS, a unified multilingual, multi-speaker text-to-speech (TTS) synthesis model that can produce high-quality speech. Benefiting from a modularized training paradigm exploiting self-supervised speech representations,…

声音 · 计算机科学 2023-05-23 Neil Shah , Vishal Tambrahalli , Saiteja Kosgi , Niranjan Pedanekar , Vineet Gandhi

Style transfer for out-of-domain (OOD) speech synthesis aims to generate speech samples with unseen style (e.g., speaker identity, emotion, and prosody) derived from an acoustic reference, while facing the following challenges: 1) The…

音频与语音处理 · 电气工程与系统科学 2022-10-14 Rongjie Huang , Yi Ren , Jinglin Liu , Chenye Cui , Zhou Zhao

Thanks to improvements in machine learning techniques, including deep learning, speech synthesis is becoming a machine learning task. To accelerate speech synthesis research, we are developing Japanese voice corpora reasonably accessible…

We present eCat, a novel end-to-end multispeaker model capable of: a) generating long-context speech with expressive and contextually appropriate prosody, and b) performing fine-grained prosody transfer between any pair of seen speakers.…

Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, existing CSS systems are…

声音 · 计算机科学 2025-02-28 Weihao wu , Zhiwei Lin , Yixuan Zhou , Jingbei Li , Rui Niu , Qinghua Wu , Songjun Cao , Long Ma , Zhiyong Wu

The performance of speaker verification systems degrades significantly under language mismatch, a critical challenge exacerbated by the field's reliance on English-centric data. To address this, we propose the TidyVoice Challenge for…

音频与语音处理 · 电气工程与系统科学 2026-01-30 Aref Farhadipour , Jan Marquenie , Srikanth Madikeri , Teodora Vukovic , Volker Dellwo , Kathy Reid , Francis M. Tyers , Ingo Siegert , Eleanor Chodroff

A judicious combination of dictionary learning methods, block sparsity and source recovery algorithm are used in a hierarchical manner to identify the noises and the speakers from a noisy conversation between two people. Conversations are…

声音 · 计算机科学 2016-10-31 K V Vijay Girish , A G Ramakrishnan , T V Ananthapadmanabha

We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive auditory experiences that align with the real physical world.…

声音 · 计算机科学 2026-02-04 Chengyuan Ma , Jiawei Jin , Ruijie Xiong , Chunxiang Jin , Canxiang Yan , Wenming Yang

Multilingual end-to-end (E2E) models have shown great promise in expansion of automatic speech recognition (ASR) coverage of the world's languages. They have shown improvement over monolingual systems, and have simplified training and…

音频与语音处理 · 电气工程与系统科学 2019-09-13 Anjuli Kannan , Arindrima Datta , Tara N. Sainath , Eugene Weinstein , Bhuvana Ramabhadran , Yonghui Wu , Ankur Bapna , Zhifeng Chen , Seungji Lee

Recent advancements in conversational systems have significantly enhanced human-machine interactions across various domains. However, training these systems is challenging due to the scarcity of specialized dialogue data. Traditionally,…

计算与语言 · 计算机科学 2026-05-29 Heydar Soudani , Roxana Petcu , Evangelos Kanoulas , Faegheh Hasibi

Singing voice synthesis has made remarkable progress in generating natural and high-quality voices. However, existing methods rarely provide precise control over vocal techniques such as intensity, mixed voice, falsetto, bubble, and breathy…

声音 · 计算机科学 2025-04-22 Wenxiang Guo , Yu Zhang , Changhao Pan , Rongjie Huang , Li Tang , Ruiqi Li , Zhiqing Hong , Yongqi Wang , Zhou Zhao

Currently, many multi-speaker speech synthesis and voice conversion systems address speaker variations with an embedding vector. Modeling it directly allows new voices outside of training data to be synthesized. GMM based approaches such as…

声音 · 计算机科学 2023-09-26 Yao Shi , Ming Li

Emotion is essential in spoken communication, yet most existing frameworks in speech emotion modeling rely on predefined categories or low-dimensional continuous attributes, which offer limited expressive capacity. Recent advances in speech…

音频与语音处理 · 电气工程与系统科学 2026-04-07 Tianhua Qi , Wenming Zheng , Björn W. Schuller , Zhaojie Luo , Haizhou Li

The ability to envisage the visual of a talking face based just on hearing a voice is a unique human capability. There have been a number of works that have solved for this ability recently. We differ from these approaches by enabling a…

计算机视觉与模式识别 · 计算机科学 2020-11-24 Ravindra Yadav , Ashish Sardana , Vinay P Namboodiri , Rajesh M Hegde

In this paper, we present CopyCat2 (CC2), a novel model capable of: a) synthesizing speech with different speaker identities, b) generating speech with expressive and contextually appropriate prosody, and c) transferring prosody at…

音频与语音处理 · 电气工程与系统科学 2022-06-28 Sri Karlapati , Penny Karanasou , Mateusz Lajszczak , Ammar Abbas , Alexis Moinet , Peter Makarov , Ray Li , Arent van Korlaar , Simon Slangen , Thomas Drugman

Novel text-to-speech systems can generate entirely new voices that were not seen during training. However, it remains a difficult task to efficiently create personalized voices from a high-dimensional speaker space. In this work, we use…