中文
相关论文

相关论文: Improving Emotional Speech Synthesis by Using SUS-…

200 篇论文

Retrieving unlabeled videos by textual queries, known as Ad-hoc Video Search (AVS), is a core theme in multimedia data management and retrieval. The success of AVS counts on cross-modal representation learning that encodes both query…

计算机视觉与模式识别 · 计算机科学 2020-11-25 Xirong Li , Fangming Zhou , Chaoxi Xu , Jiaqi Ji , Gang Yang

Speech Neuroprostheses have the potential to enable communication for people with dysarthria or anarthria. Recent advances have demonstrated high-quality text decoding and speech synthesis from electrocorticographic grids placed on the…

Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse emotional speech…

音频与语音处理 · 电气工程与系统科学 2025-06-04 Xiaoxue Gao , Huayun Zhang , Nancy F. Chen

Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human emotional…

声音 · 计算机科学 2024-12-13 Weizhen Bian , Yubo Zhou , Kaitai Zhang , Xiaohan Gu

Emotion is a complicated notion present in music that is hard to capture even with fine-tuned feature engineering. In this paper, we investigate the utility of state-of-the-art pre-trained deep audio embedding methods to be used in the…

声音 · 计算机科学 2021-04-15 Eunjeong Koh , Shlomo Dubnov

In this paper, we propose a feature reinforcement method under the sequence-to-sequence neural text-to-speech (TTS) synthesis framework. The proposed method utilizes the multiple input encoder to take three levels of text information, i.e.,…

声音 · 计算机科学 2019-03-07 Huaiping Ming , Lei He , Haohan Guo , Frank K. Soong

Acoustic word embeddings (AWEs) are fixed-dimensional vector representations of speech segments that encode phonetic content so that different realisations of the same word have similar embeddings. In this paper we explore semantic AWE…

音频与语音处理 · 电气工程与系统科学 2023-07-06 Christiaan Jacobs , Herman Kamper

Images evoke emotions that profoundly influence perception, often prioritized over content. Current Image Emotional Synthesis (IES) approaches artificially separate generation and editing tasks, creating inefficiencies and limiting…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Yingjie Xia , Xi Wang , Jinglei Shi , Vicky Kalogeiton , Jian Yang

This paper introduces EmoSSLSphere, a novel framework for multilingual emotional text-to-speech (TTS) synthesis that combines spherical emotion vectors with discrete token features derived from self-supervised learning (SSL). By encoding…

音频与语音处理 · 电气工程与系统科学 2025-10-07 Joonyong Park , Kenichi Nakamura

Pre-trained models (PTMs) have shown great promise in the speech and audio domain. Embeddings leveraged from these models serve as inputs for learning algorithms with applications in various downstream tasks. One such crucial task is Speech…

音频与语音处理 · 电气工程与系统科学 2023-04-25 Orchid Chetia Phukan , Arun Balaji Buduru , Rajesh Sharma

Emotions play a central role in human communication, shaping trust, engagement, and social interaction. As artificial intelligence systems powered by large language models become increasingly integrated into everyday life, enabling them to…

音频与语音处理 · 电气工程与系统科学 2026-03-11 Soumya Dutta

Affect is an emotional characteristic encompassing valence, arousal, and intensity, and is a crucial attribute for enabling authentic conversations. While existing text-to-speech (TTS) and speech-to-speech systems rely on strength embedding…

Cross-speaker emotion transfer speech synthesis aims to synthesize emotional speech for a target speaker by transferring the emotion from reference speech recorded by another (source) speaker. In this task, extracting speaker-independent…

声音 · 计算机科学 2022-07-05 Tao Li , Xinsheng Wang , Qicong Xie , Zhichao Wang , Mingqi Jiang , Lei Xie

Previous multilingual text-to-speech (TTS) approaches have considered leveraging monolingual speaker data to enable cross-lingual speech synthesis. However, such data-efficient approaches have ignored synthesizing emotional aspects of…

音频与语音处理 · 电气工程与系统科学 2023-08-01 Xinfa Zhu , Yi Lei , Tao Li , Yongmao Zhang , Hongbin Zhou , Heng Lu , Lei Xie

Multimodal Sentiment Analysis (MSA) stands as a critical research frontier, seeking to comprehensively unravel human emotions by amalgamating text, audio, and visual data. Yet, discerning subtle emotional nuances within audio and video…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Sheng Wu , Xiaobao Wang , Longbiao Wang , Dongxiao He , Jianwu Dang

Emotional voice conversion (EVC) aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. In this paper, we study the disentanglement and recomposition of emotional…

声音 · 计算机科学 2020-11-05 Kun Zhou , Berrak Sisman , Haizhou Li

Word embedding models such as GloVe are widely used in natural language processing (NLP) research to convert words into vectors. Here, we provide a preliminary guide to probe latent emotions in text through GloVe word vectors. First, we…

计算与语言 · 计算机科学 2019-08-22 Zhengxuan Wu , Yueyi Jiang

We propose a Text-to-Speech method to create an unseen expressive style using one utterance of expressive speech of around one second. Specifically, we enhance the disentanglement capabilities of a state-of-the-art sequence-to-sequence…

机器学习 · 计算机科学 2020-02-18 Vatsal Aggarwal , Marius Cotescu , Nishant Prateek , Jaime Lorenzo-Trueba , Roberto Barra-Chicote

Speaker embedding extractors significantly influence the performance of clustering-based speaker diarisation systems. Conventionally, only one embedding is extracted from each speech segment. However, because of the sliding window approach,…

声音 · 计算机科学 2022-11-09 Hee-Soo Heo , Youngki Kwon , Bong-Jin Lee , You Jin Kim , Jee-weon Jung

Expressive speech synthesis requires vibrant prosody and well-timed pauses. We propose an effective strategy to augment a small dataset to train an expressive end-to-end Text-to-Speech model. We merge audios of emotionally congruent text…

声音 · 计算机科学 2026-02-12 Raymond Chung