中文
相关论文

相关论文: Affectron: Emotional Speech Synthesis with Affecti…

200 篇论文

It is important for machines to interpret human emotions properly for better human-machine communications, as emotion is an essential part of human-to-human communications. One aspect of emotion is reflected in the language we use. How to…

计算与语言 · 计算机科学 2018-08-23 Ji Ho Park

Voice conversion (VC) transforms an utterance to sound like another person without changing the linguistic content. A recently proposed generative adversarial network-based VC method, StarGANv2-VC is very successful in generating…

音频与语音处理 · 电气工程与系统科学 2023-09-15 Arnab Das , Suhita Ghosh , Tim Polzehl , Sebastian Stober

In this paper, we have defined a novel task of affective feedback synthesis that deals with generating feedback for input text & corresponding image in a similar way as humans respond towards the multimodal data. A feedback synthesis system…

多媒体 · 计算机科学 2022-04-01 Puneet Kumar , Gaurav Bhat , Omkar Ingle , Daksh Goyal , Balasubramanian Raman

As a common way of emotion signaling via non-linguistic vocalizations, vocal burst (VB) plays an important role in daily social interaction. Understanding and modeling human vocal bursts are indispensable for developing robust and general…

音频与语音处理 · 电气工程与系统科学 2023-03-15 Jinchao Li , Xixin Wu , Kaitao Song , Dongsheng Li , Xunying Liu , Helen Meng

Emergent communication (EmCom) with deep neural network-based agents promises to yield insights into the nature of human language, but remains focused primarily on a few subfield-specific goals and metrics that prioritize communication…

计算与语言 · 计算机科学 2025-10-22 Miles Gilberti , Shane Storks , Huteng Dai

Audio Sentiment Analysis is a popular research area which extends the conventional text-based sentiment analysis to depend on the effectiveness of acoustic features extracted from speech. However, current progress on audio sentiment…

音频与语音处理 · 电气工程与系统科学 2019-08-01 Feiyang Chen , Ziqian Luo

Recent advances in emotional voice conversion (EVC) have enabled the generation of expressive synthetic speech, raising new concerns in audio deepfake detection. Existing approaches treat speech as a homogeneous signal and largely overlook…

声音 · 计算机科学 2026-05-06 Vamshi Nallaguntla , Shruti Kshirsagar , Anderson R. Avila

Human conversational styles are measured by the sense of humor, personality, and tone of voice. These characteristics have become essential for conversational intelligent virtual assistants. However, most of the state-of-the-art intelligent…

Storytelling's captivating potential makes it a fascinating research area, with implications for entertainment, education, therapy, and cognitive studies. In this paper, we propose Affective Story Generator (AffGen) for generating…

计算与语言 · 计算机科学 2024-04-03 Tenghao Huang , Ehsan Qasemi , Bangzheng Li , He Wang , Faeze Brahman , Muhao Chen , Snigdha Chaturvedi

Dimensional representations of speech emotions such as the arousal-valence (AV) representation provide a continuous and fine-grained description and control than their categorical counterparts. They have wide applications in tasks such as…

音频与语音处理 · 电气工程与系统科学 2024-02-07 Enting Zhou , You Zhang , Zhiyao Duan

Dynamic facial expression recognition in the wild remains challenging due to data scarcity and long-tail distributions, which hinder models from effectively learning the temporal dynamics of scarce emotions. To address these limitations, we…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Huanzhen Wang , Ziheng Zhou , Jiaqi Song , Li He , Yunshi Lan , Yan Wang , Wenqiang Zhang

Voice conversion (VC) consists of digitally altering the voice of an individual to manipulate part of its content, primarily its identity, while maintaining the rest unchanged. Research in neural VC has accomplished considerable…

声音 · 计算机科学 2021-07-28 Laurent Benaroya , Nicolas Obin , Axel Roebel

Purpose: Emotion is a fundamental component of human communication, shaping understanding, trust, and engagement across domains such as education, healthcare, and mental health. While large language models (LLMs) exhibit strong reasoning…

计算与语言 · 计算机科学 2025-10-15 Yurui Dong , Luozhijie Jin , Yao Yang , Bingjie Lu , Jiaxi Yang , Zhi Liu

Emotional talking head synthesis aims to generate talking portrait videos with vivid expressions. Existing methods still exhibit limitations in control flexibility, motion naturalness, and expression quality. Moreover, currently available…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Yiguo Jiang , Xiaodong Cun , Yong Zhang , Yudian Zheng , Fan Tang , Chi-Man Pun

Despite great advances, achieving high-fidelity emotional voice conversion (EVC) with flexible and interpretable control remains challenging. This paper introduces ClapFM-EVC, a novel EVC framework capable of generating high-quality…

声音 · 计算机科学 2025-05-21 Yu Pan , Yanni Hu , Yuguang Yang , Jixun Yao , Jianhao Ye , Hongbin Zhou , Lei Ma , Jianjun Zhao

In recent years, neural vocoders have surpassed classical speech generation approaches in naturalness and perceptual quality of the synthesized speech. Computationally heavy models like WaveNet and WaveGlow achieve best results, while…

音频与语音处理 · 电气工程与系统科学 2021-02-15 Ahmed Mustafa , Nicola Pia , Guillaume Fuchs

Neural text-to-speech (TTS) generally consists of cascaded architecture with separately optimized acoustic model and vocoder, or end-to-end architecture with continuous mel-spectrograms or self-extracted speech frames as the intermediate…

音频与语音处理 · 电气工程与系统科学 2023-03-09 Ruiqing Xue , Yanqing Liu , Lei He , Xu Tan , Linquan Liu , Edward Lin , Sheng Zhao

Recent approaches in text-to-speech (TTS) synthesis employ neural network strategies to vocode perceptually-informed spectrogram representations directly into listenable waveforms. Such vocoding procedures create a computational bottleneck…

声音 · 计算机科学 2019-07-29 Paarth Neekhara , Chris Donahue , Miller Puckette , Shlomo Dubnov , Julian McAuley

Traditional RGB-based speech generation faces Temporal Granularity Mismatch since fixed camera exposure times inevitably blur the high-frequency articulatory transients essential for rendering emotional speech. To break this ceiling, we…

多媒体 · 计算机科学 2026-05-27 Jingping Fang , Lin Chen , Chenyang Xu , Tong Zhao , Weidong Cai , Xiaoming Chen

Generative deep neural networks are widely used for speech synthesis, but most existing models directly generate waveforms or spectral outputs. Humans, however, produce speech by controlling articulators, which results in the production of…

声音 · 计算机科学 2023-05-10 Gašper Beguš , Alan Zhou , Peter Wu , Gopala K Anumanchipalli