English
Related papers

Related papers: EmoShift: Lightweight Activation Steering for Enha…

200 papers

Recent advances in expressive text-to-speech (TTS) have introduced diverse methods based on style embedding extracted from reference speech. However, synthesizing high-quality expressive speech remains challenging. We propose SpotlightTTS,…

Sound · Computer Science 2025-11-20 Nam-Gyu Kim

Speech emotion recognition (SER) systems are constrained by existing datasets that typically cover only 6-10 basic emotions, lack scale and diversity, and face ethical challenges when collecting sensitive emotional states. We introduce…

This paper proposes a unified model to conduct emotion transfer, control and prediction for sequence-to-sequence based fine-grained emotional speech synthesis. Conventional emotional speech synthesis often needs manual labels or reference…

Sound · Computer Science 2020-11-18 Yi Lei , Shan Yang , Lei Xie

We present a meta-learning approach for adaptive text-to-speech (TTS) with few data. During training, we learn a multi-speaker model using a shared conditional WaveNet core and independent learned embeddings for each speaker. The aim of…

Recent advances in expressive text-to-speech (TTS) have introduced diverse methods based on style embedding extracted from reference speech. However, synthesizing high-quality expressive speech remains challenging. We propose Spotlight-TTS,…

Sound · Computer Science 2025-07-01 Nam-Gyu Kim , Deok-Hyeon Cho , Seung-Bin Kim , Seong-Whan Lee

Large language models have revolutionized sign language generation by automatically transforming text into high-quality sign language videos, providing accessible communication for the Deaf community. However, existing LLM-based approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Yanchao Zhao , Jihao Zhu , Yu Liu , Weizhuo Chen , Yuling Yang , Kun Peng

Zero-shot emotion transfer in cross-lingual speech synthesis aims to transfer emotion from an arbitrary speech reference in the source language to the synthetic speech in the target language. Building such a system faces challenges of…

Sound · Computer Science 2023-10-09 Yuke Li , Xinfa Zhu , Yi Lei , Hai Li , Junhui Liu , Danming Xie , Lei Xie

Visual Emotion Analysis (VEA) aims to bridge the affective gap between visual content and human emotional responses. Despite its promise, progress in this field remains limited by the lack of open-source and interpretable datasets. Most…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Yijie Guo , Dexiang Hong , Weidong Chen , Zihan She , Cheng Ye , Xiaojun Chang , Zhendong Mao

The creation of increasingly vivid 3D talking face has become a hot topic in recent years. Currently, most speech-driven works focus on lip synchronisation but neglect to effectively capture the correlations between emotions and facial…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Yihong Lin , Liang Peng , Zhaoxin Fan , Xianjia Wu , Jianqiao Hu , Xiandong Li , Wenxiong Kang , Songju Lei

Depressive and anxiety disorders are widespread, necessitating timely identification and management. Recent advances in Large Language Models (LLMs) offer potential solutions, yet high costs and ethical concerns about training data remain…

Computation and Language · Computer Science 2025-01-28 June M. Liu , Mengxia Gao , Sahand Sabour , Zhuang Chen , Minlie Huang , Tatia M. C. Lee

Zero-shot emotion transfer in cross-lingual speech synthesis refers to generating speech in a target language, where the emotion is expressed based on reference speech from a different source language. However, this task remains challenging…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-13 Tianlun Zuo , Jingbin Hu , Yuke Li , Xinfa Zhu , Hai Li , Ying Yan , Junhui Liu , Danming Xie , Lei Xie

This paper introduces Easy One-Step Text-to-Speech (E1 TTS), an efficient non-autoregressive zero-shot text-to-speech system based on denoising diffusion pretraining and distribution matching distillation. The training of E1 TTS is…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Zhijun Liu , Shuai Wang , Pengcheng Zhu , Mengxiao Bi , Haizhou Li

Despite advances in deep learning, current state-of-the-art speech emotion recognition (SER) systems still have poor performance due to a lack of speech emotion datasets. This paper proposes augmenting SER systems with synthetic emotional…

Sound · Computer Science 2023-01-11 Abdullah Shahid , Siddique Latif , Junaid Qadir

Language models (LMs) automatically learn word embeddings during pre-training on language corpora. Although word embeddings are usually interpreted as feature vectors for individual words, their roles in language model generation remain…

Computation and Language · Computer Science 2024-06-07 Chi Han , Jialiang Xu , Manling Li , Yi Fung , Chenkai Sun , Nan Jiang , Tarek Abdelzaher , Heng Ji

Robustness against temporal variations is important for emotion recognition from speech audio, since emotion is ex-pressed through complex spectral patterns that can exhibit significant local dilation and compression on the time axis…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-10 Eric Guizzo , Tillman Weyde , Jack Barnett Leveson

Head avatars animated by visual signals have gained popularity, particularly in cross-driving synthesis where the driver differs from the animated character, a challenging but highly practical approach. The recently presented MegaPortraits…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Nikita Drobyshev , Antoni Bigata Casademunt , Konstantinos Vougioukas , Zoe Landgraf , Stavros Petridis , Maja Pantic

Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse emotional speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-04 Xiaoxue Gao , Huayun Zhang , Nancy F. Chen

Emotion recognition in conversations (ERC) is challenging due to the multimodal nature of the emotion expression. In this paper, we propose to pretrain a text-based recognition model from unsupervised speech transcripts with LLM guidance.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-22 Soumya Dutta , Sriram Ganapathy

Micro-expression recognition can obtain the real emotion of the individual at the current moment. Although deep learning-based methods, especially Transformer-based methods, have achieved impressive results, these methods have high…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Junbo Wang , Liangyu Fu , Yuke Li , Yining Zhu , Xuecheng Wu , Kun Hu

Emotion detection from text seeks to identify an individual's emotional or mental state - positive, negative, or neutral - based on linguistic cues. While significant progress has been made for English and other high-resource languages,…

Computation and Language · Computer Science 2025-11-11 Abdullah Al Maruf , Aditi Golder , Zakaria Masud Jiyad , Abdullah Al Numan , Tarannum Shaila Zaman