English
Related papers

Related papers: EmojiVoice: Towards long-term controllable express…

200 papers

Social robots are starting to become incorporated into daily lives by assisting in the promotion of physical and mental wellbeing. This paper investigates the use of social robots for delivering mindfulness sessions. We created a…

Human-Computer Interaction · Computer Science 2020-06-12 Indu P. Bodala , Nikhil Churamani , Hatice Gunes

The capability of generating speech with specific type of emotion is desired for many applications of human-computer interaction. Cross-speaker emotion transfer is a common approach to generating emotional speech when speech with emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-05 Guangyan Zhang , Ying Qin , Wenjie Zhang , Jialun Wu , Mei Li , Yutao Gai , Feijun Jiang , Tan Lee

Data availability is crucial for advancing artificial intelligence applications, including voice-based technologies. As content creation, particularly in social media, experiences increasing demand, translation and text-to-speech (TTS)…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-27 Ahmet Gunduz , Kamer Ali Yuksel , Kareem Darwish , Golara Javadi , Fabio Minazzi , Nicola Sobieski , Sebastien Bratieres

Cloud-based Large Language Models (LLMs) such as ChatGPT have become increasingly integral to daily operations. Nevertheless, they also introduce privacy concerns: firstly, numerous studies underscore the risks to user privacy posed by…

Computation and Language · Computer Science 2025-03-24 Sam Lin , Wenyue Hua , Zhenting Wang , Mingyu Jin , Lizhou Fan , Yongfeng Zhang

This study proposes FlexiVoice, a text-to-speech (TTS) synthesis system capable of flexible style control with zero-shot voice cloning. The speaking style is controlled by a natural-language instruction and the voice timbre is provided by a…

Sound · Computer Science 2026-01-09 Dekun Chen , Xueyao Zhang , Yuancheng Wang , Kenan Dai , Li Ma , Zhizheng Wu

Zoomorphic robots have the potential to offer companionship and well-being as accessible, low-maintenance alternatives to pet ownership. Many such robots, however, feature limited emotional expression, restricting their potential for rich…

Human-Computer Interaction · Computer Science 2024-10-22 Shaun Macdonald , Robin Bretin , Salma ElSayed

Affective tactile interaction constitutes a fundamental component of human communication. In natural human-human encounters, touch is seldom experienced in isolation; rather, it is inherently multisensory. Individuals not only perceive the…

Robotics · Computer Science 2025-10-09 Qiaoqiao Ren , Tony Belpaeme

Enabling digital humans to express rich emotions has significant applications in dialogue systems, gaming, and other interactive scenarios. While recent advances in talking head synthesis have achieved impressive results in lip…

Artificial Intelligence · Computer Science 2025-10-21 Haidong Xu , Meishan Zhang , Hao Ju , Zhedong Zheng , Erik Cambria , Min Zhang , Hao Fei

Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industrial demand and breakthroughs in deep learning, e.g.,…

Computation and Language · Computer Science 2025-08-26 Tianxin Xie , Yan Rong , Pengfei Zhang , Wenwu Wang , Li Liu

Voice cloning is the task of learning to synthesize the voice of an unseen speaker from a few samples. While current voice cloning methods achieve promising results in Text-to-Speech (TTS) synthesis for a new voice, these approaches lack…

Sound · Computer Science 2021-02-02 Paarth Neekhara , Shehzeen Hussain , Shlomo Dubnov , Farinaz Koushanfar , Julian McAuley

Most existing Zero-Shot Text-To-Speech(ZS-TTS) systems generate the unseen speech based on single prompt, such as reference speech or text descriptions, which limits their flexibility. We propose a customized emotion ZS-TTS system based on…

Sound · Computer Science 2025-05-27 Zhichao Wu , Yueteng Kang , Songjun Cao , Long Ma , Qiulin Li , Qun Yang

The style transfer task in Text-to-Speech refers to the process of transferring style information into text content to generate corresponding speech with a specific style. However, most existing style transfer approaches are either based on…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-01 Wenhao Guan , Yishuang Li , Tao Li , Hukai Huang , Feng Wang , Jiayan Lin , Lingyan Huang , Lin Li , Qingyang Hong

This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeling continuous data by autoregressively generating latent…

Computation and Language · Computer Science 2025-08-27 Zhiliang Peng , Jianwei Yu , Wenhui Wang , Yaoyao Chang , Yutao Sun , Li Dong , Yi Zhu , Weijiang Xu , Hangbo Bao , Zehua Wang , Shaohan Huang , Yan Xia , Furu Wei

This paper introduces EmpathyEar, a pioneering open-source, avatar-based multimodal empathetic chatbot, to fill the gap in traditional text-only empathetic response generation (ERG) systems. Leveraging the advancements of a large language…

Multimedia · Computer Science 2024-06-24 Hao Fei , Han Zhang , Bin Wang , Lizi Liao , Qian Liu , Erik Cambria

Robot-moderated group discussions have the potential to facilitate engaging and productive interactions among human participants. Previous work on topic management in conversational agents has predominantly focused on human engagement and…

Robotics · Computer Science 2025-04-04 Georgios Hadjiantonis , Sarah Gillet , Marynel Vázquez , Iolanda Leite , Fethiye Irmak Dogan

Socially assistive robots (SARs) have shown great potential for supplementing well-being support. However, prior studies have found that existing dialogue pipelines for SARs remain limited in real-time latency, back-channeling, and…

Robotics · Computer Science 2025-07-22 Mengxue Fu , Zhonghao Shi , Minyu Huang , Siqi Liu , Mina Kian , Yirui Song , Maja J. Matarić

Generating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessitates disentangled…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Xuli Shen , Hua Cai , Dingding Yu , Weilin Shen , Qing Xu , Xiangyang Xue

Emoji are a contemporary and extremely popular way to enhance electronic communication. Without rigid semantics attached to them, emoji symbols take on different meanings based on the context of a message. Thus, like the word sense…

Computation and Language · Computer Science 2016-10-26 Sanjaya Wijeratne , Lakshika Balasuriya , Amit Sheth , Derek Doran

We generalized a voice morphing algorithm capable of handling temporally variable, multiple-attributes, and multiple instances. The generalized morphing provides a new strategy for investigating speech diversity. However, excessive…

Human-Computer Interaction · Computer Science 2024-04-23 Hideki Kawahara , Masanori Morise

In this paper, we explore the new design space of extra-linguistic cues inspired by graphical tropes used in graphic novels and animation to enhance the expressiveness of social robots. To achieve this, we identified a set of cues that can…

Robotics · Computer Science 2023-06-29 Amy Koike , Bilge Mutlu