中文
相关论文

相关论文: EmojiVoice: Towards long-term controllable express…

200 篇论文

Imbuing machines with the ability to talk has been a longtime pursuit of artificial intelligence (AI) research. From the very beginning, the community has not only aimed to synthesise high-fidelity speech that accurately conveys the…

计算与语言 · 计算机科学 2025-04-11 Andreas Triantafyllopoulos , Björn W. Schuller

While recent advances in Text-to-Speech (TTS) technology produce natural and expressive speech, they lack the option for users to select emotion and control intensity. We propose EmoKnob, a framework that allows fine-grained emotion control…

计算与语言 · 计算机科学 2024-10-02 Haozhe Chen , Run Chen , Julia Hirschberg

Recent development in developing humanoid robot poses new challenges to human-machine interaction communication. A major challenge is to develop robots that can behave like and interact with human in the most natural way possible. This…

机器人学 · 计算机科学 2014-12-03 Ong Sing Goh , Lance Fung

People employ expressive behaviors to effectively communicate and coordinate their actions with others, such as nodding to acknowledge a person glancing at them or saying "excuse me" to pass people in a busy corridor. We would like robots…

Advances in generative AI, speech synthesis, and embodied avatars enable systems that not only assist communication, but can act as proxies on users' behalf. Prior work in HCI has largely focused on systems as external tools, with less…

人机交互 · 计算机科学 2026-03-09 Jieying Zhang , Steeven Villa , Abdallah El Ali

Equipping robotic faces with singing capabilities is crucial for empathetic Human-Robot Interaction. However, existing robotic face driving research primarily focuses on conversations or mimicking static expressions, struggling to meet the…

机器人学 · 计算机科学 2026-01-06 Zhuoxiong Xu , Xuanchen Li , Yuhao Cheng , Fei Xu , Yichao Yan , Xiaokang Yang

This study aimed to develop a system that provides vibrotactile feedback corresponding to the emotional content of text when a communication robot speaks. We used OpenAI's "GPT-4o Mini" for emotion estimation, extracting valence and arousal…

人机交互 · 计算机科学 2024-11-11 Yuki Konishi , Yoshihiro Tanaka

Displaying a written transcript of what a human said (i.e. producing an "automatic speech recognition transcript") is a common feature for smartphone vocal assistants: the utterance produced by a human speaker (e.g. a question) is displayed…

人机交互 · 计算机科学 2025-04-08 Damien Rudaz , Christian Licoppe

In this project, we aim to build a Text-to-Speech system able to produce speech with a controllable emotional expressiveness. We propose a methodology for solving this problem in three main steps. The first is the collection of emotional…

音频与语音处理 · 电气工程与系统科学 2019-07-08 Noé Tits

Large-scale generative models such as GPT and DALL-E have revolutionized the research community. These models not only generate high fidelity outputs, but are also generalists which can solve tasks not explicitly taught. In contrast, speech…

音频与语音处理 · 电气工程与系统科学 2023-10-20 Matthew Le , Apoorv Vyas , Bowen Shi , Brian Karrer , Leda Sari , Rashel Moritz , Mary Williamson , Vimal Manohar , Yossi Adi , Jay Mahadeokar , Wei-Ning Hsu

Recent years have witnessed a trend that large language model (LLM) based text-to-speech (TTS) emerges into the mainstream due to their high naturalness and zero-shot capacity. In this paradigm, speech signals are discretized into token…

Human collaboration with robotics is dependant on the development of a relationship between human and robot, without which performance and utilization can decrease. Emotion and personality conveyance has been shown to enhance robotic…

机器人学 · 计算机科学 2020-10-13 Richard Savery , Lisa Zahray , Gil Weinberg

Artificial speech synthesis has made a great leap in terms of naturalness as recent Text-to-Speech (TTS) systems are capable of producing speech with similar quality to human recordings. However, not all speaking styles are easy to model:…

Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance text-to-speech (TTS). Key requirements include accurate…

Embodied robots which can interact with their environment and neighbours are increasingly being used as a test case to develop Artificial Intelligence. This creates a need for multimodal robot controllers that can operate across different…

机器人学 · 计算机科学 2025-02-05 William Hunt , Sarvapali D. Ramchurn , Mohammad D. Soorati

We propose augmenting the empathetic capacities of social robots by integrating non-verbal cues. Our primary contribution is the design and labeling of four types of empathetic non-verbal cues, abbreviated as SAFE: Speech, Action (gesture),…

机器人学 · 计算机科学 2023-09-01 Yoon Kyung Lee , Yoonwon Jung , Gyuyi Kang , Sowon Hahn

Mindfulness-based therapies have been shown to be effective in improving mental health, and technology-based methods have the potential to expand the accessibility of these therapies. To enable real-time personalized content generation for…

人机交互 · 计算机科学 2024-01-09 Zhonghao Shi , Han Chen , Anna-Maria Velentza , Siqi Liu , Nathaniel Dennler , Allison O'Connell , Maja Matarić

We present BreezyVoice, a Text-to-Speech (TTS) system specifically adapted for Taiwanese Mandarin, highlighting phonetic control abilities to address the unique challenges of polyphone disambiguation in the language. Building upon…

Voice-based communication is often cited as one of the most `natural' ways in which humans and robots might interact, and the recent availability of accurate automatic speech recognition and intelligible speech synthesis has enabled…

机器人学 · 计算机科学 2022-03-17 Roger K. Moore

Understanding the intentions of robots is essential for natural and seamless human-robot collaboration. Ensuring that robots have means for non-verbal communication is a basis for intuitive and implicit interaction. For this, we contribute…

机器人学 · 计算机科学 2024-10-03 Jan Leusmann , Steeven Villa , Thomas Liang , Chao Wang , Albrecht Schmidt , Sven Mayer