中文
相关论文

相关论文: EmojiVoice: Towards long-term controllable express…

200 篇论文

While the use of social robots for language teaching has been explored, there remains limited work on a task-specific synthesized voices for language teaching robots. Given that language is a verbal task, this gap may have severe…

机器人学 · 计算机科学 2025-07-31 Paige Tuttösí

Human speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the…

音频与语音处理 · 电气工程与系统科学 2025-08-14 Guanrou Yang , Chen Yang , Qian Chen , Ziyang Ma , Wenxi Chen , Wen Wang , Tianrui Wang , Yifan Yang , Zhikang Niu , Wenrui Liu , Fan Yu , Zhihao Du , Zhifu Gao , ShiLiang Zhang , Xie Chen

Speech is a natural interface for humans to interact with robots. Yet, aligning a robot's voice to its appearance is challenging due to the rich vocabulary of both modalities. Previous research has explored a few labels to describe robots…

人机交互 · 计算机科学 2024-02-09 Pol van Rijn , Silvan Mertes , Kathrin Janowski , Katharina Weitz , Nori Jacoby , Elisabeth André

Emoji have become ubiquitous in written communication, on the Web and beyond. They can emphasize or clarify emotions, add details to conversations, or simply serve decorative purposes. This casual use, however, barely scratches the surface…

计算与语言 · 计算机科学 2024-03-08 Lars Henning Klein , Roland Aydin , Robert West

Large speech generation models are evolving from single-speaker, short sentence synthesis to multi-speaker, long conversation geneartion. Current long-form speech generation models are predominately constrained to dyadic, turn-based…

When people try to influence others to do something, they subconsciously adjust their speech to include appropriate emotional information. In order for a robot to influence people in the same way, the robot should be able to imitate the…

Recent studies have outlined the accessibility challenges faced by blind or visually impaired, and less-literate people, in interacting with social networks, in-spite of facilitating technologies such as monotone text-to-speech (TTS) screen…

社会与信息网络 · 计算机科学 2024-10-28 Suparna De , Ionut Bostan , Nishanth Sastry

The frequent use of Emojis on social media platforms has created a new form of multimodal social interaction. Developing methods for the study and representation of emoji semantics helps to improve future multimodal communication systems.…

计算与语言 · 计算机科学 2018-05-03 Francesco Barbieri , Luis Marujo , Pradeep Karuturi , William Brendel , Horacio Saggion

Robotic assistants reduce the manual efforts being put in by humans in their day-to-day tasks. In this paper, we develop a voice-controlled personal assistant robot. The robot takes the human voice commands by its own built-in microphone.…

机器人学 · 计算机科学 2024-02-07 Vineeth Teeda , K Sujatha , Rakesh Mutukuru

State-of-the-art speech synthesis models try to get as close as possible to the human voice. Hence, modelling emotions is an essential part of Text-To-Speech (TTS) research. In our work, we selected FastSpeech2 as the starting point and…

音频与语音处理 · 电气工程与系统科学 2023-07-04 Daria Diatlova , Vitaly Shutov

Modern TTS systems are capable of creating highly realistic and natural-sounding speech. Despite these developments, the process of customizing TTS voices remains a complex task, mostly requiring the expertise of specialists within the…

Social robots aim to establish long-term bonds with humans through engaging conversation. However, traditional conversational approaches, reliant on scripted interactions, often fall short in maintaining engaging conversations. This paper…

机器人学 · 计算机科学 2024-02-20 Zining Wang , Paul Reisert , Eric Nichols , Randy Gomez

Recent advances in text-to-speech (TTS) have been driven by large, multi-domain speech corpora, yet the expressive potential of audiobook data remains underexamined. We argue that human-narrated audiobooks, particularly fictional works,…

音频与语音处理 · 电气工程与系统科学 2026-04-22 Gaspard Michel , Elena V. Epure , Christophe Cerisara

Generating emotional language is a key step towards building empathetic natural language processing agents. However, a major challenge for this line of research is the lack of large-scale labeled training data, and previous studies are…

计算与语言 · 计算机科学 2018-05-15 Xianda Zhou , William Yang Wang

Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selected from a predefined set, average variation learned from…

音频与语音处理 · 电气工程与系统科学 2021-11-10 Sarath Sivaprasad , Saiteja Kosgi , Vineet Gandhi

Collaborative robots must quickly adapt to their partner's intent and preferences to proactively identify helpful actions. This is especially true in situated settings where human partners can continually teach robots new high-level…

机器人学 · 计算机科学 2025-06-17 Jennifer Grannen , Siddharth Karamcheti , Blake Wulfe , Dorsa Sadigh

Spoken language interaction is at the heart of interpersonal communication, and people flexibly adapt their speech to different individuals and environments. It is surprising that robots, and by extension other digital devices, are not…

机器人学 · 计算机科学 2024-05-17 Qiaoqiao Ren , Yuanbo Hou , Dick Botteldooren , Tony Belpaeme

Novel text-to-speech systems can generate entirely new voices that were not seen during training. However, it remains a difficult task to efficiently create personalized voices from a high-dimensional speaker space. In this work, we use…

Speech synthesis has significantly advanced from statistical methods to deep neural network architectures, leading to various text-to-speech (TTS) models that closely mimic human speech patterns. However, capturing nuances such as emotion…

声音 · 计算机科学 2025-01-14 Shaozuo Zhang , Ambuj Mehrish , Yingting Li , Soujanya Poria

People change their tones of voice, often accompanied by nonverbal vocalizations (NVs) such as laughter and cries, to convey rich emotions. However, most text-to-speech (TTS) systems lack the capability to generate speech with rich…

音频与语音处理 · 电气工程与系统科学 2024-09-18 Haibin Wu , Xiaofei Wang , Sefik Emre Eskimez , Manthan Thakker , Daniel Tompkins , Chung-Hsien Tsai , Canrun Li , Zhen Xiao , Sheng Zhao , Jinyu Li , Naoyuki Kanda
‹ 上一页 1 2 3 10 下一页 ›