中文
相关论文

相关论文: Robot Synesthesia: A Sound and Emotion Guided AI P…

200 篇论文

The ability to modulate vocal sounds and generate speech is one of the features which set humans apart from other living beings. The human voice can be characterized by several attributes such as pitch, timbre, loudness, and vocal tone. It…

计算机视觉与模式识别 · 计算机科学 2017-10-30 Poorna Banerjee Dasgupta

This work proposes an interactive art installation "Mood spRing" designed to reflect the mood of the environment through interpretation of language and tone. Mood spRing consists of an AI program that controls an immersive 3D animation of…

人机交互 · 计算机科学 2023-04-03 Nina Marhamati , Sena Clara Creston

Positive human-perception of robots is critical to achieving sustained use of robots in shared environments. One key factor affecting human-perception of robots are their sounds, especially the consequential sounds which robots (as…

机器人学 · 计算机科学 2025-02-05 Aimee Allen , Tom Drummond , Dana Kulić

World models have demonstrated impressive performance on robotic learning tasks. Many such tasks inherently demand multimodal reasoning; for example, filling a bottle with water will lead to visual information alone being ambiguous or…

机器人学 · 计算机科学 2025-12-10 Fan Zhang , Michael Gienger

Prior robot painting and drawing work, such as FRIDA, has focused on decreasing the sim-to-real gap and expanding input modalities for users, but the interaction with these systems generally exists only in the input stages. To support…

机器人学 · 计算机科学 2024-02-22 Peter Schaldenbrand , Gaurav Parmar , Jun-Yan Zhu , James McCann , Jean Oh

Text-writing robots have been used in assistive writing and drawing applications. However, robots do not convey emotional tones in the writing process due to the lack of behaviors humans typically adopt. To examine how people interpret…

人机交互 · 计算机科学 2023-02-14 Yanheng Li , Lin Luoying , Xinyan Li , Yaxuan Mao , Ray Lc

We are witnessing a revolution in conditional image synthesis with the recent success of large scale text-to-image generation methods. This success also opens up new opportunities in controlling the generation and editing process using…

计算机视觉与模式识别 · 计算机科学 2024-05-03 Burak Can Biner , Farrin Marouf Sofian , Umur Berkay Karakaş , Duygu Ceylan , Erkut Erdem , Aykut Erdem

We present a method for automatically producing human-like vocal imitations of sounds: the equivalent of "sketching," but for auditory rather than visual representation. Starting with a simulated model of the human vocal tract, we first try…

图形学 · 计算机科学 2024-09-23 Matthew Caren , Kartik Chandra , Joshua B. Tenenbaum , Jonathan Ragan-Kelley , Karima Ma

To the best of our knowledge, we first present a live system that generates personalized photorealistic talking-head animation only driven by audio signals at over 30 fps. Our system contains three stages. The first stage is a deep neural…

图形学 · 计算机科学 2021-09-27 Yuanxun Lu , Jinxiang Chai , Xun Cao

Indeed, these are exciting times. We are in the heart of a digital renaissance. Automation and computer technology allow engineers and scientists to fabricate processes that amalgamate quality of life. We anticipate much growth in medical…

计算机视觉与模式识别 · 计算机科学 2013-01-16 H. J. Moukalled

Social robots are required not only to understand human intentions but also to effectively communicate their intentions or own internal states to users. This study explores the use of sonification to provide explicit auditory feedback,…

机器人学 · 计算机科学 2024-11-15 Simone Arreghini , Antonio Paolillo , Gabriele Abbate , Alessandro Giusti

The creation of virtual humans increasingly leverages automated synthesis of speech and gestures, enabling expressive, adaptable agents that effectively engage users. However, the independent development of voice and gesture generation…

图形学 · 计算机科学 2025-07-02 Haoyang Du , Kiran Chhatre , Christopher Peters , Brian Keegan , Rachel McDonnell , Cathy Ennis

Sonification is the science of communication of data and events to users through sounds. Auditory icons, earcons, and speech are the common auditory display schemes utilized in sonification, or more specifically in the use of audio to…

Sound effects model design commonly uses digital signal processing techniques with full control ability, but it is difficult to achieve realism within a limited number of parameters. Recently, neural sound effects synthesis methods have…

声音 · 计算机科学 2025-03-13 Yisu Zong , Joshua Reiss

When operating a machine, the operator needs to know some spatial relations, like the relative location of the target or the nearest obstacle. Often, sensors are used to derive this spatial information, and visual displays are deployed as…

人机交互 · 计算机科学 2021-03-01 Tim Ziemer , Nuttawut Nuchprayoon , Holger Schultheis

We present READ Avatars, a 3D-based approach for generating 2D avatars that are driven by audio input with direct and granular control over the emotion. Previous methods are unable to achieve realistic animation due to the many-to-many…

计算机视觉与模式识别 · 计算机科学 2023-03-02 Jack Saunders , Vinay Namboodiri

Direct speech-to-image translation without text is an interesting and useful topic due to the potential applications in human-computer interaction, art creation, computer-aided design. etc. Not to mention that many languages have no writing…

多媒体 · 计算机科学 2020-07-15 Jiguo Li , Xinfeng Zhang , Chuanmin Jia , Jizheng Xu , Li Zhang , Yue Wang , Siwei Ma , Wen Gao

Robots make compulsory machine sounds, known as `consequential sounds', as they move and operate. As robots become more prevalent in workplaces, homes and public spaces, understanding how sounds produced by robots affect human-perceptions…

机器人学 · 计算机科学 2025-02-27 Aimee Allen , Tom Drummond , Dana Kulić

As artificial intelligence shifts from pure tool for delegation toward agentic collaboration, its use in the arts can shift beyond the exploration of machine autonomy toward synergistic co-creation. While our earlier robotic works utilized…

人机交互 · 计算机科学 2026-03-09 Patrick Tresset , Markus Wulfmeier

In this study, we propose AniPortrait, a novel framework for generating high-quality animation driven by audio and a reference portrait image. Our methodology is divided into two stages. Initially, we extract 3D intermediate representations…

计算机视觉与模式识别 · 计算机科学 2024-03-27 Huawei Wei , Zejun Yang , Zhisheng Wang