中文
相关论文

相关论文: SingingBot: An Avatar-Driven System for Robotic Fa…

200 篇论文

Robots and artificial agents that interact with humans should be able to do so without bias and inequity, but facial perception systems have notoriously been found to work more poorly for certain groups of people than others. In our work,…

计算机视觉与模式识别 · 计算机科学 2022-08-17 Saba Akhyani , Mehryar Abbasi Boroujeni , Mo Chen , Angelica Lim

Accurate facial expression imitation on human-face robots is crucial for achieving natural human-robot interaction. Most existing methods have achieved photorealistic expression imitation through mapping 2D facial landmarks to a robot's…

机器人学 · 计算机科学 2026-03-10 Xu Chen , Rui Gao , Che Sun , Zhehang Liu , Yuwei Wu , Shuo Yang , Yunde Jia

Breath is a significant component in singing performance, which is still underresearched in most singing-related music interfaces. In this paper, we present a multimodal system that detects the learner's singing pitch and breathing states…

人机交互 · 计算机科学 2022-07-06 Ziyue Piao , Gus Xia

Audio-driven portrait animation has made significant advances with diffusion-based models, improving video quality and lipsync accuracy. However, the increasing complexity of these models has led to inefficiencies in training and inference,…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Xuyang Cao , Guoxin Wang , Sheng Shi , Jun Zhao , Yang Yao , Jintao Fei , Minyu Gao , Pei Xie

Detecting singing-voice in polyphonic instrumental music is critical to music information retrieval. To train a robust vocal detector, a large dataset marked with vocal or non-vocal label at frame-level is essential. However, frame-level…

音频与语音处理 · 电气工程与系统科学 2020-08-12 Yuanbo Hou , Frank K. Soong , Jian Luan , Shengchen Li

With the development of the artificial intelligence (AI), the AI applications have influenced and changed people's daily life greatly. Here, a wearable affective robot that integrates the affective robot, social robot, brain wearable, and…

人机交互 · 计算机科学 2018-10-26 Min Chen , Jun Zhou , Guangming Tao , Jun Yang , Long Hu

Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate poorly with human perception. We propose FlowPortrait, a…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Weiting Tan , Andy T. Liu , Ming Tu , Xinghua Qu , Philipp Koehn , Lu Lu

Generating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessitates disentangled…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Xuli Shen , Hua Cai , Dingding Yu , Weilin Shen , Qing Xu , Xiangyang Xue

Recognition of social signals, from human facial expressions or prosody of speech, is a popular research topic in human-robot interaction studies. There is also a long line of research in the spoken dialogue community that investigates user…

机器人学 · 计算机科学 2017-06-12 Jekaterina Novikova , Christian Dondrup , Ioannis Papaioannou , Oliver Lemon

Various parametric representations have been proposed to model the speech signal. While the performance of such vocoders is well-known in the context of speech processing, their extrapolation to singing voice synthesis might not be…

音频与语音处理 · 电气工程与系统科学 2020-06-09 Onur Babacan , Thomas Drugman , Tuomo Raitio , Daniel Erro , Thierry Dutoit

Natural human-robot interaction in complex and unpredictable environments is one of the main research lines in robotics. In typical real-world scenarios, humans are at some distance from the robot and the acquired signals are strongly…

机器人学 · 计算机科学 2015-09-04 Xavier Alameda-Pineda , Radu Horaud

Emotion recognition and generation have emerged as crucial topics in Artificial Intelligence research, playing a significant role in enhancing human-computer interaction within healthcare, customer service, and other fields. Although…

机器学习 · 计算机科学 2025-02-12 Rebecca Mobbs , Dimitrios Makris , Vasileios Argyriou

Previous work on emotion recognition demonstrated a synergistic effect of combining several modalities such as auditory, visual, and transcribed text to estimate the affective state of a speaker. Among these, the linguistic modality is…

计算与语言 · 计算机科学 2019-03-01 Egor Lakomkin , Mohammad Ali Zamani , Cornelius Weber , Sven Magg , Stefan Wermter

In today's high-pressure and isolated society, the demand for emotional support has surged, necessitating innovative solutions. Socially Assistive Robots (SARs) offer a technological approach to providing emotional assistance by leveraging…

机器人学 · 计算机科学 2024-11-11 Leanne Oon Hui Yee , Siew Sui Fun , Thit Sar Zin , Zar Nie Aung , Kian Meng Yap , Jiehan Teoh

Emotions can provide a natural communication modality to complement the existing multi-modal capabilities of social robots, such as text and speech, in many domains. We conducted three online studies with 112, 223, and 151 participants to…

机器人学 · 计算机科学 2022-08-23 Sami Alperen Akgun , Moojan Ghafurian , Mark Crowley , Kerstin Dautenhahn

Voice is an essential modality for human-robot interaction (HRI). The way a robot sounds plays a central role in shaping how humans perceive and engage with it, influencing factors such as intelligibility, understandability, and likability.…

人机交互 · 计算机科学 2026-01-21 Amy Koike , Yuki Okafuji , Sichao Song

The virtual world is being established in which digital humans are created indistinguishable from real humans. Producing their audio-related capabilities is crucial since voice conveys extensive personal characteristics. We aim to create a…

声音 · 计算机科学 2023-05-10 Wei Xue , Yiwen Wang , Qifeng Liu , Yike Guo

Current audio-driven facial animation methods achieve impressive results for short videos but suffer from error accumulation and identity drift when extended to longer durations. Existing methods attempt to mitigate this through external…

We present a wav-to-wav generative model for the task of singing voice conversion from any identity. Our method utilizes both an acoustic model, trained for the task of automatic speech recognition, together with melody extracted features…

音频与语音处理 · 电气工程与系统科学 2020-08-10 Adam Polyak , Lior Wolf , Yossi Adi , Yaniv Taigman

Detecting singing voice deepfakes, or SingFake, involves determining the authenticity and copyright of a singing voice. Existing models for speech deepfake detection have struggled to adapt to unseen attacks in this unique singing voice…

音频与语音处理 · 电气工程与系统科学 2025-06-04 Xuanjun Chen , Haibin Wu , Jyh-Shing Roger Jang , Hung-yi Lee