中文
相关论文

相关论文: Evaluating Speech-in-Speech Perception via a Human…

200 篇论文

Does speaking style variation affect humans' ability to distinguish individuals from their voices? How do humans compare with automatic systems designed to discriminate between voices? In this paper, we attempt to answer these questions by…

音频与语音处理 · 电气工程与系统科学 2020-08-11 Amber Afshan , Jody Kreiman , Abeer Alwan

To enable humanoid robots to share our social space we need to develop technology for easy interaction with the robots using multiple modes such as speech, gestures and share our emotions with them. We have targeted this research towards…

机器学习 · 计算机科学 2020-01-14 Shruti Jaiswal , Gora Chand Nandi

Speaker recognition, recognizing speaker identities based on voice alone, enables important downstream applications, such as personalization and authentication. Learning speaker representations, in the context of supervised learning,…

机器学习 · 计算机科学 2022-07-13 Metehan Cekic , Ruirui Li , Zeya Chen , Yuguang Yang , Andreas Stolcke , Upamanyu Madhow

Effectively recognising and applying emotions to interactions is a highly desirable trait for social robots. Implicitly understanding how subjects experience different kinds of actions and objects in the world is crucial for natural HRI…

机器人学 · 计算机科学 2021-03-09 Henrique Siqueira , Alexander Sutherland , Pablo Barros , Mattias Kerzel , Sven Magg , Stefan Wermter

When people try to influence others to do something, they subconsciously adjust their speech to include appropriate emotional information. In order for a robot to influence people in the same way, the robot should be able to imitate the…

Non-native speakers (NNSs) frequently encounter speaking difficulties in multilingual communication, where existing approaches have shown promise in facilitating NNSs' comprehension and participation in real-time communication. However,…

人机交互 · 计算机科学 2026-04-21 Peinuan Qin , Justin Peng , Zhengtao Xu , Jiting Cheng , Zicheng Zhu , Naomi Yamashita , Yi-Chieh Lee

The rise of machine-learning systems that process sensory input has brought with it a rise in comparisons between human and machine perception. But such comparisons face a challenge: Whereas machine perception of some stimulus can often be…

音频与语音处理 · 电气工程与系统科学 2022-08-04 Michael A Lepori , Chaz Firestone

Recognition of social signals, from human facial expressions or prosody of speech, is a popular research topic in human-robot interaction studies. There is also a long line of research in the spoken dialogue community that investigates user…

机器人学 · 计算机科学 2017-06-12 Jekaterina Novikova , Christian Dondrup , Ioannis Papaioannou , Oliver Lemon

As for the humanoid robots, the internal noise, which is generated by motors, fans and mechanical components when the robot is moving or shaking its body, severely degrades the performance of the speech recognition accuracy. In this paper,…

声音 · 计算机科学 2018-08-28 Moa Lee , Joon Hyuk Chang

Voice-based communication is often cited as one of the most `natural' ways in which humans and robots might interact, and the recent availability of accurate automatic speech recognition and intelligible speech synthesis has enabled…

机器人学 · 计算机科学 2022-03-17 Roger K. Moore

To enhance human-robot social interaction, it is essential for robots to process multiple social cues in a complex real-world environment. However, incongruency of input information across modalities is inevitable and could be challenging…

机器人学 · 计算机科学 2023-03-14 Di Fu , Fares Abawi , Hugo Carneiro , Matthias Kerzel , Ziwei Chen , Erik Strahl , Xun Liu , Stefan Wermter

Today, as seen in smart speakers, spoken dialogue technology is rapidly advancing to enable human-like interaction. However, current dialogue systems cannot pay attention not only to the content of speech, but also to the way of speaking…

机器人学 · 计算机科学 2022-10-20 Koki Inoue , Shuichiro Ogake , Hayato Kawamura , Naoki Igo

A robot needs contextual awareness, effective speech production and complementing non-verbal gestures for successful communication in society. In this paper, we present our end-to-end system that tries to enhance the effectiveness of…

机器人学 · 计算机科学 2024-10-01 Bishal Ghosh , Abhinav Dhall , Ekta Singla

Recent work on speech representation models jointly pre-trained with text has demonstrated the potential of improving speech representations by encoding speech and text in a shared space. In this paper, we leverage such shared…

计算与语言 · 计算机科学 2023-10-10 Chung-Ming Chien , Mingjiamei Zhang , Ju-Chieh Chou , Karen Livescu

Speech tokenization is the task of representing speech signals as a sequence of discrete units. Such representations can be later used for various downstream tasks including automatic speech recognition, text-to-speech, etc. More relevant…

声音 · 计算机科学 2024-06-18 Shoval Messica , Yossi Adi

Speech recognition is very challenging in student learning environments that are characterized by significant cross-talk and background noise. To address this problem, we present a bilingual speech recognition system that uses an…

Language technologies have a racial bias, committing greater errors for Black users than for white users. However, little work has evaluated what effect these disparate error rates have on users themselves. The present study aims to…

人机交互 · 计算机科学 2023-02-27 Kimi Wenzel , Nitya Devireddy , Cam Davidson , Geoff Kaufman

Humans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known as selective auditory attention. Recent studies in auditory neuroscience indicate a strong correlation…

音频与语音处理 · 电气工程与系统科学 2023-07-27 Zexu Pan , Marvin Borsdorf , Siqi Cai , Tanja Schultz , Haizhou Li

Older adults living alone have a number of challenges, and robots can help with some of them--by providing reminders, initiating activity, or offering comfort. As part of developing a cat robot with limited assistive functions, we designed…

人机交互 · 计算机科学 2026-05-05 Vivienne Bihe Chi , Claudia B. Rébola , Bertram F. Malle

Testing humanoid robots with users is slow, causes wear, and limits iteration and diversity. Yet screening agents must master conversational timing, prosody, backchannels, and what to attend to in faces and speech for Depression and PTSD.…

机器学习 · 计算机科学 2025-12-11 Filippo Cenacchi , Deborah Richards , Longbing Cao