中文
相关论文

相关论文: Single-Channel Robot Ego-Speech Filtering during H…

200 篇论文

Self-anthropomorphism in robots manifests itself through their display of human-like characteristics in dialogue, such as expressing preferences and emotions. Our study systematically analyzes self-anthropomorphic expression within various…

计算与语言 · 计算机科学 2024-10-08 Yu Li , Devamanyu Hazarika , Di Jin , Julia Hirschberg , Yang Liu

Gender/ing guides how we view ourselves, the world around us, and each other--including non-humans. Critical voices have raised the alarm about stereotyped gendering in the design of socially embodied artificial agents like voice…

机器人学 · 计算机科学 2022-12-09 Katie Seaborn , Alexa Frank

Voice-based communication is often cited as one of the most `natural' ways in which humans and robots might interact, and the recent availability of accurate automatic speech recognition and intelligible speech synthesis has enabled…

机器人学 · 计算机科学 2022-03-17 Roger K. Moore

We explore the use of inner speech as a mechanism to enhance transparency and trust in social robots for dietary advice. In humans, inner speech structures thought processes and decision-making; in robotics, it improves explainability by…

机器人学 · 计算机科学 2026-02-03 Valerio Belcamino , Alessandro Carfì , Valeria Seidita , Fulvio Mastrogiovanni , Antonio Chella

This research investigates the Statistical Machine Translation approaches to translate speech in real time automatically. Such systems can be used in a pipeline with speech recognition and synthesis software in order to produce a real-time…

计算与语言 · 计算机科学 2015-10-01 Krzysztof Wołk , Krzysztof Marasek

With its strong modeling capacity that comes from a multi-head and multi-layer structure, Transformer is a very powerful model for learning a sequential representation and has been successfully applied to speech separation recently.…

声音 · 计算机科学 2020-10-26 Sanyuan Chen , Yu Wu , Zhuo Chen , Takuya Yoshioka , Shujie Liu , Jinyu Li

This paper aims to achieve single-channel target speech extraction (TSE) in enclosures by solely utilizing distance information. This is the first work that utilizes only distance cues without using speaker physiological information for…

音频与语音处理 · 电气工程与系统科学 2024-12-31 Runwu Shi , Benjamin Yen , Kazuhiro Nakadai

Spoken language interaction is at the heart of interpersonal communication, and people flexibly adapt their speech to different individuals and environments. It is surprising that robots, and by extension other digital devices, are not…

机器人学 · 计算机科学 2024-05-17 Qiaoqiao Ren , Yuanbo Hou , Dick Botteldooren , Tony Belpaeme

The objective of this paper is to learn representations of speaker identity without access to manually annotated data. To do so, we develop a self-supervised learning objective that exploits the natural cross-modal synchrony between faces…

音频与语音处理 · 电气工程与系统科学 2020-05-05 Arsha Nagrani , Joon Son Chung , Samuel Albanie , Andrew Zisserman

Voice-activated systems are integrated into a variety of desktop, mobile, and Internet-of-Things (IoT) devices. However, voice spoofing attacks, such as impersonation and replay attacks, in which malicious attackers synthesize the voice of…

声音 · 计算机科学 2022-05-31 Hanqing Guo , Qiben Yan , Nikolay Ivanov , Ying Zhu , Li Xiao , Eric J. Hunter

Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recently, TF-GridNet has…

音频与语音处理 · 电气工程与系统科学 2025-02-27 Peter Vieting , Simon Berger , Thilo von Neumann , Christoph Boeddeker , Ralf Schlüter , Reinhold Haeb-Umbach

Speech separation in realistic acoustic environments remains challenging because overlapping speakers, background noise, and reverberation must be resolved simultaneously. Although recent time-frequency (TF) domain models have shown strong…

音频与语音处理 · 电气工程与系统科学 2026-05-15 Ui-Hyeop Shin , Hyung-Min Park

Human language is a combination of elemental languages/domains/styles that change across and sometimes within discourses. Language models, which play a crucial role in speech recognizers and machine translation systems, are particularly…

计算与语言 · 计算机科学 2013-03-22 Damianos Karakos , Mark Dredze , Sanjeev Khudanpur

We present ESPnet-SE, which is designed for the quick development of speech enhancement and speech separation systems in a single framework, along with the optional downstream speech recognition module. ESPnet-SE is a new project which…

Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent advances in diffusion and flow-matching models have improved…

音频与语音处理 · 电气工程与系统科学 2025-12-23 Riki Shimizu , Xilin Jiang , Nima Mesgarani

Spoken language is the most natural way for a human to communicate with a robot. It may seem intuitive that a robot should communicate with users in their native language. However, it is not clear if a user's perception of a robot is…

机器人学 · 计算机科学 2023-10-25 Barbara Sienkiewicz , Gabriela Sejnova , Paul Gajewski , Michal Vavrecka , Bipin Indurkhya

Target speech separation refers to extracting the target speaker's speech from mixed signals. Despite the recent advances in deep learning based close-talk speech separation, the applications to real-world are still an open issue. Two main…

声音 · 计算机科学 2020-01-03 Rongzhi Gu , Yuexian Zou

During human-robot interaction (HRI), we want the robot to understand us, and we want to intuitively understand the robot. In order to communicate with and understand the robot, we can leverage interactions, where the human and robot…

机器人学 · 计算机科学 2019-02-05 Dylan P. Losey , Marcia K. O'Malley

In service robotics, there is an interest to identify the user by voice alone. However, in application scenarios where a service robot acts as a waiter or a store clerk, new users are expected to enter the environment frequently. Typically,…

音频与语音处理 · 电气工程与系统科学 2018-09-13 Ivette Vélez , Caleb Rascon , Gibrán Fuentes-Pineda

Many approaches can derive information about a single speaker's identity from the speech by learning to recognize consistent characteristics of acoustic parameters. However, it is challenging to determine identity information when there are…

音频与语音处理 · 电气工程与系统科学 2020-08-07 Hyewon Han , Soo-Whan Chung , Hong-Goo Kang