English
Related papers

Related papers: Single-Channel Robot Ego-Speech Filtering during H…

200 papers

Self-anthropomorphism in robots manifests itself through their display of human-like characteristics in dialogue, such as expressing preferences and emotions. Our study systematically analyzes self-anthropomorphic expression within various…

Computation and Language · Computer Science 2024-10-08 Yu Li , Devamanyu Hazarika , Di Jin , Julia Hirschberg , Yang Liu

Gender/ing guides how we view ourselves, the world around us, and each other--including non-humans. Critical voices have raised the alarm about stereotyped gendering in the design of socially embodied artificial agents like voice…

Robotics · Computer Science 2022-12-09 Katie Seaborn , Alexa Frank

Voice-based communication is often cited as one of the most `natural' ways in which humans and robots might interact, and the recent availability of accurate automatic speech recognition and intelligible speech synthesis has enabled…

Robotics · Computer Science 2022-03-17 Roger K. Moore

We explore the use of inner speech as a mechanism to enhance transparency and trust in social robots for dietary advice. In humans, inner speech structures thought processes and decision-making; in robotics, it improves explainability by…

This research investigates the Statistical Machine Translation approaches to translate speech in real time automatically. Such systems can be used in a pipeline with speech recognition and synthesis software in order to produce a real-time…

Computation and Language · Computer Science 2015-10-01 Krzysztof Wołk , Krzysztof Marasek

With its strong modeling capacity that comes from a multi-head and multi-layer structure, Transformer is a very powerful model for learning a sequential representation and has been successfully applied to speech separation recently.…

Sound · Computer Science 2020-10-26 Sanyuan Chen , Yu Wu , Zhuo Chen , Takuya Yoshioka , Shujie Liu , Jinyu Li

This paper aims to achieve single-channel target speech extraction (TSE) in enclosures by solely utilizing distance information. This is the first work that utilizes only distance cues without using speaker physiological information for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-31 Runwu Shi , Benjamin Yen , Kazuhiro Nakadai

Spoken language interaction is at the heart of interpersonal communication, and people flexibly adapt their speech to different individuals and environments. It is surprising that robots, and by extension other digital devices, are not…

Robotics · Computer Science 2024-05-17 Qiaoqiao Ren , Yuanbo Hou , Dick Botteldooren , Tony Belpaeme

The objective of this paper is to learn representations of speaker identity without access to manually annotated data. To do so, we develop a self-supervised learning objective that exploits the natural cross-modal synchrony between faces…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-05 Arsha Nagrani , Joon Son Chung , Samuel Albanie , Andrew Zisserman

Voice-activated systems are integrated into a variety of desktop, mobile, and Internet-of-Things (IoT) devices. However, voice spoofing attacks, such as impersonation and replay attacks, in which malicious attackers synthesize the voice of…

Sound · Computer Science 2022-05-31 Hanqing Guo , Qiben Yan , Nikolay Ivanov , Ying Zhu , Li Xiao , Eric J. Hunter

Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recently, TF-GridNet has…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-27 Peter Vieting , Simon Berger , Thilo von Neumann , Christoph Boeddeker , Ralf Schlüter , Reinhold Haeb-Umbach

Speech separation in realistic acoustic environments remains challenging because overlapping speakers, background noise, and reverberation must be resolved simultaneously. Although recent time-frequency (TF) domain models have shown strong…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-15 Ui-Hyeop Shin , Hyung-Min Park

Human language is a combination of elemental languages/domains/styles that change across and sometimes within discourses. Language models, which play a crucial role in speech recognizers and machine translation systems, are particularly…

Computation and Language · Computer Science 2013-03-22 Damianos Karakos , Mark Dredze , Sanjeev Khudanpur

We present ESPnet-SE, which is designed for the quick development of speech enhancement and speech separation systems in a single framework, along with the optional downstream speech recognition module. ESPnet-SE is a new project which…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-18 Chenda Li , Jing Shi , Wangyou Zhang , Aswin Shanmugam Subramanian , Xuankai Chang , Naoyuki Kamo , Moto Hira , Tomoki Hayashi , Christoph Boeddeker , Zhuo Chen , Shinji Watanabe

Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent advances in diffusion and flow-matching models have improved…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-23 Riki Shimizu , Xilin Jiang , Nima Mesgarani

Spoken language is the most natural way for a human to communicate with a robot. It may seem intuitive that a robot should communicate with users in their native language. However, it is not clear if a user's perception of a robot is…

Robotics · Computer Science 2023-10-25 Barbara Sienkiewicz , Gabriela Sejnova , Paul Gajewski , Michal Vavrecka , Bipin Indurkhya

Target speech separation refers to extracting the target speaker's speech from mixed signals. Despite the recent advances in deep learning based close-talk speech separation, the applications to real-world are still an open issue. Two main…

Sound · Computer Science 2020-01-03 Rongzhi Gu , Yuexian Zou

During human-robot interaction (HRI), we want the robot to understand us, and we want to intuitively understand the robot. In order to communicate with and understand the robot, we can leverage interactions, where the human and robot…

Robotics · Computer Science 2019-02-05 Dylan P. Losey , Marcia K. O'Malley

In service robotics, there is an interest to identify the user by voice alone. However, in application scenarios where a service robot acts as a waiter or a store clerk, new users are expected to enter the environment frequently. Typically,…

Audio and Speech Processing · Electrical Eng. & Systems 2018-09-13 Ivette Vélez , Caleb Rascon , Gibrán Fuentes-Pineda

Many approaches can derive information about a single speaker's identity from the speech by learning to recognize consistent characteristics of acoustic parameters. However, it is challenging to determine identity information when there are…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Hyewon Han , Soo-Whan Chung , Hong-Goo Kang