English
Related papers

Related papers: Multichannel Robot Speech Recognition Database: MC…

200 papers

Advancing human-robot communication is crucial for autonomous systems operating in dynamic environments, where accurate real-time interpretation of human signals is essential. RoboCup provides a compelling scenario for testing these…

Some speech recognition tasks, such as automatic speech recognition (ASR), are approaching or have reached human performance in many reported metrics. Yet, they continue to struggle in complex, real-world, situations, such as with distanced…

Computation and Language · Computer Science 2025-07-31 Paige Tuttösí , Mantaj Dhillon , Luna Sang , Shane Eastwood , Poorvi Bhatia , Quang Minh Dinh , Avni Kapoor , Yewon Jin , Angelica Lim

This paper describes our RoyalFlush system for the track of multi-speaker automatic speech recognition (ASR) in the M2MeT challenge. We adopted the serialized output training (SOT) based multi-speakers ASR system with large-scale simulation…

Sound · Computer Science 2022-02-25 Shuaishuai Ye , Peiyao Wang , Shunfei Chen , Xinhui Hu , Xinkang Xu

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-13 Otavio Braga , Olivier Siohan

Emotion recognition and touch gesture decoding are crucial for advancing human-robot interaction (HRI), especially in social environments where emotional cues and tactile perception play important roles. However, many humanoid robots, such…

Human-Computer Interaction · Computer Science 2025-01-03 Yuanbo Hou , Qiaoqiao Ren , Wenwu Wang , Dick Botteldooren

Multimodal speech emotion recognition (SER) has emerged as pivotal for improving human-machine interaction. Researchers are increasingly leveraging both speech and textual information obtained through automatic speech recognition (ASR) to…

Human-Computer Interaction · Computer Science 2025-09-24 Jiajun He , Xiaohan Shi , Cheng-Hung Hu , Jinyi Mi , Xingfeng Li , Tomoki Toda

Maintaining situational awareness (SA) is critical in human-robot teams. Yet, under high workload and dynamic conditions, operators often experience SA gaps. Automated detection of SA gaps could provide timely assistance for operators.…

Most approaches to multi-talker overlapped speech separation and recognition assume that the number of simultaneously active speakers is given, but in realistic situations, it is typically unknown. To cope with this, we extend an iterative…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-22 Thilo von Neumann , Christoph Boeddeker , Lukas Drude , Keisuke Kinoshita , Marc Delcroix , Tomohiro Nakatani , Reinhold Haeb-Umbach

For many small- and medium-vocabulary tasks, audio-visual speech recognition can significantly improve the recognition rates compared to audio-only systems. However, there is still an ongoing debate regarding the best combination strategy…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-29 Wentao Yu , Steffen Zeiler , Dorothea Kolossa

Multi-robot systems hold significant promise for social environments such as homes and hospitals, yet existing multi-robot works treat robots as functionally identical, overlooking how robots individual identity shape user perception and…

Robotics · Computer Science 2026-04-15 Shaid Hasan , Breenice Lee , Sujan Sarker , Tariq Iqbal

Recent development of speech processing, such as speech recognition, speaker diarization, etc., has inspired numerous applications of speech technologies. The meeting scenario is one of the most valuable and, at the same time, most…

Sound · Computer Science 2022-02-28 Fan Yu , Shiliang Zhang , Yihui Fu , Lei Xie , Siqi Zheng , Zhihao Du , Weilong Huang , Pengcheng Guo , Zhijie Yan , Bin Ma , Xin Xu , Hui Bu

Human-in-the-loop approaches are of great importance for robot applications. In the presented study, we implemented a multimodal human-robot interaction (HRI) scenario, in which a simulated robot communicates with its human partner through…

Robotics · Computer Science 2024-03-14 Su Kyoung Kim , Michael Maurus , Mathias Trampler , Marc Tabie , Elsa Andrea Kirchner

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is suffering from the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Qiushi Zhu , Jie Zhang , Yu Gu , Yuchen Hu , Lirong Dai

In the development of acoustic signal processing algorithms, their evaluation in various acoustic environments is of utmost importance. In order to advance evaluation in realistic and reproducible scenarios, several high-quality acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-15 Thomas Dietzen , Randall Ali , Maja Taseska , Toon van Waterschoot

Automatic Speech Recognition (ASR) systems have proliferated over the recent years to the point that free platforms such as YouTube now provide speech recognition services. Given the wide selection of ASR systems, we contribute to the field…

Previous work on emotion recognition demonstrated a synergistic effect of combining several modalities such as auditory, visual, and transcribed text to estimate the affective state of a speaker. Among these, the linguistic modality is…

Computation and Language · Computer Science 2019-03-01 Egor Lakomkin , Mohammad Ali Zamani , Cornelius Weber , Sven Magg , Stefan Wermter

Police departments around the world use two-way radio for coordination. These broadcast police communications (BPC) are a unique source of information about everyday police activity and emergency response. Yet BPC are not transcribed, and…

Sound · Computer Science 2024-09-18 Tejes Srivastava , Ju-Chieh Chou , Priyank Shroff , Karen Livescu , Christopher Graziul

Multidimensional human understanding is essential for real-world applications such as film analysis and virtual digital humans, yet current LVLM benchmarks largely focus on single-task settings and lack fine-grained, human-centric…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Kangkang Wang , Qinting Jiang , Wanping Zhang , Bowen Ren , Shengzhao Wen

Visual context has been shown to be useful for automatic speech recognition (ASR) systems when the speech signal is noisy or corrupted. Previous work, however, has only demonstrated the utility of visual context in an unrealistic setting,…

Computation and Language · Computer Science 2020-10-20 Tejas Srinivasan , Ramon Sanabria , Florian Metze , Desmond Elliott

With the development of deep learning, automatic speech recognition (ASR) has made significant progress. To further enhance the performance of ASR, revising recognition results is one of the lightweight but efficient manners. Various…

Computation and Language · Computer Science 2024-06-14 Yi-Wei Wang , Ke-Han Lu , Kuan-Yu Chen