中文
相关论文

相关论文: Gesture-Aware Zero-Shot Speech Recognition for Pat…

200 篇论文

In this study, we address the importance of modeling behavior style in virtual agents for personalized human-agent interaction. We propose a machine learning approach to synthesize gestures, driven by prosodic features and text, in the…

音频与语音处理 · 电气工程与系统科学 2023-05-23 Mireille Fares , Catherine Pelachaud , Nicolas Obin

Dysarthric speech recognition (DSR) presents a formidable challenge due to inherent inter-speaker variability, leading to severe performance degradation when applying DSR models to new dysarthric speakers. Traditional speaker adaptation…

声音 · 计算机科学 2024-09-25 Shiyao Wang , Shiwan Zhao , Jiaming Zhou , Aobo Kong , Yong Qin

This paper proposes a powerful Visual Speech Recognition (VSR) method for multiple languages, especially for low-resource languages that have a limited number of labeled data. Different from previous methods that tried to improve the VSR…

计算机视觉与模式识别 · 计算机科学 2024-01-15 Jeong Hun Yeo , Minsu Kim , Shinji Watanabe , Yong Man Ro

Isolated Sign Language Recognition (ISLR) is crucial for scalable sign language technology, yet language-specific approaches limit current models. To address this, we propose a one-shot learning approach that generalises across languages…

计算与语言 · 计算机科学 2025-02-28 Toon Vandendriessche , Mathieu De Coster , Annelies Lejon , Joni Dambre

Automatic speech recognition (ASR) systems struggle with dysarthric speech due to high inter-speaker variability and slow speaking rates. To address this, we explore dysarthric-to-healthy speech conversion for improved ASR performance. Our…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Karl El Hajal , Enno Hermann , Sevada Hovsepyan , Mathew Magimai. -Doss

Dysarthric speech recognition faces challenges from severity variations and disparities relative to normal speech. Conventional approaches individually fine-tune ASR models pre-trained on normal speech per patient to prevent feature…

声音 · 计算机科学 2025-08-27 Qing Xiao , Yingshan Peng , PeiPei Zhang

Alzheimer's Disease is the most common form of dementia. Automatic detection from speech could help to identify symptoms at early stages, so that preventive actions can be carried out. This research is a contribution to the ADReSSo…

计算与语言 · 计算机科学 2021-11-01 Joan Codina-Filbà , Guillermo Cámbara , Jordi Luque , Mireia Farrús

Interactions based on automatic speech recognition (ASR) have become widely used, with speech input being increasingly utilized to create documents. However, as there is no easy way to distinguish between commands being issued and text…

人机交互 · 计算机科学 2022-08-24 Jun Rekimoto

Audio-Visual Speech Recognition (AVSR) combines auditory and visual speech cues to enhance the accuracy and robustness of speech recognition systems. Recent advancements in AVSR have improved performance in noisy environments compared to…

音频与语音处理 · 电气工程与系统科学 2025-04-29 Zhaofeng Lin , Naomi Harte

In Speech Emotion Recognition (SER), textual data is often used alongside audio signals to address their inherent variability. However, the reliance on human annotated text in most research hinders the development of practical SER systems.…

音频与语音处理 · 电气工程与系统科学 2023-05-30 Yuanchao Li , Zeyu Zhao , Ondrej Klejch , Peter Bell , Catherine Lai

Lip Reading, or Visual Automatic Speech Recognition (V-ASR), is a complex task requiring the interpretation of spoken language exclusively from visual cues, primarily lip movements and facial expressions. This task is especially challenging…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Marshall Thomas , Edward Fish , Richard Bowden

Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart…

音频与语音处理 · 电气工程与系统科学 2024-01-22 Ju Lin , Niko Moritz , Yiteng Huang , Ruiming Xie , Ming Sun , Christian Fuegen , Frank Seide

Millions of people live with cognitive impairment from Alzheimer's disease and related dementias (ADRD). Voice-enabled smart home systems offer promise for supporting daily living but rely on automatic speech recognition (ASR) to transcribe…

人机交互 · 计算机科学 2026-03-02 Michelle Cohn , Alyssa Lanzi , Yui Ishihara , Chen-Nee Chuah , Georgia Zellou , Alyssa Weakley

This paper proposes an interactive system for mobile devices controlled by hand gestures aimed at helping people with visual impairments. This system allows the user to interact with the device by making simple static and dynamic hand…

计算机视觉与模式识别 · 计算机科学 2022-07-27 Samer Alashhab , Antonio Javier Gallego , Miguel Ángel Lozano

Automatic speech recognition (ASR) systems are well known to perform poorly on dysarthric speech. Previous works have addressed this by speaking rate modification to reduce the mismatch with typical speech. Unfortunately, these approaches…

音频与语音处理 · 电气工程与系统科学 2025-01-20 Karl El Hajal , Enno Hermann , Ajinkya Kulkarni , Mathew Magimai. -Doss

Automatic speech recognition (ASR) systems often need to be developed for extremely low-resource languages to serve end-uses such as audio content categorization and search. While universal phone recognition is natural to consider when no…

Structured hand gestures that incorporate visual motions and signs are used in sign language. Sign language is a valuable means of daily communication for individuals who are deaf or have speech impairments, but it is still rare among…

Automatic speech recognition (ASR) research has achieved impressive performance in recent years and has significant potential for enabling access for people with dysarthria (PwD) in augmentative and alternative communication (AAC) and home…

声音 · 计算机科学 2024-06-14 Wing-Zin Leung , Mattias Cross , Anton Ragni , Stefan Goetze

Stuttering -- characterized by involuntary disfluencies such as blocks, prolongations, and repetitions -- is often misinterpreted by automatic speech recognition (ASR) systems, resulting in elevated word error rates and making voice-driven…

声音 · 计算机科学 2025-08-22 Dena Mujtaba , Nihar Mahapatra

Automatic speech recognition (ASR) is a key area in computational linguistics, focusing on developing technologies that enable computers to convert spoken language into text. This field combines linguistics and machine learning. ASR models,…

计算与语言 · 计算机科学 2024-06-27 Anish Saha , A. G. Ramakrishnan