中文
相关论文

相关论文: Multilingual Audio-Visual Speech Recognition with …

200 篇论文

The goal of self-supervised learning (SSL) for automatic speech recognition (ASR) is to learn good speech representations from a large amount of unlabeled speech for the downstream ASR task. However, most SSL frameworks do not consider…

计算与语言 · 计算机科学 2022-01-27 Yiming Wang , Jinyu Li , Heming Wang , Yao Qian , Chengyi Wang , Yu Wu

Audio-visual speech recognition (AVSR) attracts a surge of research interest recently by leveraging multimodal signals to understand human speech. Mainstream approaches addressing this task have developed sophisticated architectures and…

音频与语音处理 · 电气工程与系统科学 2023-06-21 Yuchen Hu , Chen Chen , Ruizhe Li , Heqing Zou , Eng Siong Chng

Audio and visual signals complement each other in human speech perception, so do they in speech recognition. The visual hint is less evident than the acoustic hint, but more robust in a complex acoustic environment, as far as speech…

多媒体 · 计算机科学 2022-09-07 Jiadong Wang , Xinyuan Qian , Haizhou Li

Wav2vec2.0 is a popular self-supervised pre-training framework for learning speech representations in the context of automatic speech recognition (ASR). It was shown that wav2vec2.0 has a good robustness against the domain shift, while the…

音频与语音处理 · 电气工程与系统科学 2022-05-10 Qiu-Shi Zhu , Jie Zhang , Zi-Qiang Zhang , Ming-Hui Wu , Xin Fang , Li-Rong Dai

Separating target speech from mixed signals containing flexible speaker quantities presents a challenging task. While existing methods demonstrate strong separation performance and noise robustness, they predominantly assume prior knowledge…

音频与语音处理 · 电气工程与系统科学 2025-07-18 Daning Zhang , Ying Wei

Visual speech recognition models extract visual features in a hierarchical manner. At the lower level, there is a visual front-end with a limited temporal receptive field that processes the raw pixels depicting the lips or faces. At the…

机器学习 · 计算机科学 2023-12-14 Oscar Chang , Hank Liao , Dmitriy Serdyuk , Ankit Shah , Olivier Siohan

Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However, in a more realistic setting, when multiple faces are…

音频与语音处理 · 电气工程与系统科学 2022-05-12 Otavio Braga , Takaki Makino , Olivier Siohan , Hank Liao

We present a frontend for improving robustness of automatic speech recognition (ASR), that jointly implements three modules within a single model: acoustic echo cancellation, speech enhancement, and speech separation. This is achieved by…

音频与语音处理 · 电气工程与系统科学 2021-11-22 Tom O'Malley , Arun Narayanan , Quan Wang , Alex Park , James Walker , Nathan Howard

The goal of this paper is to learn strong lip reading models that can recognise speech in silent videos. Most prior works deal with the open-set visual speech recognition problem by adapting existing automatic speech recognition techniques…

计算机视觉与模式识别 · 计算机科学 2021-12-06 K R Prajwal , Triantafyllos Afouras , Andrew Zisserman

Automatic speech recognition (ASR) needs to be robust to speaker differences. Voice Conversion (VC) modifies speaker characteristics of input speech. This is an attractive feature for ASR data augmentation. In this paper, we demonstrate…

Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer outputs. Among these, word-level timestamp prediction is…

音频与语音处理 · 电气工程与系统科学 2026-04-28 Xulin Fan , Vishal Sunder , Samuel Thomas , Mark Hasegawa-Johnson , Brian Kingsbury , George Saon

Anomaly recognition plays a vital role in surveillance, transportation, healthcare, and public safety. However, most existing approaches rely solely on visual data, making them unreliable under challenging conditions such as occlusion, low…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Amjid Ali , Zulfiqar Ahmad Khan , Altaf Hussain , Muhammad Munsif , Adnan Hussain , Sung Wook Baik

In this work, we propose a new automatic speech recognition (ASR) system based on feature learning and an end-to-end training procedure for air traffic control (ATC) systems. The proposed model integrates the feature learning block,…

声音 · 计算机科学 2021-11-05 Peng Fan , Dongyue Guo , Yi Lin , Bo Yang , Jianwei Zhang

Bootstrapping speech recognition on limited data resources has been an area of active research for long. The recent transition to all-neural models and end-to-end (E2E) training brought along particular challenges as these models are known…

音频与语音处理 · 电气工程与系统科学 2021-06-21 Manuel Giollo , Deniz Gunceler , Yulan Liu , Daniel Willett

The Aduio-visual Speech Recognition (AVSR) which employs both the video and audio information to do Automatic Speech Recognition (ASR) is one of the application of multimodal leaning making ASR system more robust and accuracy. The…

计算机视觉与模式识别 · 计算机科学 2017-03-20 Chunlin Tian , Weijun Ji

This paper introduces an efficient and accurate pipeline for text-dependent speaker verification (TDSV), designed to address the need for high-performance biometric systems. The proposed system incorporates a Fast-Conformer-based ASR module…

声音 · 计算机科学 2024-11-26 Mohammadreza Molavi , Reza Khodadadi

This paper delineates the visual speech recognition (VSR) system introduced by the NPU-ASLP (Team 237) in the second Chinese Continuous Visual Speech Recognition Challenge (CNVSRC 2024), engaging in all four tracks, including the fixed and…

计算机视觉与模式识别 · 计算机科学 2024-09-13 He Wang , Lei Xie

Visual Speech Recognition (VSR) stands at the intersection of computer vision and speech recognition, aiming to interpret spoken content from visual cues. A prominent challenge in VSR is the presence of homophenes-visually similar lip…

人工智能 · 计算机科学 2024-06-19 Young Jin Ahn , Jungwoo Park , Sangha Park , Jonghyun Choi , Kee-Eung Kim

Under noisy conditions, speech recognition systems suffer from high Word Error Rates (WER). In such cases, information from the visual modality comprising the speaker lip movements can help improve the performance. In this work, we propose…

音频与语音处理 · 电气工程与系统科学 2020-01-30 Rohith Aralikatti , Sharad Roy , Abhinav Thanda , Dilip Kumar Margam , Pujitha Appan Kandala , Tanay Sharma , Shankar M Venkatesan

This paper presents the architecture and performance of a novel Multilingual Automatic Speech Recognition (ASR) system developed by the Transsion Speech Team for Track 1 of the MLC-SLM 2025 Challenge. The proposed system comprises three key…

音频与语音处理 · 电气工程与系统科学 2025-08-22 Xiaoxiao Li , An Zhu , Youhai Jiang , Fengjie Zhu