中文
相关论文

相关论文: Audio-Visual Decision Fusion for WFST-based and se…

200 篇论文

Attention-based sequence-to-sequence models for automatic speech recognition jointly train an acoustic model, language model, and alignment mechanism. Thus, the language model component is only trained on transcribed audio-text pairs. This…

音频与语音处理 · 电气工程与系统科学 2017-12-07 Anjuli Kannan , Yonghui Wu , Patrick Nguyen , Tara N. Sainath , Zhifeng Chen , Rohit Prabhavalkar

Voice disorders negatively impact the quality of daily life in various ways. However, accurately recognizing the category of pathological features from raw audio remains a considerable challenge due to the limited dataset. A promising…

声音 · 计算机科学 2024-10-08 Lipeng Shen , Yifan Xiong , Dongyue Guo , Wei Mo , Lingyu Yu , Hui Yang , Yi Lin

Audio-visual target speech extraction, which aims to extract a certain speaker's speech from the noisy mixture by looking at lip movements, has made significant progress combining time-domain speech separation models and visual feature…

多媒体 · 计算机科学 2023-03-07 Zhongweiyang Xu , Xulin Fan , Mark Hasegawa-Johnson

For a speech-enhancement algorithm, it is highly desirable to simultaneously improve perceptual quality and recognition rate. Thanks to computational costs and model complexities, it is challenging to train a model that effectively…

机器学习 · 计算机科学 2018-02-19 Rasool Fakoor , Xiaodong He , Ivan Tashev , Shuayb Zarar

Speech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and integrating features from various layers of pretrained model…

One of the many tasks facing the typically-developing child language learner is learning to discriminate between the distinctive sounds that make up words in their native language. Here we investigate whether multimodal…

计算与语言 · 计算机科学 2024-07-24 Sophia Zhi , Roger P. Levy , Stephan C. Meylan

Audio-Visual Speech Recognition (AVSR) models have surpassed their audio-only counterparts in terms of performance. However, the interpretability of AVSR systems, particularly the role of the visual modality, remains under-explored. In this…

音频与语音处理 · 电气工程与系统科学 2026-05-06 Aristeidis Papadopoulos , Naomi Harte

In this contribution, we investigate the effectiveness of deep fusion of text and audio features for categorical and dimensional speech emotion recognition (SER). We propose a novel, multistage fusion method where the two information…

机器学习 · 计算机科学 2023-03-27 Andreas Triantafyllopoulos , Uwe Reichel , Shuo Liu , Stephan Huber , Florian Eyben , Björn W. Schuller

Non-intrusive assessment of speech quality and intelligibility is essential when clean reference signals are unavailable. In this work, we propose a multimodal framework that integrates audio features and visual cues to predict PESQ and…

音频与语音处理 · 电气工程与系统科学 2025-06-12 Shafique Ahmed , Ryandhimas E. Zezario , Nasir Saleem , Amir Hussain , Hsin-Min Wang , Yu Tsao

In this study, we propose a novel multi-modal end-to-end neural approach for automated assessment of non-native English speakers' spontaneous speech using attention fusion. The pipeline employs Bi-directional Recurrent Convolutional Neural…

计算与语言 · 计算机科学 2021-11-30 Manraj Singh Grover , Yaman Kumar , Sumit Sarin , Payman Vafaee , Mika Hama , Rajiv Ratn Shah

Autonomous soundscape augmentation systems typically use trained models to pick optimal maskers to effect a desired perceptual change. While acoustic information is paramount to such systems, contextual information, including participant…

声音 · 计算机科学 2024-07-03 Kenneth Ooi , Karn N. Watcharasupat , Bhan Lam , Zhen-Ting Ong , Woon-Seng Gan

We present an approach to Audio-Visual Speech Recognition that builds on a pre-trained Whisper model. To infuse visual information into this audio-only model, we extend it with an AV fusion module and LoRa adapters, one of the most…

声音 · 计算机科学 2025-02-05 Christopher Simic , Korbinian Riedhammer , Tobias Bocklet

We consider the problem of recognizing speech utterances spoken to a device which is generating a known sound waveform; for example, recognizing queries issued to a digital assistant which is generating responses to previous user inputs.…

音频与语音处理 · 电气工程与系统科学 2021-06-03 Nathan Howard , Alex Park , Turaj Zakizadeh Shabestary , Alexander Gruenstein , Rohit Prabhavalkar

For multi-channel speech recognition, speech enhancement techniques such as denoising or dereverberation are conventionally applied as a front-end processor. Deep learning-based front-ends using such techniques require aligned clean and…

声音 · 计算机科学 2020-07-28 Hyeongju Kim , Hyeonseung Lee , Woo Hyun Kang , Hyung Yong Kim , Nam Soo Kim

Comparing spoken segments is a central operation to speech processing. Traditional approaches in this area have favored frame-level dynamic programming algorithms, such as dynamic time warping, because they require no supervision, but they…

计算与语言 · 计算机科学 2023-08-30 Shane Settle

Multimodal speech emotion recognition (SER) has emerged as pivotal for improving human-machine interaction. Researchers are increasingly leveraging both speech and textual information obtained through automatic speech recognition (ASR) to…

人机交互 · 计算机科学 2025-09-24 Jiajun He , Xiaohan Shi , Cheng-Hung Hu , Jinyi Mi , Xingfeng Li , Tomoki Toda

This paper proposes a novel, resource-efficient approach to Visual Speech Recognition (VSR) leveraging speech representations produced by any trained Automatic Speech Recognition (ASR) model. Moving away from the resource-intensive trends…

计算机视觉与模式识别 · 计算机科学 2023-12-18 Hendrik Laux , Emil Mededovic , Ahmed Hallawa , Lukas Martin , Arne Peine , Anke Schmeink

Research in auditory, visual, and audiovisual speech recognition (ASR, VSR, and AVSR, respectively) has traditionally been conducted independently. Even recent self-supervised studies addressing two or all three tasks simultaneously tend to…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Alexandros Haliassos , Rodrigo Mira , Honglie Chen , Zoe Landgraf , Stavros Petridis , Maja Pantic

With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Trevine Oorloff , Surya Koppisetti , Nicolò Bonettini , Divyaraj Solanki , Ben Colman , Yaser Yacoob , Ali Shahriyari , Gaurav Bharaj

Audio-visual speech enhancement system is regarded as one of promising solutions for isolating and enhancing speech of desired speaker. Typical methods focus on predicting clean speech spectrum via a naive convolution neural network based…

音频与语音处理 · 电气工程与系统科学 2022-07-01 Xinmeng Xu , Yang Wang , Jie Jia , Binbin Chen , Dejun Li