English
Related papers

Related papers: Predict-and-Update Network: Audio-Visual Speech Re…

200 papers

Identifying mistakes (i.e., miscues) made while reading aloud is commonly approached post-hoc by comparing automatic speech recognition (ASR) transcriptions to the target reading text. However, post-hoc methods perform poorly when ASR…

Machine Learning · Computer Science 2025-05-30 Griffin Dietz Smith , Dianna Yee , Jennifer King Chen , Leah Findlater

Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to leverage visual information to improve speech recognition…

Sound · Computer Science 2026-01-27 Junli Chen , Changli Tang , Yixuan Li , Guangzhi Sun , Chao Zhang

Comprehending the overall intent of an utterance helps a listener recognize the individual words spoken. Inspired by this fact, we perform a novel study of the impact of explicitly incorporating intent representations as additional…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-22 Swayambhu Nath Ray , Minhua Wu , Anirudh Raju , Pegah Ghahremani , Raghavendra Bilgi , Milind Rao , Harish Arsikere , Ariya Rastrow , Andreas Stolcke , Jasha Droppo

Audio-visual speech enhancement (AV-SE) aims to enhance degraded speech along with extra visual information such as lip videos, and has been shown to be more effective than audio-only speech enhancement. This paper proposes further…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-21 Rui-Chen Zheng , Yang Ai , Zhen-Hua Ling

Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often yield suboptimal…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Chihyun Liu , Jiaxuan Fan , Mingtung Sun , Michael Anthony , Mingsian R. Bai , Yu Tsao

In this paper, we propose long short term memory speech enhancement network (LSTMSE-Net), an audio-visual speech enhancement (AVSE) method. This innovative method leverages the complementary nature of visual and audio information to boost…

Lip reading is used to understand or interpret speech without hearing it, a technique especially mastered by people with hearing difficulties. The ability to lip read enables a person with a hearing impairment to communicate with others and…

Computer Vision and Pattern Recognition · Computer Science 2014-09-05 Ahmad B. A. Hassanat

Active speaker detection (ASD) in egocentric videos presents unique challenges due to unstable viewpoints, motion blur, and off-screen speech sources - conditions under which traditional visual-centric methods degrade significantly. We…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yu Wang , Juhyung Ha , David J. Crandall

Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and end-to-end learning…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Tao Feng , Yifan Xie , Xun Guan , Jiyuan Song , Zhou Liu , Fei Ma , Fei Yu

Accurate recognition of cocktail party speech containing overlapping speakers, noise and reverberation remains a highly challenging task to date. Motivated by the invariance of visual modality to acoustic signal corruption, an audio-visual…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-07 Guinan Li , Jiajun Deng , Mengzhe Geng , Zengrui Jin , Tianzi Wang , Shujie Hu , Mingyu Cui , Helen Meng , Xunying Liu

While significant advancements in artificial intelligence (AI) have catalyzed progress across various domains, its full potential in understanding visual perception remains underexplored. We propose an artificial neural network dubbed…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Ruixing Liang , Xiangyu Zhang , Qiong Li , Lai Wei , Hexin Liu , Avisha Kumar , Kelley M. Kempski Leadingham , Joshua Punnoose , Leibny Paola Garcia , Amir Manbachi

Automatic Speech Recognition (ASR) has achieved remarkable success with deep learning, driving advancements in conversational artificial intelligence, media transcription, and assistive technologies. However, ASR systems still struggle in…

Sound · Computer Science 2026-03-17 Haoyuan Yang , Yue Zhang , Liqiang Jing , John H. L. Hansen

Audio and visual modalities are inherently connected in speech signals: lip movements and facial expressions are correlated with speech sounds. This motivates studies that incorporate the visual modality to enhance an acoustic speech signal…

Sound · Computer Science 2023-06-02 Juan F. Montesinos , Daniel Michelsanti , Gloria Haro , Zheng-Hua Tan , Jesper Jensen

Visual signals can enhance audiovisual speech recognition accuracy by providing additional contextual information. Given the complexity of visual signals, an audiovisual speech recognition model requires robust generalization capabilities…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-20 Yihan Wu , Yifan Peng , Yichen Lu , Xuankai Chang , Ruihua Song , Shinji Watanabe

Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the majority of current research concentrates on the examination…

Sound · Computer Science 2025-04-03 Xinyuan Qian , Jiaran Gao , Yaodan Zhang , Qiquan Zhang , Hexin Liu , Leibny Paola Garcia , Haizhou Li

This paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system. To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy audio speech with a…

Sound · Computer Science 2022-07-14 Joanna Hong , Minsu Kim , Daehun Yoo , Yong Man Ro

Speech enhancement can potentially benefit from the visual information from the target speaker, such as lip movement and facial expressions, because the visual aspect of speech is essentially unaffected by acoustic environment. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-24 Xinmeng Xu , Jianjun Hao

Accurate prediction of the user intent to interact with a voice assistant (VA) on a device (e.g. on the phone) is critical for achieving naturalistic, engaging, and privacy-centric interactions with the VA. To this end, we present a novel…

Computation and Language · Computer Science 2022-10-24 Pranay Dighe , Prateeth Nayak , Oggi Rudovic , Erik Marchi , Xiaochuan Niu , Ahmed Tewfik

Self-supervised sound source localization is usually challenged by the modality inconsistency. In recent studies, contrastive learning based strategies have shown promising to establish such a consistent correspondence between audio and…

Computer Vision and Pattern Recognition · Computer Science 2023-08-10 Tianyu Liu , Peng Zhang , Wei Huang , Yufei Zha , Tao You , Yanning Zhang

Most neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement.…

Sound · Computer Science 2022-10-31 Shulin He , Wei Rao , Jinjiang Liu , Jun Chen , Yukai Ju , Xueliang Zhang , Yannan Wang , Shidong Shang
‹ Prev 1 4 5 6 7 8 10 Next ›