English
Related papers

Related papers: WhisperNetV2: SlowFast Siamese Network For Lip-Bas…

200 papers

Epileptic seizures cause abnormal brain activity, and their unpredictability can lead to accidents, underscoring the need for long-term seizure prediction. Although seizures can be predicted by analyzing electroencephalogram (EEG) signals,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-03 Guorui Lu , Jing Peng , Bingyuan Huang , Chang Gao , Todor Stefanov , Yong Hao , Qinyu Chen

Lipreading involves using visual data to recognize spoken words by analyzing the movements of the lips and surrounding area. It is a hot research topic with many potential applications, such as human-machine interaction and enhancing audio…

Computer Vision and Pattern Recognition · Computer Science 2024-09-20 Samar Daou , Achraf Ben-Hamadou , Ahmed Rekik , Abdelaziz Kallel

We present Audiovisual SlowFast Networks, an architecture for integrated audiovisual perception. AVSlowFast has Slow and Fast visual pathways that are deeply integrated with a Faster Audio pathway to model vision and sound in a unified…

Computer Vision and Pattern Recognition · Computer Science 2020-03-10 Fanyi Xiao , Yong Jae Lee , Kristen Grauman , Jitendra Malik , Christoph Feichtenhofer

Speculative decoding speeds up autoregressive generation in Large Language Models (LLMs) through a two-step procedure, where a lightweight draft model proposes tokens which the target model then verifies in a single forward pass. Although…

Machine Learning · Computer Science 2026-05-12 Anton Plaksin , Sergei Krutikov , Sergei Skvortsov , Alexander Samarin

Talking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and visual quality. Despite…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Jiadong Wang , Xinyuan Qian , Malu Zhang , Robby T. Tan , Haizhou Li

Due to the widespread deployment of fingerprint/face/speaker recognition systems, attacking deep learning based biometric systems has drawn more and more attention. Previous research mainly studied the attack to the vision-based system,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-08 Jiguo Li , Xinfeng Zhang , Jizheng Xu , Li Zhang , Yue Wang , Siwei Ma , Wen Gao

Lip segmentation plays a crucial role in various domains, such as lip synchronization, lipreading, and diagnostics. However, the effectiveness of supervised lip segmentation is constrained by the availability of lip contour in the training…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Hanie Moghaddasi , Christina Chambers , Sarah N. Mattson , Jeffrey R. Wozniak , Claire D. Coles , Raja Mukherjee , Michael Suttie

Emotion recognition from facial expressions is tremendously useful, especially when coupled with smart devices and wireless multimedia applications. However, the inadequate network bandwidth often limits the spatial resolution of the…

Computer Vision and Pattern Recognition · Computer Science 2017-09-12 Bowen Cheng , Zhangyang Wang , Zhaobin Zhang , Zhu Li , Ding Liu , Jianchao Yang , Shuai Huang , Thomas S. Huang

Creating realistic or stylized facial and lip sync animation is a tedious task. It requires lot of time and skills to sync the lips with audio and convey the right emotion to the character's face. To allow animators to spend more time on…

Graphics · Computer Science 2024-06-03 Bastien Arcelin , Nicolas Chaverou

Speech production is a dynamic procedure, which involved multi human organs including the tongue, jaw and lips. Modeling the dynamics of the vocal tract deformation is a fundamental problem to understand the speech, which is the most common…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-23 Haiyang Liu , Jihan Zhang

Previous studies have confirmed the effectiveness of incorporating visual information into speech enhancement (SE) systems. Despite improved denoising performance, two problems may be encountered when implementing an audio-visual SE (AVSE)…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-19 Shang-Yi Chuang , Yu Tsao , Chen-Chou Lo , Hsin-Min Wang

Many people with some form of hearing loss consider lipreading as their primary mode of day-to-day communication. However, finding resources to learn or improve one's lipreading skills can be challenging. This is further exacerbated in the…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Aditya Agarwal , Bipasha Sen , Rudrabha Mukhopadhyay , Vinay Namboodiri , C. V Jawahar

The lip is a dominant dynamic facial unit when a person is speaking. Detecting lip events is beneficial to speech analysis and support for the hearing impaired. This paper proposes a 3D lip event detection pipeline that automatically…

Computer Vision and Pattern Recognition · Computer Science 2021-11-19 Jie Zhang , Robert B. Fisher

Siamese-network-based self-supervised learning (SSL) suffers from slow convergence and instability in training. To alleviate this, we propose a framework to exploit intermediate self-supervisions in each stage of deep nets, called the…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Ryota Yoshihashi , Shuhei Nishimura , Dai Yonebayashi , Yuya Otsuka , Tomohiro Tanaka , Takashi Miyazaki

Despite the advancement in the domain of audio and audio-visual speech recognition, visual speech recognition systems are still quite under-explored due to the visual ambiguity of some phonemes. In this work, we propose a new lip-reading…

Computer Vision and Pattern Recognition · Computer Science 2021-08-10 Shahd Elashmawy , Marian Ramsis , Hesham M. Eraqi , Farah Eldeshnawy , Hadeel Mabrouk , Omar Abugabal , Nourhan Sakr

Lip-to-speech (L2S) synthesis for Mandarin is a significant challenge, hindered by complex viseme-to-phoneme mappings and the critical role of lexical tones in intelligibility. To address this issue, we propose Lexical Tone-Aware…

Sound · Computer Science 2025-10-01 Kang Yang , Yifan Liang , Fangkun Liu , Zhenping Xie , Chengshi Zheng

Visual Automatic Speech Recognition (V-ASR) is a challenging task that involves interpreting spoken language solely from visual information, such as lip movements and facial expressions. This task is notably challenging due to the absence…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Matthew Kit Khinn Teng , Haibo Zhang , Takeshi Saitoh

Visual speech recognition is a technique to identify spoken content in silent speech videos, which has raised significant attention in recent years. Advancements in data-driven deep learning methods have significantly improved both the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Lei Yang , Junshan Jin , Mingyuan Zhang , Yi He , Bofan Chen , Shilin Wang

Over the last few decades, many aspects of human life have been enhanced with virtual domains, from the advent of digital assistants such as Amazon's Alexa and Apple's Siri to the latest metaverse efforts of the rebranded Meta. These trends…

Computer Vision and Pattern Recognition · Computer Science 2023-03-27 Siddarth Ravichandran , Ondřej Texler , Dimitar Dinev , Hyun Jae Kang

As a key component of talking face generation, lip movements generation determines the naturalness and coherence of the generated talking face video. Prior literature mainly focuses on speech-to-lip generation while there is a paucity in…

Multimedia · Computer Science 2021-12-21 Jinglin Liu , Zhiying Zhu , Yi Ren , Wencan Huang , Baoxing Huai , Nicholas Yuan , Zhou Zhao