English
Related papers

Related papers: One-shot lip-based biometric authentication: exten…

200 papers

Lipreading involves using visual data to recognize spoken words by analyzing the movements of the lips and surrounding area. It is a hot research topic with many potential applications, such as human-machine interaction and enhancing audio…

Computer Vision and Pattern Recognition · Computer Science 2024-09-20 Samar Daou , Achraf Ben-Hamadou , Ahmed Rekik , Abdelaziz Kallel

The task of video-to-speech aims to translate silent video of lip movement to its corresponding audio signal. Previous approaches to this task are generally limited to the case of a single speaker, but a method that accounts for multiple…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-21 Dan Oneata , Adriana Stan , Horia Cucu

Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved. This is due to the lack of available 3D datasets, models, and standard evaluation metrics. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-09 Daniel Cudeiro , Timo Bolkart , Cassidy Laidlaw , Anurag Ranjan , Michael J. Black

Lip synchronization is the task of aligning a speaker's lip movements in video with corresponding speech audio, and it is essential for creating realistic, expressive video content. However, existing methods often rely on reference frames…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Ziqiao Peng , Jiwen Liu , Haoxian Zhang , Xiaoqiang Liu , Songlin Tang , Pengfei Wan , Di Zhang , Hongyan Liu , Jun He

Existing audio-driven facial animation methods face critical challenges, including expression leakage, ineffective subtle expression transfer, and imprecise audio-driven synchronization. We discovered that these issues stem from limitations…

Computer Vision and Pattern Recognition · Computer Science 2024-10-21 Bin Lin , Yanzhen Yu , Jianhao Ye , Ruitao Lv , Yuguang Yang , Ruoye Xie , Pan Yu , Hongbin Zhou

Audio-driven talking head generation is critical for applications such as virtual assistants, video games, and films, where natural lip movements are essential. Despite progress in this field, challenges remain in producing both consistent…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Yucheng Wang , Dan Xu

Deep speech classification tasks, including keyword spotting and speaker verification, are vital in speech-based human-computer interaction. Recently, the security of these technologies has been revealed to be susceptible to backdoor…

Sound · Computer Science 2025-06-11 Wenhan Yao , Fen Xiao , Xiarun Chen , Jia Liu , YongQiang He , Weiping Wen

Imagine a robot is shown new concepts visually together with spoken tags, e.g. "milk", "eggs", "butter". After seeing one paired audio-visual example per class, it is shown a new set of unseen instances of these objects, and asked to pick…

Computation and Language · Computer Science 2019-04-16 Ryan Eloff , Herman A. Engelbrecht , Herman Kamper

We use the term re-identification to refer to the process of recovering the original speaker's identity from anonymized speech outputs. Speaker de-identification systems aim to reduce the risk of re-identification, but most evaluations…

Sound · Computer Science 2025-09-19 Seungmin Seo , Oleg Aulov , P. Jonathon Phillips

Vision-language-action (VLA) models have demonstrated strong semantic understanding and zero-shot generalization, yet most existing systems assume an accurate low-level controller with hand-crafted action "vocabulary" such as end-effector…

The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data. Recent advances in self-supervised learning have shown that effective representations can be learnt from natural cross-modal…

Sound · Computer Science 2020-11-05 Soo-Whan Chung , Hong Goo Kang , Joon Son Chung

Low-Rank Adaptation (LoRA) has emerged as an efficient method for fine-tuning large language models (LLMs) and is widely adopted within the open-source community. However, the decentralized dissemination of LoRA adapters through platforms…

Cryptography and Security · Computer Science 2025-12-23 Linzhi Chen , Yang Sun , Hongru Wei , Yuqi Chen

Lip Reading, or Visual Automatic Speech Recognition (V-ASR), is a complex task requiring the interpretation of spoken language exclusively from visual cues, primarily lip movements and facial expressions. This task is especially challenging…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Marshall Thomas , Edward Fish , Richard Bowden

Modern media firms require automated and efficient methods to identify content that is most engaging and appealing to users. Leveraging a large-scale dataset from Upworthy (a news publisher), which includes 17,681 headline A/B tests, we…

Machine Learning · Computer Science 2024-11-27 Zikun Ye , Hema Yoganarasimhan , Yufeng Zheng

The task of audio-driven portrait animation involves generating a talking head video using an identity image and an audio track of speech. While many existing approaches focus on lip synchronization and video quality, few tackle the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Jian Zhang , Weijian Mai , Zhijun Zhang

Automatic lip-reading (ALR) aims to automatically transcribe spoken content from a speaker's silent lip motion captured in video. Current mainstream lip-reading approaches only use a single visual encoder to model input videos of a single…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 He Wang , Pengcheng Guo , Xucheng Wan , Huan Zhou , Lei Xie

When we speak, the prosody and content of the speech can be inferred from the movement of our lips. In this work, we explore the task of lip to speech synthesis, i.e., learning to generate speech given only the lip movements of a speaker…

Computer Vision and Pattern Recognition · Computer Science 2022-06-29 Christen Millerdurai , Lotfy Abdel Khaliq , Timon Ulrich

Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate poorly with human perception. We propose FlowPortrait, a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Weiting Tan , Andy T. Liu , Ming Tu , Xinghua Qu , Philipp Koehn , Lu Lu

Synthesizing personalized talking faces that uphold and highlight a speaker's unique style while maintaining lip-sync accuracy remains a significant challenge. A primary limitation of existing approaches is the intrinsic confounding of…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Renjie Lu , Xulong Zhang , Xiaoyang Qu , Jianzong Wang , Shangfei Wang

Lip reading is a challenging task that has many potential applications in speech recognition, human-computer interaction, and security systems. However, existing lip reading systems often suffer from low accuracy due to the limitations of…

Computer Vision and Pattern Recognition · Computer Science 2023-07-20 Javad Peymanfard , Vahid Saeedi , Mohammad Reza Mohammadi , Hossein Zeinali , Nasser Mozayani