English
Related papers

Related papers: Towards Pose-invariant Lip-Reading

200 papers

Although recent years have witnessed significant advancements in image editing thanks to the remarkable progress of text-to-image diffusion models, the problem of non-rigid image editing still presents its complexities and challenges.…

Computer Vision and Pattern Recognition · Computer Science 2024-06-26 Aoyang Liu , Qingnan Fan , Shuai Qin , Hong Gu , Yansong Tang

Many people with some form of hearing loss consider lipreading as their primary mode of day-to-day communication. However, finding resources to learn or improve one's lipreading skills can be challenging. This is further exacerbated in the…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Aditya Agarwal , Bipasha Sen , Rudrabha Mukhopadhyay , Vinay Namboodiri , C. V Jawahar

The aim of this work is to investigate the impact of crossmodal self-supervised pre-training for speech reconstruction (video-to-audio) by leveraging the natural co-occurrence of audio and visual streams in videos. We propose LipSound2…

Sound · Computer Science 2022-09-13 Leyuan Qu , Cornelius Weber , Stefan Wermter

Talking head synthesis, also known as speech-to-lip synthesis, reconstructs the facial motions that align with the given audio tracks. The synthesized videos are evaluated on mainly two aspects, lip-speech synchronization and image…

Machine Learning · Computer Science 2025-03-18 Xulin Fan , Heting Gao , Ziyi Chen , Peng Chang , Mei Han , Mark Hasegawa-Johnson

Natural language processing for document scans and PDFs has the potential to enormously improve the efficiency of business processes. Layout-aware word embeddings such as LayoutLM have shown promise for classification of and information…

Computation and Language · Computer Science 2021-09-02 Anik Saha , Catherine Finegan-Dollak , Ashish Verma

In this paper we present a deep learning architecture for extracting word embeddings for visual speech recognition. The embeddings summarize the information of the mouth region that is relevant to the problem of word recognition, while…

Computer Vision and Pattern Recognition · Computer Science 2017-11-01 Themos Stafylakis , Georgios Tzimiropoulos

The performance of modern face recognition systems is a function of the dataset on which they are trained. Most datasets are largely biased toward "near-frontal" views with benign lighting conditions, negatively effecting recognition…

Computer Vision and Pattern Recognition · Computer Science 2017-04-17 Daniel Crispell , Octavian Biris , Nate Crosswhite , Jeffrey Byrne , Joseph L. Mundy

3D face reconstruction is an important task in the field of computer vision. Although 3D face reconstruction has being developing rapidly in recent years, it is still a challenge for face reconstruction under large pose. That is because…

Computer Vision and Pattern Recognition · Computer Science 2018-11-14 Lei Jiang , XiaoJun Wu , Josef Kittler

Audio-Driven Talking Face Generation aims at generating realistic videos of talking faces, focusing on accurate audio-lip synchronization without deteriorating any identity-related visual details. Recent state-of-the-art methods are based…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazım Kemal Ekenel , Alexander Waibel

Lip sync has emerged as a promising technique for generating mouth movements from audio signals. However, synthesizing a high-resolution and photorealistic virtual news anchor is still challenging. Lack of natural appearance, visual…

Computer Vision and Pattern Recognition · Computer Science 2021-05-06 Ruobing Zheng , Zhou Zhu , Bo Song , Changjiang Ji

Available 3D human pose estimation approaches leverage different forms of strong (2D/3D pose) or weak (multi-view or depth) paired supervision. Barring synthetic or in-studio domains, acquiring such supervision for each new target…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Jogendra Nath Kundu , Siddharth Seth , Anirudh Jamkhandi , Pradyumna YM , Varun Jampani , Anirban Chakraborty , R. Venkatesh Babu

Visual speech recognition is the task to decode the speech content from a video based on visual information, especially the movements of lips. It is also referenced as lipreading. Motivated by two problems existing in lipreading, words with…

Computer Vision and Pattern Recognition · Computer Science 2019-01-14 Jingyun Xiao

Pose-invariant face recognition has become a challenging problem for modern AI-based face recognition systems. It aims at matching a profile face captured in the wild with a frontal face registered in a database. Existing methods perform…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Nikolay Stanishev , Yuhang Lu , Touradj Ebrahimi

Generating 3D human poses from multimodal inputs such as images or text requires models to capture both rich spatial and semantic correspondences. While pose-specific multimodal large language models (MLLMs) have shown promise in this task,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Bao Li , Xiaomei Zhang , Miao Xu , Zhaoxin Fan , Xiangyu Zhu , Zhen Lei

NeRFs have enabled highly realistic synthesis of human faces including complex appearance and reflectance effects of hair and skin. These methods typically require a large number of multi-view input images, making the process hardware…

Recent advances in deep learning methods have increased the performance of face detection and recognition systems. The accuracy of these models relies on the range of variation provided in the training data. Creating a dataset that…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Shubhajit Basak , Hossein Javidnia , Faisal Khan , Rachel McDonnell , Michael Schukat

Lip-reading is the operation of recognizing speech from lip movements. This is a difficult task because the movements of the lips when pronouncing the words are similar for some of them. Viseme is used to describe lip movements during a…

Computer Vision and Pattern Recognition · Computer Science 2021-11-09 Javad Peymanfard , Mohammad Reza Mohammadi , Hossein Zeinali , Nasser Mozayani

We study the problem of facial analysis in videos. We propose a novel weakly supervised learning method that models the video event (expression, pain etc.) as a sequence of automatically mined, discriminative sub-events (eg. onset and…

Computer Vision and Pattern Recognition · Computer Science 2016-04-07 Karan Sikka , Gaurav Sharma , Marian Bartlett

Most of the face recognition works focus on specific modules or demonstrate a research idea. This paper presents a pose-invariant 3D-aided 2D face recognition system (UR2D) that is robust to pose variations as large as 90? by leveraging…

Computer Vision and Pattern Recognition · Computer Science 2017-09-20 Xiang Xu , Pengfei Dou , Ha A. Le , Ioannis A. Kakadiaris

Lip region-of-interest (ROI) is conventionally used for visual input in the lipreading task. Few works have adopted the entire face as visual input because lip-excluded parts of the face are usually considered to be redundant and irrelevant…

Computer Vision and Pattern Recognition · Computer Science 2022-06-03 Jing-Xuan Zhang , Gen-Shun Wan , Jia Pan
‹ Prev 1 4 5 6 7 8 10 Next ›