English
Related papers

Related papers: Resolution limits on visual speech recognition

200 papers

The challenge of talking face generation from speech lies in aligning two different modal information, audio and video, such that the mouth region corresponds to input audio. Previous methods either exploit audio-visual representation…

Computer Vision and Pattern Recognition · Computer Science 2022-11-04 Se Jin Park , Minsu Kim , Joanna Hong , Jeongsoo Choi , Yong Man Ro

End-to-end audio-conditioned latent diffusion models (LDMs) have been widely adopted for audio-driven portrait animation, demonstrating their effectiveness in generating lifelike and high-resolution talking videos. However, direct…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Chunyu Li , Chao Zhang , Weikai Xu , Jingyu Lin , Jinghui Xie , Weiguo Feng , Bingyue Peng , Cunjian Chen , Weiwei Xing

Natural language provides a widely accessible and expressive interface for robotic agents. To understand language in complex environments, agents must reason about the full range of language inputs and their correspondence to the world.…

Computation and Language · Computer Science 2017-10-03 Stephanie Zhou , Alane Suhr , Yoav Artzi

Our objective is an audio-visual model for separating a single speaker from a mixture of sounds such as other speakers and background noise. Moreover, we wish to hear the speaker even when the visual cues are temporarily absent due to…

Computer Vision and Pattern Recognition · Computer Science 2019-07-12 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Video recordings of speech contain correlated audio and visual information, providing a strong signal for speech representation learning from the speaker's lip movements and the produced sound. We introduce Audio-Visual Hidden Unit BERT…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-15 Bowen Shi , Wei-Ning Hsu , Kushal Lakhotia , Abdelrahman Mohamed

Humans learn language by interaction with their environment and listening to other humans. It should also be possible for computational models to learn language directly from speech but so far most approaches require text. We improve on…

Computation and Language · Computer Science 2019-09-25 Danny Merkx , Stefan L. Frank , Mirjam Ernestus

We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets generated by…

Machine Learning · Computer Science 2023-04-06 Alexandros Haliassos , Pingchuan Ma , Rodrigo Mira , Stavros Petridis , Maja Pantic

It has been shown that learning audiovisual features can lead to improved speech recognition performance over audio-only features, especially for noisy speech. However, in many common applications, the visual features are partially or…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-20 Oscar Chang , Otavio Braga , Hank Liao , Dmitriy Serdyuk , Olivier Siohan

Augmented reality display performance depends strongly on features of the human visual system. This is especially true for retinal scan glasses, which use laser beam scanning and transparent holographic optical combiners. Human-centered…

Optics · Physics 2024-12-03 Maximilian Rutz , Pia Neuberger , Simon Pick , Torsten Straßer

Recent advances in speech technologies have produced new tools that can be used to improve the performance and flexibility of speaker recognition While there are few degrees of freedom or alternative methods when using fingerprint or iris…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-28 Marcos Faundez-Zanuy , Enric Monte-Moreno

How to extract more and useful information for single image super resolution is an imperative and difficult problem. Learning-based method is a representative method for such task. However, the results are not so stable as there may exist…

Image and Video Processing · Electrical Eng. & Systems 2020-03-25 Hu Liang , Shengrong Zhao

Recognition of document images have important applications in restoring old and classical texts. The problem involves quality improvement before passing it to a properly trained OCR to get accurate recognition of the text. The image…

Computer Vision and Pattern Recognition · Computer Science 2017-02-01 Ram Krishna Pandey , A G Ramakrishnan

Audio-driven lip sync has recently drawn significant attention due to its widespread application in the multimedia domain. Individuals exhibit distinct lip shapes when speaking the same utterance, attributed to the unique speaking styles of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Weizhi Zhong , Jichang Li , Yinqi Cai , Ming Li , Feng Gao , Liang Lin , Guanbin Li

Lip-reading aims to recognize speech content from videos via visual analysis of speakers' lip movements. This is a challenging task due to the existence of homophemes-words which involve identical or highly similar lip movements, as well as…

Computer Vision and Pattern Recognition · Computer Science 2019-09-04 Chenhao Wang

This paper proposes a powerful Visual Speech Recognition (VSR) method for multiple languages, especially for low-resource languages that have a limited number of labeled data. Different from previous methods that tried to improve the VSR…

Computer Vision and Pattern Recognition · Computer Science 2024-01-15 Jeong Hun Yeo , Minsu Kim , Shinji Watanabe , Yong Man Ro

Visualization in the virtual image formed by dielectric microparticles has been shown to enable the distinction of objects that remain indistinguishable under direct observation. We perform the resolution analysis based on a full…

A main goal in developing video-compression algorithms is to enhance human-perceived visual quality while maintaining file size. But modern video-analysis efforts such as detection and recognition, which are integral to video surveillance…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Mikhail Dremin , Konstantin Kozhemyakov , Ivan Molodetskikh , Malakhov Kirill , Artur Sagitov , Dmitriy Vatolin

Audio super resolution aims to predict the missing high resolution components of the low resolution audio signals. While audio in nature is a continuous signal, current approaches treat it as discrete data (i.e., input is defined on…

Sound · Computer Science 2022-03-31 Jaechang Kim , Yunjoo Lee , Seunghoon Hong , Jungseul Ok

Talking head video compression has advanced with neural rendering and keypoint-based methods, but challenges remain, especially at low bit rates, including handling large head movements, suboptimal lip synchronization, and distorted facial…

Image and Video Processing · Electrical Eng. & Systems 2025-06-17 Riku Takahashi , Ryugo Morita , Jinjia Zhou

Vision model have gained increasing attention due to their simplicity and efficiency in Scene Text Recognition (STR) task. However, due to lacking the perception of linguistic knowledge and information, recent vision models suffer from two…

Computer Vision and Pattern Recognition · Computer Science 2023-05-11 Boqiang Zhang , Hongtao Xie , Yuxin Wang , Jianjun Xu , Yongdong Zhang