English
Related papers

Related papers: NeuroLip: An Event-driven Spatiotemporal Learning …

200 papers

Visual Speech Recognition (VSR) aims to recognize corresponding text by analyzing visual information from lip movements. Due to the high variability and weak information of lip movements, VSR tasks require effectively utilizing any…

Sound · Computer Science 2024-10-23 Zehua Liu , Xiaolou Li , Chen Chen , Li Guo , Lantian Li , Dong Wang

Human language processing relies on the brain's capacity for predictive inference. We present a machine learning framework for decoding neural (EEG) responses to dynamic visual language stimuli in Deaf signers. Using coherence between…

Neurons and Cognition · Quantitative Biology 2025-12-25 Sean C. Borneman , Julia Krebs , Ronnie B. Wilbur , Evie A. Malaia

Reconstructing visual information from brain activity via computer vision technology provides an intuitive understanding of visual neural mechanisms. Despite progress in decoding fMRI data with generative models, achieving accurate…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Shiyi Zhang , Dong Liang , Yihang Zhou

Recent advancements in audio-driven talking face generation have made great progress in lip synchronization. However, current methods often lack sufficient control over facial animation such as speaking style and emotional expression,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Baiqin Wang , Xiangyu Zhu , Fan Shen , Hao Xu , Zhen Lei

We present HighSync, an end-to-end diffusion-based framework for high-fidelity lip synchronization that generates photorealistic talking-face videos aligned with arbitrary input audio. Existing approaches consistently struggle to reconcile…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Saeed Firouzi Daghigh , Majid Iranpour Mobarekeh , Mostafa Alavi , Mehdi Bagheri

This paper introduces an unsupervised compact architecture that can extract features and classify the contents of dynamic scenes from the temporal output of a neuromorphic asynchronous event-based camera. Event-based cameras are clock-less…

Computer Vision and Pattern Recognition · Computer Science 2018-04-26 Germain Haessig , Ryad Benosman

The goal of this project is to develop a limited lip reading algorithm for a subset of the English language. We consider a scenario in which no audio information is available. The raw video is processed and the position of the lips in each…

Computer Vision and Pattern Recognition · Computer Science 2017-08-04 Jithin Donny George , Ronan Keane , Conor Zellmer

Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leverages multimodal self-distillation, has demonstrated…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-11 Jing-Xuan Zhang , Tingzhi Mao , Longjiang Guo , Jin Li , Lichen Zhang

Real-time video dubbing that preserves identity consistency while achieving accurate lip synchronization remains a critical challenge. Existing approaches face a trilemma: diffusion-based methods achieve high visual fidelity but suffer from…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Yue Zhang , Zhizhou Zhong , Minhao Liu , Zhaokang Chen , Bin Wu , Yubin Zeng , Chao Zhan , Yingjie He , Junxin Huang , Wenjiang Zhou

The integration of human-intuitive interactions into autonomous systems has been limited. Traditional Natural Language Processing (NLP) systems struggle with context and intent understanding, severely restricting human-robot interaction.…

Robotics · Computer Science 2025-04-29 Amogh Joshi , Sourav Sanyal , Kaushik Roy

Effective human behavior modeling is critical for successful human-robot interaction. Current state-of-the-art approaches for predicting listening head behavior during dyadic conversations employ continuous-to-discrete representations,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Tri Tung Nguyen Nguyen , Quang Tien Dam , Dinh Tuan Tran , Joo-Ho Lee

Event cameras offer a promising sensing modality for face recognition due to their inherent advantages in illumination robustness and privacy-friendliness. However, because event streams lack the stable photometric appearance relied upon by…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Qingguo Meng , Xingbo Dong , Zhe Jin , Massimo Tistarelli

Humans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known as selective auditory attention. Recent studies in auditory neuroscience indicate a strong correlation…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-27 Zexu Pan , Marvin Borsdorf , Siqi Cai , Tanja Schultz , Haizhou Li

Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Yuqi Hou , Zhongqun Zhang , Nora Horanyi , Jaewon Moon , Yihua Cheng , Hyung Jin Chang

Dense visual prediction tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbounded. While Vision-Language Models (VLMs) like CLIP have…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Junjie Wang , Bin Chen , Yulin Li , Bin Kang , Yichi Chen , Zhuotao Tian

Vision-and-language pretraining (VLP) in the medical field utilizes contrastive learning on image-text pairs to achieve effective transfer across tasks. Yet, current VLP approaches with the masked modeling strategy face two challenges when…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Biao Wu , Yutong Xie , Zeyu Zhang , Minh Hieu Phan , Qi Chen , Ling Chen , Qi Wu

In this paper, we introduce a novel approach to address the task of synthesizing speech from silent videos of any in-the-wild speaker solely based on lip movements. The traditional approach of directly generating speech from lip videos…

Multimedia · Computer Science 2024-03-05 Sindhu Hegde , Rudrabha Mukhopadhyay , C. V. Jawahar , Vinay Namboodiri

Event-based sensors are well suited for real-time processing due to their fast response times and encoding of the sensory data as successive temporal differences. These and other valuable properties, such as a high dynamic range, are…

Machine Learning · Computer Science 2024-10-10 Mark Schöne , Neeraj Mohan Sushma , Jingyue Zhuge , Christian Mayr , Anand Subramoney , David Kappel

High-resolution neural datasets enable foundation models for the next generation of brain-computer interfaces and neurological treatments. The community requires rigorous benchmarks to discriminate between competing modeling approaches, yet…

The study of talking face generation mainly explores the intricacies of synchronizing facial movements and crafting visually appealing, temporally-coherent animations. However, due to the limited exploration of global audio perception,…