English
Related papers

Related papers: Modelling Lips-State Detection Using CNN for Non-V…

200 papers

This paper presents an efficient visual speech encoder for lip reading. While most recent lip reading studies have been based on the ResNet architecture and have achieved significant success, they are not sufficiently suitable for…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Young-Hu Park , Rae-Hong Park , Hyung-Min Park

Lip sync is a fundamental audio-visual task. However, existing lip sync methods fall short of being robust in the wild. One important cause could be distracting factors on the visual input side, making extracting lip motion information…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Chun Wang

Open-vocabulary semantic segmentation aims to segment an image into semantic regions according to text descriptions, which may not have been seen during training. Recent two-stage methods first generate class-agnostic mask proposals and…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Feng Liang , Bichen Wu , Xiaoliang Dai , Kunpeng Li , Yinan Zhao , Hang Zhang , Peizhao Zhang , Peter Vajda , Diana Marculescu

Silent speech interfaces (SSI) has been an exciting area of recent interest. In this paper, we present a non-invasive silent speech interface that uses inaudible acoustic signals to capture people's lip movements when they speak. We exploit…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-24 Jian Luo , Jianzong Wang , Ning Cheng , Guilin Jiang , Jing Xiao

Vision-language pre-training like CLIP has shown promising performance on various downstream tasks such as zero-shot image classification and image-text retrieval. Most of the existing CLIP-alike works usually adopt relatively large image…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Ying Nie , Wei He , Kai Han , Yehui Tang , Tianyu Guo , Fanyi Du , Yunhe Wang

Open-world detection poses significant challenges, as it requires the detection of any object using either object class labels or free-form texts. Existing related works often use large-scale manual annotated caption datasets for training,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Fanjie Kong , Yanbei Chen , Jiarui Cai , Davide Modolo

Despite the success of large-scale pretrained Vision-Language Models (VLMs) especially CLIP in various open-vocabulary tasks, their application to semantic segmentation remains challenging, producing noisy segmentation maps with…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Mengcheng Lan , Chaofeng Chen , Yiping Ke , Xinjiang Wang , Litong Feng , Wayne Zhang

Vision model have gained increasing attention due to their simplicity and efficiency in Scene Text Recognition (STR) task. However, due to lacking the perception of linguistic knowledge and information, recent vision models suffer from two…

Computer Vision and Pattern Recognition · Computer Science 2023-05-11 Boqiang Zhang , Hongtao Xie , Yuxin Wang , Jianjun Xu , Yongdong Zhang

Interpreting human neural signals to decode static speech intentions such as text or images and dynamic speech intentions such as audio or video is showing great potential as an innovative communication tool. Human communication accompanies…

Artificial Intelligence · Computer Science 2025-01-22 Ji-Ha Park , Seo-Hyun Lee , Soowon Kim , Seong-Whan Lee

Owing to large-scale image-text contrastive training, pre-trained vision language model (VLM) like CLIP shows superior open-vocabulary recognition ability. Most existing open-vocabulary object detectors attempt to utilize the pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Xiangyu Gao , Yu Dai , Benliu Qiu , Lanxiao Wang , Heqian Qiu , Hongliang Li

Recently, convolutional neural networks (CNNs)-based facial landmark detection methods have achieved great success. However, most of existing CNN-based facial landmark detection methods have not attempted to activate multiple correlated…

Computer Vision and Pattern Recognition · Computer Science 2020-11-17 Jun Wan , Zhihui Lai , Linlin Shen , Jie Zhou , Can Gao , Gang Xiao , Xianxu Hou

Top-performing landmark estimation algorithms are based on exploiting the excellent ability of large convolutional neural networks (CNNs) to represent local appearance. However, it is well known that they can only learn weak spatial…

Computer Vision and Pattern Recognition · Computer Science 2022-10-14 Andrés Prados-Torreblanca , José M. Buenaposada , Luis Baumela

Speaker embedding models that utilize neural networks to map utterances to a space where distances reflect similarity between speakers have driven recent progress in the speaker recognition task. However, there is still a significant…

Machine Learning · Computer Science 2019-02-08 Jixuan Wang , Kuan-Chieh Wang , Marc Law , Frank Rudzicz , Michael Brudno

Most current speech technology systems are designed to operate well even in the presence of multiple active speakers. However, most solutions assume that the number of co-current speakers is known. Unfortunately, this information might not…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-02 Midia Yousefi , John H. L. Hansen

Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human communication, co-speech gestures are also temporally aligned…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-28 Zexu Pan , Xinyuan Qian , Shengkui Zhao , Kun Zhou , Bin Ma

Hard of hearing or profoundly deaf people make use of cued speech (CS) as a communication tool to understand spoken language. By delivering cues that are relevant to the phonetic information, CS offers a way to enhance lipreading. In…

Computation and Language · Computer Science 2023-06-16 Sanjana Sankar , Denis Beautemps , Frédéric Elisei , Olivier Perrotin , Thomas Hueber

Lip reading is vital for robots in social settings, improving their ability to understand human communication. This skill allows them to communicate more easily in crowded environments, especially in caregiving and customer service roles.…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Ali Farshian Abbasi , Aghil Yousefi-Koma , Soheil Dehghani Firouzabadi , Parisa Rashidi , Alireza Naeini

When we experience a visual stimulus as beautiful, how much of that experience derives from perceptual computations we cannot describe versus conceptual knowledge we can readily translate into natural language? Disentangling perception from…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Colin Conwell , Christopher Hamblin , Chelsea Boccagno , David Mayo , Jesse Cummings , Leyla Isik , Andrei Barbu

Speech is the most used communication method between humans and it involves the perception of auditory and visual channels. Automatic speech recognition focuses on interpreting the audio signals, although the video can provide information…

Computer Vision and Pattern Recognition · Computer Science 2017-04-27 Adriana Fernandez-Lopez , Oriol Martinez , Federico M. Sukno

Deep Convolutional Neural Networks (CNNs) achieve high accuracy but often rely on purely global, gradient-based optimisation, which can lead to overfitting, redundant filters, and reduced interpretability. To address these limitations, we…

Machine Learning · Computer Science 2025-08-28 Davorin Miličević , Ratko Grbić