中文
相关论文

相关论文: Automatic Viseme Vocabulary Construction to Enhanc…

200 篇论文

Despite the remarkable success of Vision-Language Models (VLMs), their performance on a range of complex visual tasks is often hindered by a "visual processing bottleneck": a propensity to lose grounding in visual evidence and exhibit a…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Xinlei Yu , Chengming Xu , Guibin Zhang , Zhangquan Chen , Yudong Zhang , Yongbo He , Peng-Tao Jiang , Jiangning Zhang , Xiaobin Hu , Shuicheng Yan

Visual speech recognition models traditionally consist of two stages, feature extraction and classification. Several deep learning approaches have been recently presented aiming to replace the feature extraction stage by automatically…

计算机视觉与模式识别 · 计算机科学 2019-07-10 Stavros Petridis , Yujiang Wang , Pingchuan Ma , Zuwei Li , Maja Pantic

In this work, we present SynesLM, an unified model which can perform three multimodal language understanding tasks: audio-visual automatic speech recognition(AV-ASR) and visual-aided speech/machine translation(VST/VMT). Unlike previous…

音频与语音处理 · 电气工程与系统科学 2024-08-02 Yichen Lu , Jiaqi Song , Xuankai Chang , Hengwei Bian , Soumi Maiti , Shinji Watanabe

Nonverbal communication (NVC) plays an integral role in human language, but studying NVC in general is challenging because of its broad scope and high variance in interpretation among individuals and cultures. However, mime -- the…

计算与语言 · 计算机科学 2025-08-08 Hyundong Cho , Spencer Lin , Tejas Srinivasan , Michael Saxon , Deuksin Kwon , Natali T. Chavez , Jonathan May

Lipreading, also known as visual speech recognition, aims to identify the speech content from videos by analyzing the visual deformations of lips and nearby areas. One of the significant obstacles for research in this field is the lack of…

计算机视觉与模式识别 · 计算机科学 2021-09-15 Evgeniy Egorov , Vasily Kostyumov , Mikhail Konyk , Sergey Kolesnikov

Universal phoneme recognition typically requires analyzing long speech segments and language-specific patterns. Many speech processing tasks require pure phoneme representations free from contextual influence, which motivated our…

计算与语言 · 计算机科学 2025-08-22 Abdul Rehman , Jian-Jun Zhang , Xiaosong Yang

We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target speech signal. We show…

Brain-computer interface uses brain signals to control external devices without actual control behavior. Recently, speech imagery has been studied for direct communication using language. Speech imagery uses brain signals generated when the…

人机交互 · 计算机科学 2020-12-08 Byeong-Hoo Lee , Byeong-Hee Kwon , Do-Yeun Lee , Ji-Hoon Jeong

In Linguistics, a grapheme is a written unit of a writing system corresponding to a phonological sound. In Natural Language Processing tasks, written language is analysed through two different mediums, word analysis, and character analysis.…

计算与语言 · 计算机科学 2024-04-03 Samuel Rose , Chandrasekhar Kambhampati

In this paper, we address the problem of lip-voice synchronisation in videos containing human face and voice. Our approach is based on determining if the lips motion and the voice in a video are synchronised or not, depending on their…

计算机视觉与模式识别 · 计算机科学 2022-07-01 Venkatesh S. Kadandale , Juan F. Montesinos , Gloria Haro

Many speech segments in movies are re-recorded in a studio during postproduction, to compensate for poor sound quality as recorded on location. Manual alignment of the newly-recorded speech with the original lip movements is a tedious task.…

计算机视觉与模式识别 · 计算机科学 2018-08-21 Tavi Halperin , Ariel Ephrat , Shmuel Peleg

Vision-language models and their adaptations to image segmentation tasks present enormous potential for producing highly accurate and interpretable results. However, implementations based on CLIP and BiomedCLIP are still lagging behind more…

图像与视频处理 · 电气工程与系统科学 2025-09-08 Julia Dietlmeier , Oluwabukola Grace Adegboro , Vayangi Ganepola , Claudia Mazo , Noel E. O'Connor

We introduce a novel and inexpensive approach for the temporal alignment of speech to highly imperfect transcripts from automatic speech recognition (ASR). Transcripts are generated for extended lecture and presentation videos, which in…

声音 · 计算机科学 2007-05-23 Alexander Haubold , John R. Kender

In this paper some of the different techniques used to localize the lips from the face are discussed and compared along with its processing steps. Lip localization is the basic step needed to read the lips for extracting visual information…

计算机视觉与模式识别 · 计算机科学 2020-09-29 S. D. Lalitha , K. K. Thyagharajan

Lipreading involves using visual data to recognize spoken words by analyzing the movements of the lips and surrounding area. It is a hot research topic with many potential applications, such as human-machine interaction and enhancing audio…

计算机视觉与模式识别 · 计算机科学 2024-09-20 Samar Daou , Achraf Ben-Hamadou , Ahmed Rekik , Abdelaziz Kallel

This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to…

Recently reported state-of-the-art results in visual speech recognition (VSR) often rely on increasingly large amounts of video data, while the publicly available transcribed video datasets are limited in size. In this paper, for the first…

Even in highly-developed countries, as many as 15-30\% of the population can only understand texts written using a basic vocabulary. Their understanding of everyday texts is limited, which prevents them from taking an active role in society…

计算与语言 · 计算机科学 2022-09-13 Sanja Stajner , Daniel Ferres , Matthew Shardlow , Kai North , Marcos Zampieri , Horacio Saggion

Speaker extraction seeks to extract the target speech in a multi-talker scenario given an auxiliary reference. Such reference can be auditory, i.e., a pre-recorded speech, visual, i.e., lip movements, or contextual, i.e., phonetic sequence.…

计算机视觉与模式识别 · 计算机科学 2022-10-13 Junjie Li , Meng Ge , Zexu Pan , Longbiao Wang , Jianwu Dang

Automatic measures of similarity between utterances are invaluable for training speech synthesizers, evaluating machine translation, and assessing learner productions. While there exist measures for semantic similarity and prosodic…

计算与语言 · 计算机科学 2024-03-25 Nigel G. Ward , Divette Marco