中文
相关论文

相关论文: Decoding visemes: improving machine lipreading

200 篇论文

Sign language is a fundamental means of communication for the deaf and hard-of-hearing (DHH) community, enabling nuanced expression through gestures, facial expressions, and body movements. Despite its critical role in facilitating…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Alexander Brettmann , Jakob Grävinghoff , Marlene Rüschoff , Marie Westhues

Sign languages are the language of hearing-impaired people who use visuals like the hand, facial, and body movements for communication. There are different signs and gestures representing alphabets, words, and phrases. Nowadays…

机器学习 · 计算机科学 2024-11-08 Sharvani Srivastava , Sudhakar Singh , Pooja , Shiv Prakash

Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leverages multimodal self-distillation, has demonstrated…

音频与语音处理 · 电气工程与系统科学 2025-02-11 Jing-Xuan Zhang , Tingzhi Mao , Longjiang Guo , Jin Li , Lichen Zhang

The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a limited number of words or phrases, we tackle lip reading as an…

计算机视觉与模式识别 · 计算机科学 2020-11-05 Joon Son Chung , Andrew Senior , Oriol Vinyals , Andrew Zisserman

During a conversation, our brain is responsible for combining information obtained from multiple senses in order to improve our ability to understand the message we are perceiving. Different studies have shown the importance of presenting…

计算机视觉与模式识别 · 计算机科学 2023-11-22 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

Multimodal Vision Language Models (VLMs) have emerged as a transformative topic at the intersection of computer vision and natural language processing, enabling machines to perceive and reason about the world through both visual and textual…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Zongxia Li , Xiyang Wu , Hongyang Du , Fuxiao Liu , Huy Nghiem , Guangyao Shi

As pretrained large language models replace task-specific decoders in speech recognition, a critical question arises: do their text-derived priors make recognition fairer or more biased across demographic groups? We evaluate nine models…

计算与语言 · 计算机科学 2026-04-24 Srishti Ginjala , Eric Fosler-Lussier , Christopher W. Myers , Srinivasan Parthasarathy

Vision-language models (VLMs) achieve remarkable performance through large-scale image-text pretraining. However, their reliance on labeled image datasets limits scalability and leaves vast amounts of unlabeled image data underutilized. To…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Sanghyun Byun , Jung Ick Guack , Mohanad Odema , Baisub Lee , Jacob Song , Woo Seong Chung

This paper proposes a simple and effective approach for automatic recognition of Cued Speech (CS), a visual communication tool that helps people with hearing impairment to understand spoken language with the help of hand gestures that can…

计算与语言 · 计算机科学 2022-04-12 Sanjana Sankar , Denis Beautemps , Thomas Hueber

Video-to-speech synthesis (also known as lip-to-speech) refers to the translation of silent lip movements into the corresponding audio. This task has received an increasing amount of attention due to its self-supervised nature (i.e., can be…

Interactions with virtual assistants typically start with a trigger phrase followed by a command. In this work, we explore the possibility of making these interactions more natural by eliminating the need for a trigger phrase. Our goal is…

Interactions with virtual assistants typically start with a predefined trigger phrase followed by the user command. To make interactions with the assistant more intuitive, we explore whether it is feasible to drop the requirement that users…

Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speakers, background noise and room reverberation remains a…

声音 · 计算机科学 2022-04-11 Guinan Li , Jianwei Yu , Jiajun Deng , Xunying Liu , Helen Meng

This paper explores training medical vision-language models (VLMs) -- where the visual and language inputs are embedded into a common space -- with a particular focus on scenarios where training data is limited, as is often the case in…

计算机视觉与模式识别 · 计算机科学 2023-04-03 Rhydian Windsor , Amir Jamaludin , Timor Kadir , Andrew Zisserman

Speech and audio encoders developed over years of community effort are routinely excluded from video understanding pipelines -- not because they fail, but because benchmarks never required listening. We audit 10 video benchmarks and find…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Geewook Kim , Minjoon Seo

Modern neural language models (LMs) are powerful tools for modeling human sentence production and comprehension, and their internal representations are remarkably well-aligned with representations of language in the human brain. But to…

计算与语言 · 计算机科学 2024-03-27 Chengxu Zhuang , Evelina Fedorenko , Jacob Andreas

Automatic lipreading has major potential impact for speech recognition, supplementing and complementing the acoustic modality. Most attempts at lipreading have been performed on small vocabulary tasks, due to a shortfall of appropriate…

图像与视频处理 · 电气工程与系统科学 2018-05-31 George Sterpu , Naomi Harte

Lip-to-speech involves generating a natural-sounding speech synchronized with a soundless video of a person talking. Despite recent advances, current methods still cannot produce high-quality speech with high levels of intelligibility for…

音频与语音处理 · 电气工程与系统科学 2024-03-29 Yochai Yemini , Aviv Shamsian , Lior Bracha , Sharon Gannot , Ethan Fetaya

In this paper, we address the problem of enhancing the speech of a speaker of interest in a cocktail party scenario when visual information of the speaker of interest is available. Contrary to most previous studies, we do not learn visual…

计算与语言 · 计算机科学 2021-02-04 Giovanni Morrone , Luca Pasa , Vadim Tikhanoff , Sonia Bergamaschi , Luciano Fadiga , Leonardo Badino

Humans can obtain the knowledge of novel visual concepts from language descriptions, and we thus use the few-shot image classification task to investigate whether a machine learning model can have this capability. Our proposed model, LIDE…

计算与语言 · 计算机科学 2022-07-08 Kosuke Nishida , Kyosuke Nishida , Shuichi Nishioka