中文
相关论文

相关论文: LASER: Lip Landmark Assisted Speaker Detection for…

200 篇论文

The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a limited number of words or phrases, we tackle lip reading as an…

计算机视觉与模式识别 · 计算机科学 2018-12-27 Triantafyllos Afouras , Joon Son Chung , Andrew Senior , Oriol Vinyals , Andrew Zisserman

Lip reading, the process of interpreting silent speech from visual lip movements, has gained rising attention for its wide range of realistic applications. Deep learning approaches greatly improve current lip reading systems. However, lip…

人工智能 · 计算机科学 2024-05-03 Linzhi Wu , Xingyu Zhang , Yakun Zhang , Changyan Zheng , Tiejun Liu , Liang Xie , Ye Yan , Erwei Yin

Automatic lip-reading (ALR) aims to automatically transcribe spoken content from a speaker's silent lip motion captured in video. Current mainstream lip-reading approaches only use a single visual encoder to model input videos of a single…

计算机视觉与模式识别 · 计算机科学 2024-05-01 He Wang , Pengcheng Guo , Xucheng Wan , Huan Zhou , Lei Xie

Visual Speech Recognition (VSR) differs from the common perception tasks as it requires deeper reasoning over the video sequence, even by human experts. Despite the recent advances in VSR, current approaches rely on labeled data to fully…

To better model the contextual information and increase the generalization ability of Speech Activity Detection (SAD) system, this paper leverages a multi-lingual Automatic Speech Recognition (ASR) system to perform SAD. Sequence…

声音 · 计算机科学 2021-04-13 Seyyed Saeed Sarfjoo , Srikanth Madikeri , Petr Motlicek

Audio-visual active speaker detection (AV-ASD) aims to identify which visible face is speaking in a scene with one or more persons. Most existing AV-ASD methods prioritize capturing speech-lip correspondence. However, there is a noticeable…

音频与语音处理 · 电气工程与系统科学 2024-04-02 Ruijie Tao , Xinyuan Qian , Rohan Kumar Das , Xiaoxue Gao , Jiadong Wang , Haizhou Li

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these…

计算机视觉与模式识别 · 计算机科学 2022-01-07 Hao Jiang , Calvin Murdock , Vamsi Krishna Ithapu

Event cameras record luminance changes with microsecond resolution, but converting their sparse, asynchronous output into dense tensors that neural networks can exploit remains a core challenge. Conventional histograms or globally-decayed…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Paul Kielty , Timothy Hanley , Peter Corcoran

Most sound event detection (SED) systems perform well on clean datasets but degrade significantly in noisy environments. Language-queried audio source separation (LASS) models show promise for robust SED by separating target events;…

声音 · 计算机科学 2025-08-12 Yuanjian Chen , Yang Xiao , Han Yin , Yadong Guan , Xubo Liu

Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario where the target speech is highly overlapped with the…

声音 · 计算机科学 2023-09-18 Junjie Li , Ruijie Tao , Zexu Pan , Meng Ge , Shuai Wang , Haizhou Li

Active speaker detection is a challenging task in audio-visual scenario understanding, which aims to detect who is speaking in one or more speakers scenarios. This task has received extensive attention as it is crucial in applications such…

计算机视觉与模式识别 · 计算机科学 2023-03-09 Junhua Liao , Haihan Duan , Kanghui Feng , Wanbing Zhao , Yanbing Yang , Liangyin Chen

In this paper, we propose a novel method for speaker adaptation in lip reading, motivated by two observations. Firstly, a speaker's own characteristics can always be portrayed well by his/her few facial images or even a single image with…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Songtao Luo , Shuang Yang , Shiguang Shan , Xilin Chen

Sound Event Detection (SED) is challenging in noisy environments where overlapping sounds obscure target events. Language-queried audio source separation (LASS) aims to isolate the target sound events from a noisy clip. However, this…

音频与语音处理 · 电气工程与系统科学 2025-01-14 Han Yin , Yang Xiao , Jisheng Bai , Rohan Kumar Das

This paper presents a self-supervised method for visual detection of the active speaker in a multi-person spoken interaction scenario. Active speaker detection is a fundamental prerequisite for any artificial cognitive system attempting to…

计算机视觉与模式识别 · 计算机科学 2019-07-19 Kalin Stefanov , Jonas Beskow , Giampiero Salvi

Facial expression recognition (FER) remains a challenging task due to the ambiguity of expressions. The derived noisy labels significantly harm the performance in real-world scenarios. To address this issue, we present a new FER model named…

计算机视觉与模式识别 · 计算机科学 2023-07-21 Zhiyu Wu , Jinshi Cui

The presence of a corresponding talking face has been shown to significantly improve speech intelligibility in noisy conditions and for hearing impaired population. In this paper, we present a system that can generate landmark points of a…

计算机视觉与模式识别 · 计算机科学 2018-04-24 Sefik Emre Eskimez , Ross K Maddox , Chenliang Xu , Zhiyao Duan

Various autonomous applications rely on recognizing specific known landmarks in their environment. For example, Simultaneous Localization And Mapping (SLAM) is an important technique that lays the foundation for many common tasks, such as…

机器人学 · 计算机科学 2023-12-01 Maarten de Backer , Wouter Jansen , Dennis Laurijssen , Ralph Simon , Walter Daems , Jan Steckel

Conventional audio-visual approaches for active speaker detection (ASD) typically rely on visually pre-extracted face tracks and the corresponding single-channel audio to find the speaker in a video. Therefore, they tend to fail every time…

音频与语音处理 · 电气工程与系统科学 2023-12-22 Davide Berghi , Philip J. B. Jackson

Active Speaker Detection (ASD) aims to identify who is speaking in each frame of a video. ASD reasons from audio and visual information from two contexts: long-term intra-speaker context and short-term inter-speaker context. Long-term…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Xizi Wang , Feng Cheng , Gedas Bertasius , David Crandall

We propose a method to address audio-visual target speaker enhancement in multi-talker environments using event-driven cameras. State of the art audio-visual speech separation methods shows that crucial information is the movement of the…

音频与语音处理 · 电气工程与系统科学 2021-02-23 Ander Arriandiaga , Giovanni Morrone , Luca Pasa , Leonardo Badino , Chiara Bartolozzi