中文
相关论文

相关论文: MAAS: Multi-modal Assignation for Active Speaker D…

200 篇论文

Audio-visual learning has demonstrated promising results in many classical speech tasks (e.g., speech separation, automatic speech recognition, wake-word spotting). We believe that introducing visual modality will also benefit speaker…

音频与语音处理 · 电气工程与系统科学 2025-08-01 Ming Cheng , Ming Li

In this paper we address the problem of tracking multiple speakers via the fusion of visual and auditory information. We propose to exploit the complementary nature of these two modalities in order to accurately estimate smooth trajectories…

计算机视觉与模式识别 · 计算机科学 2019-10-30 Yutong Ban , Xavier Alameda-Pineda , Laurent Girin , Radu Horaud

Current Active Speaker Detection (ASD) models achieve great results on AVA-ActiveSpeaker (AVA), using only sound and facial features. Although this approach is applicable in movie setups (AVA), it is not suited for less constrained…

计算机视觉与模式识别 · 计算机科学 2023-03-10 Tiago Roxo , Joana C. Costa , Pedro R. M. Inácio , Hugo Proença

We introduce a new efficient framework, the Unified Context Network (UniCon), for robust active speaker detection (ASD). Traditional methods for ASD usually operate on each candidate's pre-cropped face track separately and do not…

计算机视觉与模式识别 · 计算机科学 2021-08-06 Yuanhang Zhang , Susan Liang , Shuang Yang , Xiao Liu , Zhongqin Wu , Shiguang Shan , Xilin Chen

Traditional reference segmentation tasks have predominantly focused on silent visual scenes, neglecting the integral role of multimodal perception and interaction in human experiences. In this work, we introduce a novel task called…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Yaoting Wang , Peiwen Sun , Dongzhan Zhou , Guangyao Li , Honggang Zhang , Di Hu

This paper addresses the issue of active speaker detection (ASD) in noisy environments and formulates a robust active speaker detection (rASD) problem. Existing ASD approaches leverage both audio and visual modalities, but non-speech sounds…

多媒体 · 计算机科学 2024-04-02 Siva Sai Nagender Vasireddy , Chenxu Zhang , Xiaohu Guo , Yapeng Tian

Interactions involving children span a wide range of important domains from learning to clinical diagnostic and therapeutic contexts. Automated analyses of such interactions are motivated by the need to seek accurate insights and offer…

音频与语音处理 · 电气工程与系统科学 2025-06-13 Anfeng Xu , Kevin Huang , Tiantian Feng , Helen Tager-Flusberg , Shrikanth Narayanan

Replay attacks remain a critical vulnerability for automatic speaker verification systems, particularly in real-time voice assistant applications. In this work, we propose acoustic maps as a novel spatial feature representation for replay…

音频与语音处理 · 电气工程与系统科学 2026-05-21 Michael Neri , Tuomas Virtanen

We present a novel approach to multilingual audio-visual speech recognition tasks by introducing a single model on a multilingual dataset. Motivated by a human cognitive system where humans can intuitively distinguish different languages…

多媒体 · 计算机科学 2023-10-24 Joanna Hong , Se Jin Park , Yong Man Ro

Recent advances in the Active Speaker Detection (ASD) problem build upon a two-stage process: feature extraction and spatio-temporal context aggregation. In this paper, we propose an end-to-end ASD workflow where feature learning and…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Juan Leon Alcazar , Moritz Cordes , Chen Zhao , Bernard Ghanem

Deepfakes are synthetic media generated using deep generative algorithms and have posed a severe societal and political threat. Apart from facial manipulation and synthetic voice, recently, a novel kind of deepfakes has emerged with either…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Vinaya Sree Katamneni , Ajita Rattani

Audiovisual active speaker detection (ASD) addresses the task of determining the speech activity of a candidate speaker given acoustic and visual data. Typically, systems model the temporal correspondence of audiovisual cues, such as the…

多媒体 · 计算机科学 2025-02-11 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world scenarios, visual cues are not always available due to various…

声音 · 计算机科学 2025-04-01 Junjie Li , Ke Zhang , Shuai Wang , Kong Aik Lee , Man-Wai Mak , Haizhou Li

Active Speaker Detection (ASD) aims to identify who is currently speaking in each frame of a video. Most state-of-the-art approaches rely on late fusion to combine visual and audio features, but late fusion often fails to capture…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Yu Wang , Juhyung Ha , Frangil M. Ramirez , Yuchen Wang , David J. Crandall

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these…

计算机视觉与模式识别 · 计算机科学 2022-01-07 Hao Jiang , Calvin Murdock , Vamsi Krishna Ithapu

We introduce a new automatic evaluation method for speaker similarity assessment, that is consistent with human perceptual scores. Modern neural text-to-speech models require a vast amount of clean training data, which is why many solutions…

声音 · 计算机科学 2022-07-04 Deja Kamil , Sanchez Ariadna , Roth Julian , Cotescu Marius

Speaker extraction algorithm relies on the speech sample from the target speaker as the reference point to focus its attention. Such a reference speech is typically pre-recorded. On the other hand, the temporal synchronization between…

音频与语音处理 · 电气工程与系统科学 2021-02-11 Zexu Pan , Ruijie Tao , Chenglin Xu , Haizhou Li

Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Yuqi Hou , Zhongqun Zhang , Nora Horanyi , Jaewon Moon , Yihua Cheng , Hyung Jin Chang

Verifying the identity of a speaker is crucial in modern human-machine interfaces, e.g., to ensure privacy protection or to enable biometric authentication. Classical speaker verification (SV) approaches estimate a fixed-dimensional…

音频与语音处理 · 电气工程与系统科学 2022-06-29 Ahmad Aloradi , Wolfgang Mack , Mohamed Elminshawi , Emanuël A. P. Habets

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…

计算与语言 · 计算机科学 2025-11-17 Tuochao Chen , Bandhav Veluri , Hongyu Gong , Shyamnath Gollakota