English

PIAVE: A Pose-Invariant Audio-Visual Speaker Extraction Network

Sound 2023-09-14 v1 Multimedia Audio and Speech Processing

Abstract

It is common in everyday spoken communication that we look at the turning head of a talker to listen to his/her voice. Humans see the talker to listen better, so do machines. However, previous studies on audio-visual speaker extraction have not effectively handled the varying talking face. This paper studies how to take full advantage of the varying talking face. We propose a Pose-Invariant Audio-Visual Speaker Extraction Network (PIAVE) that incorporates an additional pose-invariant view to improve audio-visual speaker extraction. Specifically, we generate the pose-invariant view from each original pose orientation, which enables the model to receive a consistent frontal view of the talker regardless of his/her head pose, therefore, forming a multi-view visual input for the speaker. Experiments on the multi-view MEAD and in-the-wild LRS3 dataset demonstrate that PIAVE outperforms the state-of-the-art and is more robust to pose variations.

Keywords

Cite

@article{arxiv.2309.06723,
  title  = {PIAVE: A Pose-Invariant Audio-Visual Speaker Extraction Network},
  author = {Qinghua Liu and Meng Ge and Zhizheng Wu and Haizhou Li},
  journal= {arXiv preprint arXiv:2309.06723},
  year   = {2023}
}

Comments

Interspeech 2023

R2 v1 2026-06-28T12:19:59.093Z