Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective
Sound
2024-10-01 v1 Computation and Language
Multimedia
Audio and Speech Processing
Abstract
In the field of spoken language processing, audio-visual speech processing is receiving increasing research attention. Key components of this research include tasks such as lip reading, audio-visual speech recognition, and visual-to-speech synthesis. Although significant success has been achieved, theoretical analysis is still insufficient for audio-visual tasks. This paper presents a quantitative analysis based on information theory, focusing on information intersection between different modalities. Our results show that this analysis is valuable for understanding the difficulties of audio-visual processing tasks as well as the benefits that could be obtained by modality integration.
Cite
@article{arxiv.2409.19575,
title = {Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective},
author = {Chen Chen and Xiaolou Li and Zehua Liu and Lantian Li and Dong Wang},
journal= {arXiv preprint arXiv:2409.19575},
year = {2024}
}
Comments
Accepted by ISCSLP2024