Shuffle Vision Transformer:面向驾驶员面部表情的轻量、快速且高效识别
计算机视觉与模式识别
2024-09-06 v1
摘要
现有的驾驶员面部表情识别(DFER)方法通常计算量大,使其不适用于实时应用。在本研究中,我们提出了一种基于迁移学习的双架构,命名为 ShuffViT-DFER,它优雅地结合了计算效率与准确性。这是通过利用卷积神经网络(CNN)与视觉 Transformer(ViT)这两种轻量且高效模型的优势来实现的。我们高效地融合所提取的特征,以增强模型准确识别驾驶员面部表情的性能。我们在两个公开基准数据集 KMU-FED 和 KDEF 上的实验结果表明,与现有方法相比,我们所提出的方法在实时应用中具有有效性及优越的性能。
引用
@article{arxiv.2409.03438,
title = {Shuffle Vision Transformer: Lightweight, Fast and Efficient Recognition of Driver Facial Expression},
author = {Ibtissam Saadi and Douglas W. Cunningham and Taleb-ahmed Abdelmalik and Abdenour Hadid and Yassin El Hillali},
journal= {arXiv preprint arXiv:2409.03438},
year = {2024}
}
备注
Accepted for publication in The 6th IEEE International Conference on Artificial Intelligence Circuits and Systems (IEEE AICAS 2024), 5 pages, 3 figures