中文

Kiki 长什么样? speech sounds 与 visual shapes 跨模态关联性在 vision-and-language 模型中的探讨

计算与语言 2024-07-26 v1

摘要

人类在将特定新词与视觉形状匹配时具有明确的跨模态偏好。证据表明,这些偏好在我们的语言处理、语言学习以及信号-意义映射的起源中发挥重要作用。随着视觉-语言 (VLM) 模型在 AI 中的兴起,探索这些模型编码的视觉-语言关联以及它们是否与人类表征一致变得日益重要。基于对人类实验的启发,我们探讨并比较了四个 VLM 在著名的人类跨模态偏好——bouba-kiki 效应上的表现。我们未发现该效应的结论性证据,但建议结果可能取决于模型特征,如架构设计、模型规模和训练细节。我们的发现为讨论 bouba-kiki 效应在人类认知中的起源以及未来开发与人类跨模态关联性一致的 VLM 提供了指导。

关键词

引用

@article{arxiv.2407.17974,
  title  = {What does Kiki look like? Cross-modal associations between speech sounds and visual shapes in vision-and-language models},
  author = {Tessa Verhoef and Kiana Shahrasbi and Tom Kouwenhoven},
  journal= {arXiv preprint arXiv:2407.17974},
  year   = {2024}
}

备注

Appeared at the 13th edition of the Workshop on Cognitive Modeling and Computational Linguistics (CMCL 2024)