看与说:面向越极与体中心视角的多模态基准数据集
计算机视觉与模式识别
2025-10-29 v2 计算与语言
机器人学
摘要
我们介绍 Look and Tell,一个用于研究越极视角与体中心视角之间 referential 通信的多模态数据集。使用 Meta Project Aria 智能眼镜和固定摄像机,我们记录了同步的凝视、语音和视频,共有 25 名参与者指导合作伙伴在厨房中识别食材。结合 3D 场景重建,该设置提供了一个基准,用于评估不同的空间表示(2D 与 3D;ego 与 exo)对多模态 grounding 产生的影响。该数据集包含 3.67 小时的录制,涵盖 2,707 条标注详尽的指代表达,旨在推动具身智能体理解并参与 situatted 对话的发展。
引用
@article{arxiv.2510.22672,
title = {Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views},
author = {Anna Deichler and Jonas Beskow},
journal= {arXiv preprint arXiv:2510.22672},
year = {2025}
}
备注
10 pages, 6 figures, 2 tables. Accepted to the NeurIPS 2025 Workshop on SPACE in Vision, Language, and Embodied AI (SpaVLE). Dataset: https://huggingface.co/datasets/annadeichler/KTH-ARIA-referential