中文

重新审视3D LLM基准:我们真的在测试3D能力吗?

人工智能 2025-06-09 v3

摘要

在这项工作中,我们识别出3D LLM评估中的“2D作弊”问题,即这些任务可能很容易被使用点云渲染图像的视觉语言模型(VLM)解决,从而暴露出对3D LLM独特3D能力评估的无效性。我们测试了VLM在多个3D LLM基准上的性能,并以此为参考,提出了更好地评估真正3D理解能力的原则。我们还主张在评估3D LLM时,明确将3D能力与1D或2D方面分开。代码和数据可在https://github.com/LLM-class-group/Revisiting-3D-LLM-Benchmarks获取。

关键词

引用

@article{arxiv.2502.08503,
  title  = {Revisiting 3D LLM Benchmarks: Are We Really Testing 3D Capabilities?},
  author = {Jiahe Jin and Yanheng He and Mingyan Yang},
  journal= {arXiv preprint arXiv:2502.08503},
  year   = {2025}
}

备注

Accepted to ACL 2025 Findings