发现并利用 Spelke 片段
摘要
计算机视觉中的片段通常由语义考虑定义,高度依赖于类别特定的惯例。相反,发展心理学表明,人类 perceive the world in terms of Spelke objects——即在物理力作用下可靠地一起移动的物理事物 groupings。Spelke objects thus operate on category-agnostic causal motion relationships which potentially better support tasks like manipulation and planning. 本文首先对 Spelke object 概念进行基准测试,引入包含多种 well-defined Spelke segments 的自然图像数据集 SpelkeBench。接下来,通过构建 SpelkeNet——一种训练用于预测未来运动分布的视觉世界模型来从图像中提取 Spelke segments。SpelkeNet 支持估计两个关键概念:(1) 运动 affordance map,识别在轻触时可能移动的区域;(2) 预期位移 map,捕捉其余场景的移动方式。这些概念用于“统计反事实探针”,其中对高运动-可 affordance 区域施加多样化“虚拟轻触”,并据此生成的预期位移 map 用于定义 Spelke segments 作为相关运动统计的统计聚合。我们发现 SpelkeNet 在 SpelkeBench 上超越了像 SegmentAnything (SAM) 这样的监督基线。最后,我们展示了 Spelke 概念在下游任务中具有实际用途,在 3DEditBench 基准上用于物理对象操控时,可在多种 off-the-shelf 对象操控模型中实现优异性能。
引用
@article{arxiv.2507.16038,
title = {Discovering and using Spelke segments},
author = {Rahul Venkatesh and Klemen Kotar and Lilian Naing Chen and Seungwoo Kim and Luca Thomas Wheeler and Jared Watrous and Ashley Xu and Gia Ancone and Wanhee Lee and Honglin Chen and Daniel Bear and Stefan Stojanov and Daniel Yamins},
journal= {arXiv preprint arXiv:2507.16038},
year = {2025}
}
备注
Project page at: https://neuroailab.github.io/spelke_net