English

CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding

Computer Vision and Pattern Recognition 2025-07-22 v1

Abstract

We introduce CRAFT, a neuro-symbolic framework for interpretable affordance grounding, which identifies the objects in a scene that enable a given action (e.g., "cut"). CRAFT integrates structured commonsense priors from ConceptNet and language models with visual evidence from CLIP, using an energy-based reasoning loop to refine predictions iteratively. This process yields transparent, goal-driven decisions to ground symbolic and perceptual structures. Experiments in multi-object, label-free settings demonstrate that CRAFT enhances accuracy while improving interpretability, providing a step toward robust and trustworthy scene understanding.

Keywords

Cite

@article{arxiv.2507.14426,
  title  = {CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding},
  author = {Zhou Chen and Joe Lin and Sathyanarayanan N. Aakur},
  journal= {arXiv preprint arXiv:2507.14426},
  year   = {2025}
}

Comments

Accepted to NeSy 2025

R2 v1 2026-07-01T04:08:53.507Z