English

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants

Artificial Intelligence 2025-07-24 v2 Computation and Language Computer Vision and Pattern Recognition

Abstract

Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when fine-grained spatial reasoning and precise object state detection are required. In this work, we explore LEGO Co-builder, a hybrid benchmark combining real-world LEGO assembly logic with programmatically generated multimodal scenes. The dataset captures stepwise visual states and procedural instructions, allowing controlled evaluation of instruction-following, object detection, and state detection. We introduce a unified framework and assess leading VLMs such as GPT-4o, Gemini, and Qwen-VL, under zero-shot and fine-tuned settings. Our results reveal that even advanced models like GPT-4o struggle with fine-grained assembly tasks, with a maximum F1 score of just 40.54\% on state detection, highlighting gaps in fine-grained visual understanding. We release the benchmark, codebase, and generation pipeline to support future research on multimodal assembly assistants grounded in real-world workflows.

Keywords

Cite

@article{arxiv.2507.05515,
  title  = {LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants},
  author = {Haochen Huang and Jiahuan Pei and Mohammad Aliannejadi and Xin Sun and Moonisa Ahsan and Chuang Yu and Zhaochun Ren and Pablo Cesar and Junxiao Wang},
  journal= {arXiv preprint arXiv:2507.05515},
  year   = {2025}
}

Comments

This version has been anonymized for double-blind review