English

Towards Logic-Aware Manipulation: A Knowledge Primitive for VLM-Based Assistants in Smart Manufacturing

Robotics 2025-12-15 v1

Abstract

Existing pipelines for vision-language models (VLMs) in robotic manipulation prioritize broad semantic generalization from images and language, but typically omit execution-critical parameters required for contact-rich actions in manufacturing cells. We formalize an object-centric manipulation-logic schema, serialized as an eight-field tuple {\tau}, which exposes object, interface, trajectory, tolerance, and force/impedance information as a first-class knowledge signal between human operators, VLM-based assistants, and robot controllers. We instantiate {\tau} and a small knowledge base (KB) on a 3D-printer spool-removal task in a collaborative cell, and analyze {\tau}-conditioned VLM planning using plan-quality metrics adapted from recent VLM/LLM planning benchmarks, while demonstrating how the same schema supports taxonomy-tagged data augmentation at training time and logic-aware retrieval-augmented prompting at test time as a building block for assistant systems in smart manufacturing enterprises.

Keywords

Cite

@article{arxiv.2512.11275,
  title  = {Towards Logic-Aware Manipulation: A Knowledge Primitive for VLM-Based Assistants in Smart Manufacturing},
  author = {Suchang Chen and Daqiang Guo},
  journal= {arXiv preprint arXiv:2512.11275},
  year   = {2025}
}

Comments

8 pages, 2 figures, submitted to the 2026 IFAC World Congress

R2 v1 2026-07-01T08:21:46.728Z