English

GeminiPainter's sequence-formed pipeline comprised of perception, cognition, planning, and action stages

Robotics 2026-08-01 v1

Abstract

We present an autonomous robotic portrait-generation system combining real-time face detection, AI-based sketch generation, and robotic drawing. The system captures video frames, extracts facial regions, converts them into minimalist single-line sketches using the Gemini Vision API, optimizes stroke order through graph-based path planning, and executes smooth trajectories on a 6-DoF collaborative manipulator. This perception-cognition-action pipeline integrates computer vision, neural artistic abstraction, motion optimization, and robot control. User ratings on a 5-point scale were high for sketch quality 4.33, perceived execution 4.53, and user experience 4.65, indicating recognizable, appealing, and engaging robotic portraits.

Keywords

Cite

@article{arxiv.2608.00829,
  title  = {GeminiPainter's sequence-formed pipeline comprised of perception, cognition, planning, and action stages},
  author = {Miguel Altamirano Cabrera and Aleksey Fedoseev and Iana Zhura and Dzmitry Tsetserukou},
  journal= {arXiv preprint arXiv:2608.00829},
  year   = {2026}
}