English

DA-Fusion: Deformable Attention-Based RGB-D Fusion Transformer for Unseen Object Instance Segmentation

Computer Vision and Pattern Recognition 2026-07-20 v1 Artificial Intelligence

Abstract

In logistics automation, precise segmentation of unseen objects is crucial for efficient robotic manipulation in cluttered environments. Tasks such as bin-picking and shelf-picking require robust perception to handle occlusions, varying object shapes, and complex spatial arrangements. Traditional RGB-based methods tend to over-segment objects due to their reliance on texture, while depth-based methods often under-segment by focusing primarily on geometric features. To address these limitations, we propose DA-Fusion, a deformable attention-based RGB-D fusion Transformer designed for unseen object instance segmentation. DA-Fusion effectively combines the strengths of both RGB and depth data, enhancing segmentation accuracy in cluttered and multi-layered object environments. We also introduce the Object Clutter Bin Dataset (OCBD), a benchmark dataset specifically tailored for evaluating bin-picking scenarios in top-down views. Extensive evaluations demonstrate that DA-Fusion outperforms state-of-the-art methods across diverse environments, making it particularly suited for real-world logistics tasks.

Keywords

Cite

@article{arxiv.2607.17754,
  title  = {DA-Fusion: Deformable Attention-Based RGB-D Fusion Transformer for Unseen Object Instance Segmentation},
  author = {Yesol Park and Hye-Jung Yoon and Juno Kim and Byoung-Tak Zhang},
  journal= {arXiv preprint arXiv:2607.17754},
  year   = {2026}
}

Comments

7 pages, 5 figures. Published in the Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA 2025)