English

MISCGrasp: Leveraging Multiple Integrated Scales and Contrastive Learning for Enhanced Volumetric Grasping

Robotics 2025-07-04 v1 Computer Vision and Pattern Recognition

Abstract

Robotic grasping faces challenges in adapting to objects with varying shapes and sizes. In this paper, we introduce MISCGrasp, a volumetric grasping method that integrates multi-scale feature extraction with contrastive feature enhancement for self-adaptive grasping. We propose a query-based interaction between high-level and low-level features through the Insight Transformer, while the Empower Transformer selectively attends to the highest-level features, which synergistically strikes a balance between focusing on fine geometric details and overall geometric structures. Furthermore, MISCGrasp utilizes multi-scale contrastive learning to exploit similarities among positive grasp samples, ensuring consistency across multi-scale features. Extensive experiments in both simulated and real-world environments demonstrate that MISCGrasp outperforms baseline and variant methods in tabletop decluttering tasks. More details are available at https://miscgrasp.github.io/.

Keywords

Cite

@article{arxiv.2507.02672,
  title  = {MISCGrasp: Leveraging Multiple Integrated Scales and Contrastive Learning for Enhanced Volumetric Grasping},
  author = {Qingyu Fan and Yinghao Cai and Chao Li and Chunting Jiao and Xudong Zheng and Tao Lu and Bin Liang and Shuo Wang},
  journal= {arXiv preprint arXiv:2507.02672},
  year   = {2025}
}

Comments

IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025

R2 v1 2026-07-01T03:45:00.828Z