English

DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces

Computer Vision and Pattern Recognition 2026-02-27 v2 Artificial Intelligence

Abstract

Articulated object pose estimation is a core task in embodied AI. Existing methods typically regress poses in a continuous space, but often struggle with 1) navigating a large, complex search space and 2) failing to incorporate intrinsic kinematic constraints. In this work, we introduce DICArt (DIsCrete Diffusion for Articulation Pose Estimation), a novel framework that formulates pose estimation as a conditional discrete diffusion process. Instead of operating in a continuous domain, DICArt progressively denoises a noisy pose representation through a learned reverse diffusion procedure to recover the GT pose. To improve modeling fidelity, we propose a flexible flow decider that dynamically determines whether each token should be denoised or reset, effectively balancing the real and noise distributions during diffusion. Additionally, we incorporate a hierarchical kinematic coupling strategy, estimating the pose of each rigid part hierarchically to respect the object's kinematic structure. We validate DICArt on both synthetic and real-world datasets. Experimental results demonstrate its superior performance and robustness. By integrating discrete generative modeling with structural priors, DICArt offers a new paradigm for reliable category-level 6D pose estimation in complex environments.

Keywords

Cite

@article{arxiv.2602.19565,
  title  = {DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces},
  author = {Li Zhang and Mingyu Mei and Ailing Wang and Xianhui Meng and Yan Zhong and Xinyuan Song and Liu Liu and Rujing Wang and Zaixing He and Cewu Lu},
  journal= {arXiv preprint arXiv:2602.19565},
  year   = {2026}
}
R2 v1 2026-07-01T10:46:58.143Z