English

Learning Social Navigation from Positive and Negative Demonstrations and Rule-Based Specifications

Robotics 2025-10-15 v1

Abstract

Mobile robot navigation in dynamic human environments requires policies that balance adaptability to diverse behaviors with compliance to safety constraints. We hypothesize that integrating data-driven rewards with rule-based objectives enables navigation policies to achieve a more effective balance of adaptability and safety. To this end, we develop a framework that learns a density-based reward from positive and negative demonstrations and augments it with rule-based objectives for obstacle avoidance and goal reaching. A sampling-based lookahead controller produces supervisory actions that are both safe and adaptive, which are subsequently distilled into a compact student policy suitable for real-time operation with uncertainty estimates. Experiments in synthetic and elevator co-boarding simulations show consistent gains in success rate and time efficiency over baselines, and real-world demonstrations with human participants confirm the practicality of deployment. A video illustrating this work can be found on our project page https://chanwookim971024.github.io/PioneeR/.

Keywords

Cite

@article{arxiv.2510.12215,
  title  = {Learning Social Navigation from Positive and Negative Demonstrations and Rule-Based Specifications},
  author = {Chanwoo Kim and Jihwan Yoon and Hyeonseong Kim and Taemoon Jeong and Changwoo Yoo and Seungbeen Lee and Soohwan Byeon and Hoon Chung and Matthew Pan and Jean Oh and Kyungjae Lee and Sungjoon Choi},
  journal= {arXiv preprint arXiv:2510.12215},
  year   = {2025}
}

Comments

For more videos, see https://chanwookim971024.github.io/PioneeR/

R2 v1 2026-07-01T06:35:45.017Z