跨域深化强化学习用于从农田到月球的导航迁移
摘要
在非结构化环境中实现自主导航对于田间和行星机器人至关重要,机器人必须在不可预测的条件下高效地到达目标并避免障碍物。传统的算法方法通常需要大量的环境特定调校,限制了其在新领域中的可扩展性。深度强化学习 (DRL) 提供了一种数据驱动的替代方法,使机器人能够通过直接与其环境交互来获取导航策略。本 work investigates the feasibility of DRL policy generalization across visually and topographically distinct simulated domains, where policies are trained in terrestrial settings and validated in a zero-shot manner in extraterrestrial environments. A 3D simulation of an agricultural rover is developed and trained using Proximal Policy Optimization (PPO) to achieve goal-directed navigation and obstacle avoidance in farmland settings. The learned policy is then evaluated in a lunar-like simulated environment to assess transfer performance. The results indicate that policies trained under terrestrial conditions retain a high level of effectiveness, achieving close to 50% success in lunar simulations without the need for additional training and fine-tuning. This underscores the potential of cross-domain DRL-based policy transfer as a promising approach to developing adaptable and efficient autonomous navigation for future planetary exploration missions, with the added benefit of minimizing retraining costs.
引用
@article{arxiv.2510.23329,
title = {Transferable Deep Reinforcement Learning for Cross-Domain Navigation: from Farmland to the Moon},
author = {Shreya Santra and Thomas Robbins and Kazuya Yoshida},
journal= {arXiv preprint arXiv:2510.23329},
year = {2025}
}
备注
6 pages, 7 figures. Accepted at IEEE iSpaRo 2025