中文

面向直驱双翼实验平台的全插拔式即时强化学习算法

机器学习 2024-12-23 v2 人工智能 机器人学

摘要

双翼系统产生的非线性和不稳定的气动干扰,在多种随机运行条件下为运动控制带来重大挑战。为此,已开发了 Concerto 强化学习扩展算法(CRL2E)。该算法具备插拔式、全即时、基于现场的强化学习特性,融合了新颖的基于物理的规则策略编 composer 策略,搭配扰动模块,以及针对实时控制优化的轻量级网络。为验证模块设计的性能和合理性,在六个具有挑战性的运行条件下进行了实验,比较了七种不同算法。结果表明,CRL2E 在前 500 步内即可实现安全稳定的训练,跟踪精度比 Soft Actor-Critic、Proximal Policy Optimization 和 Twin Delayed Deep Deterministic Policy Gradient 算法提升了 14 倍至 66 倍。此外,CRL2E 在各种随机运行条件下均显著提升性能,跟踪精度提升幅度从 8.3% 到 60.4%。CRL2E 的收敛速度比仅引入 Composer 扰动的 CRL 算法快 36.11% 至 57.64%,比同时引入 Composer 扰动和时间交错扰动的 CRL 算法快 43.52% 至 65.85%,尤其在标准 CRL 难以收敛的条件下效果更为显著。硬件测试表明,优化后的轻量级网络结构在权重加载和平均推理时间方面表现出色,满足实时控制要求。

关键词

引用

@article{arxiv.2410.15554,
  title  = {A Plug-and-Play Fully On-the-Job Real-Time Reinforcement Learning Algorithm for a Direct-Drive Tandem-Wing Experiment Platforms Under Multiple Random Operating Conditions},
  author = {Zhang Minghao and Song Bifeng and Yang Xiaojun and Wang Liang},
  journal= {arXiv preprint arXiv:2410.15554},
  year   = {2024}
}

备注

To prevent potential misunderstandings or negative impacts on the community, I am requesting the withdrawal of my submission due to the discovery of critical errors and major flaws in the work. Recent discussions with researchers in the field have identified significant defects that compromise the validity of the results