基于深度仲裁策略的四足机器人步态学习

Gait Learning of Quadruped Robot Based on Deep Arbitration Strategy

  • 摘要: 复现高等生物的学习过程是机器人研究的一个重要研究方向,研究人员已探索出一些常用的基于行动者评价器(actor critic,AC)网络的强化学习算法可以完成此任务,但是还存在一些不足. 针对深度确定性策略梯度(deep deterministic policy gradient,DDPG)存在着<i<Q</i<值过估计导致恶化学习效果的问题,受到大脑前额叶皮质层仲裁机制的启发,提出了一种深度仲裁行动者评价器(deep arbitration actor critic,DAAC)算法,其中包含两套评价网络,通过仲裁机制进行择优选取评价网络去更新策略参数,有效解决了<i<Q</i<值过估计的问题,该算法使得四足机器人成功复现了仿生的步态学习过程. 通过仿真实验,将DAAC算法与DDPG、软行动者评价器(soft actor critic,SAC)、近端策略优化(proximal policy optimization,PPO)三种算法进行了对比实验,实验证明经DAAC训练的四足机器人步态在奖励值、机体稳定性和速度三个方面都有更好的表现,有效验证了算法的优越性.

     

    Abstract: Reproducing the learning process of higher organisms is an important research direction in robot research. Some commonly used reinforcement learning algorithms had been explored based on actor critic (AC) networks to accomplish this task. Due to some shortcomings still existed in the reinforcement learning algorithms, some improvements were also took place. For the deep deterministic policy gradient (DDPG), an overestimated problem to <i<Q</i< value led to deterioration of the learning effect. Inspired by the arbitration mechanism in the prefrontal cortex of the brain, a deep arbitration actor critic (DAAC) algorithm was proposed, including two sets of evaluation networks. Through the arbitration mechanism, an optimal evaluation network was selected to update the policy parameters, solving the overestimated problem to <i<Q</i< value effectively. This algorithm enables the quadruped robot reproduce the bionic gait learning process. In simulation experiments, the DAAC algorithm was compared with three algorithms, DDPG, soft actor critic (SAC), and proximal policy optimization (PPO). The experiment results show that the gait of the quadruped robot trained by DAAC has better performance in three aspects, reward value, machine stability, and speed, verifying effectively the superiority of the algorithm.

     

/

返回文章
返回
Baidu
map