基于深度确定性梯度学习的集群多目标分配方法

Research on Multi-Target Assignment Method for Clusters Based on Deep Deterministic Policy Gradient Learning

  • 摘要: 针对多弹协同作战进行目标分配时,存在敌方平台和反舰导弹数量不确定性和类型多样化,导致目标分配算法难以建模的问题,为提升高动态协同攻击条件下的攻击效能,建立动态战场环境模型和多目标分配的单回合马尔可夫决策模型,提出一种改进深度确定性策略梯度的分配算法. 通过与模拟器的交互自动求解最佳分配策略,利用mask方法对动作空间进行掩码操作,实现算法对平台数量和类型的适应能力. 实验结果表明,在各种不同舰船的防御配置和红蓝双方数量配置下,算法求解得到的攻击策略相对于随机策略的性能提升约为87.5%,模型推理时间约为0.04 ms. 研究结果将加速基于深度确定性梯度学习的方法在高动态环境下智能决策中的应用,对集群自主决策方法的研究具有推动作用.

     

    Abstract: In the target assignment for multi-missile cooperative operations, there exists uncertainty in the number and variety of enemy platforms and anti-ship missiles, which makes it difficult to model the target assignment algorithm. To improve the effectiveness of attacks under high-dynamic collaborative attack conditions, a dynamic battlefield environment model and a single-round Markov decision model for multi-target assignment were established. An improved deep deterministic policy gradient (DDPG) assignment algorithm was proposed to automatically find the optimal allocation strategy through interaction with the simulator. The algorithm uses the mask method to mask the action space and adapt to the number and type of platforms. The simulation results show that under different defense configurations and configurations of red and blue sides, the performance improvement of the attack strategy obtained by the algorithm was about 87.5% compared with that of the random strategy, and the reasoning time of the model was about 0.04 ms. This research will accelerate the application of DDPG-based methods in intelligent decision-making in high-dynamic environments, and promote the research on cluster autonomous decision-making methods.

     

/

返回文章
返回
Baidu
map