一种集成规划的SARSA(λ)强化学习算法
An Integrating Planning SARSA(λ) Algorithm of Reinforcement Learning
-
摘要: 提出一种新的集成规划的 SARSA(λ)强化学习算法 .该算法的主要思想是充分利用已有的经验数据 ,在无模型学习的同时估计系统模型 ,每进行一次无模型学习的试验后 ,利用模型在所记忆的状态 /行动对组成的表中进行规划 ,同时利用该表给出了在学习和规划之间的量化折中参考 .实验结果表明 ,本算法比单纯的无模型学习SARSA(λ)算法有效Abstract: A new integrating planning SARSA (λ) algorithm of reinforcement learning is proposed. The algorithm makes extremely efficient use of the experience data. It learns the model while learning to estimate the optimal value without a model. After each episode, it plans in the table of state/action pairs being recorded, and the table can be as a quantificational trade off reference between learning and planning. The results of experiment show that the algorithm has better performance than the SARSA (λ) algorithm.
下载: