RTP-Q: A reinforcement learning system with time constraints exploration planning for accelerating the learning rate
Citation
G. Zhao et al., RTP-Q: A reinforcement learning system with time constraints exploration planning for accelerating the learning rate, IEICE T FUN, E82A(10), 1999, pp. 2266-2273
Categorie Soggetti
Eletrical & Eletronics Engineeing
Journal title
IEICE TRANSACTIONS ON FUNDAMENTALS OF ELECTRONICS COMMUNICATIONS AND COMPUTER SCIENCES
SICI code
0916-8508(199910)E82A:10<2266:RARLSW>2.0.ZU;2-Q
Abstract
Reinforcement learning is an efficient method for solving Markov Decision P
rocesses that an agent improves its performance by using scalar reward valu
es with higher capability of reactive and adaptive behaviors. Q-learning is
a representative reinforcement learning method which is guaranteed to obta
in an optimal policy but needs numerous trials to achieve it. Ic-Certainty
Exploration Learning System realizes active exploration to an environment,
but, the learning process is separated into two phases and estimate values
are not derived during the process of identifying the environment. Dyna-Q a
rchitecture makes fuller use of a limited amount of experiences and achieve
s a better policy with fewer environment interactions during identifying an
environment by learning and planning with constrained time, however, the e
xploration is not active. This paper proposes a RTP-Q reinforcement learnin
g system which varies an efficient method for exploring an environment into
time constraints exploration planning and compounds it into an integrated
system of learning, planning and reacting for aiming for the best of both m
ethods. Based on improving the performance of exploring an environment, ref
ining the model of the environment, the RTP-Q learning system accelerates t
he learning rate for obtaining an optimal policy. The results of experiment
; on navigation tasks demonstrate that the RTP-Q learning system is efficie
nt.