RTP-Q: A reinforcement learning system with time constraints exploration planning for accelerating the learning rate

Citation
G. Zhao et al., RTP-Q: A reinforcement learning system with time constraints exploration planning for accelerating the learning rate, IEICE T FUN, E82A(10), 1999, pp. 2266-2273
Citations number
12
Categorie Soggetti
Eletrical & Eletronics Engineeing
Journal title
IEICE TRANSACTIONS ON FUNDAMENTALS OF ELECTRONICS COMMUNICATIONS AND COMPUTER SCIENCES
ISSN journal
09168508 → ACNP
Volume
E82A
Issue
10
Year of publication
1999
Pages
2266 - 2273
Database
ISI
SICI code
0916-8508(199910)E82A:10<2266:RARLSW>2.0.ZU;2-Q
Abstract
Reinforcement learning is an efficient method for solving Markov Decision P rocesses that an agent improves its performance by using scalar reward valu es with higher capability of reactive and adaptive behaviors. Q-learning is a representative reinforcement learning method which is guaranteed to obta in an optimal policy but needs numerous trials to achieve it. Ic-Certainty Exploration Learning System realizes active exploration to an environment, but, the learning process is separated into two phases and estimate values are not derived during the process of identifying the environment. Dyna-Q a rchitecture makes fuller use of a limited amount of experiences and achieve s a better policy with fewer environment interactions during identifying an environment by learning and planning with constrained time, however, the e xploration is not active. This paper proposes a RTP-Q reinforcement learnin g system which varies an efficient method for exploring an environment into time constraints exploration planning and compounds it into an integrated system of learning, planning and reacting for aiming for the best of both m ethods. Based on improving the performance of exploring an environment, ref ining the model of the environment, the RTP-Q learning system accelerates t he learning rate for obtaining an optimal policy. The results of experiment ; on navigation tasks demonstrate that the RTP-Q learning system is efficie nt.