Convergence of the Q-ae learning on deterministic MDPs and its efficiency on the stochastic environment

Citation
G. Zhao et al., Convergence of the Q-ae learning on deterministic MDPs and its efficiency on the stochastic environment, IEICE T FUN, E83A(9), 2000, pp. 1786-1795
Citations number
21
Categorie Soggetti
Eletrical & Eletronics Engineeing
Journal title
IEICE TRANSACTIONS ON FUNDAMENTALS OF ELECTRONICS COMMUNICATIONS AND COMPUTER SCIENCES
ISSN journal
09168508 → ACNP
Volume
E83A
Issue
9
Year of publication
2000
Pages
1786 - 1795
Database
ISI
SICI code
0916-8508(200009)E83A:9<1786:COTQLO>2.0.ZU;2-5
Abstract
Reinforcement Learning (RL) is an efficient method for solving Markov Decis ion Processes (MDPs) without a priori knowledge about an environment, and c an be classified into the exploitation oriented method and the exploration oriented method. Q-learning is a representative RL and is classified as an exploration oriented method. It is guaranteed to obtain an optimal policy, however, Q-learning needs numerous trials to learn it because there is not action-selecting mechanism in Q-learning. For accelerating the learning rat e of the Q-learning and realizing exploitation and exploration at a learnin g process, the Q-ee,learning system has been proposed, which uses preaction -selector, action-selector and back propagation of Q values to improve the performance of Q-learning. But the Q-ee learning is merely suitable for det erministic MDPs, and its convergent guarantee to derive an optimal policy h as not been proved. In this paper, based on discussing different exploratio n methods, replacing the pre-action-selector in the Q-ce learning, we intro duce a method that can be used to implement an active exploration to an env ironment, the Active Exploration Planning (AEP), into the learning system, which we call the Q-ae learning. With this replacement, the Q-ae learning n ot only maintains advantages of the Q-ee learning but also is adapted to a stochastic environment. Moreover, under deterministic MDPs, this paper pres ents the convergent condition and its proof for an agent to obtain the opti mal policy by the method of the Q-ae learning. Further, by discussions and experiments, it is shown that by adjusting the relation between the learnin g factor and the discounted rate, the exploration process to an environment can be controlled on a stochastic environment. And, experimental results a bout the exploration rate to an environment and the correct rate of learned policies also illustrate the efficiency of the Q-ae learning on the stocha stic;environment.