Convergence of the Q-ae learning on deterministic MDPs and its efficiency on the stochastic environment
Citation
G. Zhao et al., Convergence of the Q-ae learning on deterministic MDPs and its efficiency on the stochastic environment, IEICE T FUN, E83A(9), 2000, pp. 1786-1795
Categorie Soggetti
Eletrical & Eletronics Engineeing
Journal title
IEICE TRANSACTIONS ON FUNDAMENTALS OF ELECTRONICS COMMUNICATIONS AND COMPUTER SCIENCES
SICI code
0916-8508(200009)E83A:9<1786:COTQLO>2.0.ZU;2-5
Abstract
Reinforcement Learning (RL) is an efficient method for solving Markov Decis
ion Processes (MDPs) without a priori knowledge about an environment, and c
an be classified into the exploitation oriented method and the exploration
oriented method. Q-learning is a representative RL and is classified as an
exploration oriented method. It is guaranteed to obtain an optimal policy,
however, Q-learning needs numerous trials to learn it because there is not
action-selecting mechanism in Q-learning. For accelerating the learning rat
e of the Q-learning and realizing exploitation and exploration at a learnin
g process, the Q-ee,learning system has been proposed, which uses preaction
-selector, action-selector and back propagation of Q values to improve the
performance of Q-learning. But the Q-ee learning is merely suitable for det
erministic MDPs, and its convergent guarantee to derive an optimal policy h
as not been proved. In this paper, based on discussing different exploratio
n methods, replacing the pre-action-selector in the Q-ce learning, we intro
duce a method that can be used to implement an active exploration to an env
ironment, the Active Exploration Planning (AEP), into the learning system,
which we call the Q-ae learning. With this replacement, the Q-ae learning n
ot only maintains advantages of the Q-ee learning but also is adapted to a
stochastic environment. Moreover, under deterministic MDPs, this paper pres
ents the convergent condition and its proof for an agent to obtain the opti
mal policy by the method of the Q-ae learning. Further, by discussions and
experiments, it is shown that by adjusting the relation between the learnin
g factor and the discounted rate, the exploration process to an environment
can be controlled on a stochastic environment. And, experimental results a
bout the exploration rate to an environment and the correct rate of learned
policies also illustrate the efficiency of the Q-ae learning on the stocha
stic;environment.