Dynamic programming with ARMA, Markov, and NARMA models vs. Q-learning-case study

J. Chrobak, Andrzej Pacut, Andrzej Karbowski · 2000

Two approaches to control policy synthesis for unknown systems are investigated. The indirect approach is based on the identification of ARMA, NARMA, or Markov chain models, and applications of dynamic programming to these models with or without the use of a certainty equivalence principle. The direct approach is represented by Q-learning, with the lookup table or with the use of radial basis function approximation. We implemented both methods to optimization of a stock portfolio and tested on the Warsaw stock exchange data.

Read the paper · More papers on PaperTik