Exploitation-Oriented Learning with Deep Learning – Introducing Profit Sharing to a Deep Q-Network –

Kazuteru Miyazaki, National Institution for Academic Degrees and Quality Enhancement of Higher Education 1-29-1 Gakuennishimachi, Kodaira, Tokyo 185-8587, Japan · Journal of Advanced Computational Intelligence and Intelligent Informatics · 2017

Currently, deep learning is attracting significant interest. Combining deep Q-networks (DQNs) and Q-learning has produced excellent results for several Atari 2600 games. In this paper, we propose an exploitation-oriented learning (XoL) method that incorporates deep learning to reduce the number of trial-and-error searches. We focus on a profit sharing (PS) method that is an XoL method, and combine it with a DQN to propose a DQNwithPS method. This method is compared with a DQN in Atari 2600 games. We demonstrate that the proposed DQNwithPS method can learn stably with fewer trial-and-error searches than required by only a DQN.

Read the paper · More papers on PaperTik