Reinforcement Learning for Selecting Custom Instructions Under Area Constraint

Shanshan Wang, Chenglong Xiao · IEEE Transactions on Artificial Intelligence · 2023

Extensible processors, which combine programma-bility and efficiency, are emerging as a promising approach in the field of embedded computing. Automated synthesis of custom instructions from high-level application descriptions is a vital step involved in the design of extensible processors. In automated custom instruction synthesis, selecting custom instructions from a large set of candidates under area constraint is essentially a difficult combinatorial optimization problem. In this paper, we show that the custom instruction selection problem can be formulated as a sequential decision-making problem. Based on this formulation, we present three reinforcement learning-based approaches, namely SARSA, Q-learning, and Double Q-Learning, for solving the custom instruction selection problem. Moreover, we also perform a comprehensive analysis and com-parison of various combinations of learning specifications: the algorithm type and the update strategy for ε-greedy policy. The experiments with 45 test instances reveal that the SARSA, Q-learning, and Double Q-Learning algorithms outperform the meta-heuristic algorithm in terms of the overall performance gains by 26.9%,26.1% and 26.4% respectively. Among the three reinforcement learning algorithms, the SARSA algorithm slightly overwhelms the other two reinforcement learning algorithms. Furthermore, the experimental results suggest that the strategy$F_3 : ε = κ^i,0 < κ < 1$is generally the most effective one for controlling the exploration and the exploitation of the learning processes.

Read the paper · More papers on PaperTik