Intrinsically Motivated Hierarchical Reinforcement Learning for Lifelong Learning by Computer Game Characters
Mohtasham Khani, Mahtab · UNSWorks (University of New South Wales, Sydney, Australia) · 2025
Reinforcement learning (RL) agents often face unknown challenges in dynamic environments they have not previously encountered. To overcome these newly emerging challenges, agents must have the ability to define their own goals, learn new skills, and solve problems as they arise, without explicit supervision. Intrinsically motivated RL (IMRL) allows agents to explore and learn independently, driven by curiosity rather than predefined rewards. This thesis proposes a novel approach to help intrinsically motivated reinforcement learning (IMRL) agents develop practical skill sets according to the specific environment and generate a hierarchy of behaviours throughout their lifetime. Using graph-based representations and planning algorithms, the proposed method enables agents to recognise patterns, adapt and plan future actions. The proposed methods are applied to three benchmark environments with a discrete state and action space. The thesis also introduces a new intrinsic motivation(IM) method, the `Frontier method', which helps agents explore environments more effectively. By using state-transition graphs, this method is compared with other motivation strategies, such as self-organising maps (SOM), demonstrating how different motivations shape agent's behaviour. To promote diverse behaviour, a novel algorithm is proposed that employs a hierarchical clustering approach to identify key critical points or `ready states' from which agents can plan their next moves, similar to fundamental positions in sports and games. These ready states are derived using our proposed hierarchical clustering algorithm on the agent’s state transition graph. They help agents to structure their movements and behaviour. The thesis also evaluates how agents learn over time by analysing their lifelong learning ability and behavioural competence. By breaking their lifespan into different stages, the approach tracks the progression of learning and adaptation across the phases of an agent's experience. Empirical evaluations demonstrate that agents using the IM method exhibit meaningful motivation-based behaviours, as well as their ability to plan a new skill using the knowledge acquired during their lifelong. Finally, the thesis integrates planning algorithms to help agents not only explore but also plan ahead. By applying breadth first search(BFS) to ready states, agents can create a new and behaviour, allowing them to anticipate future actions rather than just reacting to their environment. The proposed methods were evaluated across three discrete environments, including the Four Rooms benchmark, a custom River environment, and a game-like environment inspired by Minecraft, demonstrating the generalisability of behaviours and skills across varying levels of complexity and exploration challenges. This research provides a systematic framework for developing behavioural competence in RL agents, combining IM, graph based techniques, clustering and planning. The outcome of this work contributes to the development of a set of behaviours that can help the self-improving of RL agents to navigate effectively in complex environments.