Learning Heuristics for Course of Action Analysis with Reinforcement Learning

Jonathan Cawalla · 2025

Supporting an operator in analyzing Courses of Action (CoA) is challenging due to the problem's inherent complexity, which involves balancing multiple objectives, managing limited resources, dealing with unpredictable adversaries, and addressing a wide variety of scenarios. These elements contribute to a complex decision-making environment that demands thorough analysis and careful planning. By employing a graph-based representation of the terrain and action space, we simplify the problem, enabling the development of automated solutions. However, creating solutions remains difficult, as optimization methods and heuristics require significant time and expertise to develop and must be adjusted whenever mission requirements change. Building on previous work, we propose a novel hybrid model that integrates beam search with a learned policy from Reinforcement Learning. This approach offers the key advantage of being easily adaptable to new scenarios and changing requirements. Compared to our previous work, we significantly expand the original problem formulation by introducing a time-based, sequential CoA problem and demonstrate superior performance and scalability for larger scenarios. This paper was originally presented at the NATO Science and Technology Organization Symposium (ICMCIS) organized by the Information Systems Technology (IST) Panel, IST205-RSY - the ICMCIS, held in Oeiras, Portugal, 13–14 May 2025.

Read the paper · More papers on PaperTik