Spatio-Temporal Reinforcement Learning-Driven Ship Path Planning Method in Dynamic Time-Varying Environments: Research on Adaptive Decision-Making in Typhoon Scenarios

Weizheng Wang, Fenghua Liu, Kai Cheng, Zuopeng Niu, Zhengwei He · Electronics · 2025

In dynamic environments with continuous variability, such as those affected by typhoons, ship path planning must account for both navigational safety and the maneuvering characteristics of the vessel. However, current methods often struggle to accurately capture the continuous evolution of dynamic obstacles and generally lack adaptive exploration mechanisms. Consequently, the planned routes tend to be suboptimal or incompatible with the ship’s maneuvering constraints. To address this challenge, this study proposes a Space–Time Integrated Q-Learning (STIQ-Learning) algorithm for dynamic path planning under typhoon conditions. The algorithm is built upon the following key innovations: (1) Spatio-Temporal Environment Modeling: The hazardous area affected by the typhoon is decomposed into temporally and spatially dynamic obstacles. A grid-based spatio-temporal environment model is constructed by integrating forecast data on typhoon wind radii and wave heights. This enables a precise representation of the typhoon’s dynamic evolution process and the surrounding maritime risk environment. (2) Optimization of State Space and Reward Mechanism: A time dimension is incorporated to expand the state space, while a composite reward function is designed by combining three sub-reward terms: target proximity, trajectory smoothness, and heading correction. These components jointly guide the learning agent to generate navigation paths that are both safe and consistent with the maneuverability characteristics of the vessel. (3) Priority-Based Adaptive Exploration Strategy: A prioritized action selection mechanism is introduced based on collision feedback, and the exploration factor ϵ is dynamically adjusted throughout the learning process. This strategy enhances the efficiency of early exploration and effectively balances the trade-off between exploration and exploitation. Simulation experiments were conducted using real-world scenarios derived from Typhoons Pulasan and Gamei in 2024. The results demonstrate that in open-sea environments, the proposed STIQ-Learning algorithm achieves reductions in path length of 14.4% and 22.3% compared to the D* and Rapidly exploring Random Trees (RRT) algorithms, respectively. In more complex maritime environments featuring geographic constraints such as islands, STIQ-Learning reductions of 2.1%, 20.7%, and 10.6% relative to the DFQL, D*, and RRT algorithms, respectively. Furthermore, the proposed method consistently avoids the hazardous wind zones associated with typhoons throughout the entire planning process, while maintaining wave heights along the generated routes within the vessel’s safety limits.

Read the paper · More papers on PaperTik