Dynamic A* Heuristic Function Optimization Based on Deep Reinforcement Learning and Path Planning in Uncertain Environments
Wenhan Li · 2025
This paper addresses the complex problem of online truck routing in container terminals, a critical operational challenge faced by major international ports. The proposed approach models the terminal environment as a directed graph comprising warehouses, quay cranes, and yard blocks, and formulates the task scheduling process with a focus on minimizing quay crane waiting time. The problem is characterized by dynamic task arrivals and stochastic service durations, which are only revealed in real time, making it inherently online and uncertain. To tackle this, we introduce a sequential decision-making framework using deep reinforcement learning (DRL), which enables adaptive and robust policy learning across diverse operational scenarios. Decision variables capture both truck assignment and temporal scheduling, while constraints enforce task precedence, service order, and exclusivity. Experimental design highlights the framework’s scalability and practical applicability, offering significant improvements in efficiency and responsiveness in uncertain port logistics environments.