A Comparative Study of DRL Algorithms for Map-Free Robot Navigation With Zero-Shot Sim-to-Real Transfer
Taner Yılmaz, Ömür Aydoğmuş · IEEE Access · 2026
This paper presents an empirical benchmark of map-free deep reinforcement learning (DRL) for goal-driven indoor navigation using LiDAR-only perception and continuous control, together with a reproducible pipeline for zero-shot sim-to-real transfer. A custom Python simulator is introduced in which, at the start of each episode, the start pose, goal position, and obstacle layout are randomly sampled to generate diverse navigation scenarios; obstacles remain static during each rollout. Five widely used DRL baselines are examined: Proximal Policy Optimization (PPO), Soft Actor-Critic (SAC), Advantage Actor-Critic (A2C), Twin Delayed Deep Deterministic Policy Gradient (TD3), and Deep Deterministic Policy Gradient (DDPG). After an initial 10,000-episode feasibility screening, all five methods were trained from scratch for 100,000 episodes under a unified observation and reward design, and results are averaged over five independent training seeds. Hyperparameters are tuned using Optuna, and practical refinements (state normalisation, LiDAR densification, reward shaping, and cosine-annealed learning rates) are assessed via ablation studies. PPO achieves the best overall trade-off between success, collision risk, and path efficiency, reaching up to 97.6% success in simulation (seed-averaged). For sim-to-real validation, the PPO policy is deployed without fine-tuning across three aligned evaluation testbeds: the Python simulator, a ROS/Gazebo digital twin, and a TurtleBot3 Burger arena. Beyond the aligned three-domain sim-to-real testbeds, an extended real-robot evaluation of 100 repeated episodes over two fixed layouts and two predefined start/goal configurations is reported under controlled static indoor conditions. Distributional statistics (median and IQR) of path length, time-to-goal, minimum clearance, and near-miss exposure are reported to quantify safety and efficiency under repeated controlled static indoor hardware trials.