Deep Q-Network Agents for Game Playing: Systematic Evaluation Across Eight Benchmark and Custom Environments

Časlav Livada, Marko Duka, Tomislav Keser, Krešimir Nenadić · Electronics · 2026

Deep Q-Networks (DQNs) have achieved strong performance across a range of benchmark tasks; however, their reliability under varying reward structures and planning horizons remains insufficiently characterized. This study presents a systematic cross-environment analysis of DQN agents evaluated across eight environments spanning simple control, arcade, and strategic domains. Rather than pursuing state-of-the-art performance, the objective is to investigate structural conditions under which standard value-based reinforcement learning succeeds, degrades, or fails. Across controlled experiments with consistent training budgets and statistical validation, three recurring failure patterns are identified: (i) sparse-reward exploration failure, (ii) reward exploitation without functional task competence, and (iii) strategic planning limitations in long-horizon or adversarial environments. Within-environment ablation studies further demonstrate that moderate network scaling (2–4× parameter increases) does not significantly alter learning outcomes when reward functions remain unchanged, suggesting that reward alignment and task horizon dominate architectural capacity as determinants of performance. The results provide a structured diagnostic perspective on DQN reliability, clarify the limits of reward shaping in complex environments, and offer practical guidance for identifying when standard value-based methods are likely to become unstable or insufficient.

Read the paper · More papers on PaperTik