Masking Actions in Reinforcement Learning: Enhanced PPO for Optimal Harvesting
Jaime Álvarez Urueña, Sarai Cobas, Lautaro Rossi Labianca, Alfonso González‐Briones, Javier Curto Hernández · Applied Engineering in Agriculture · 2025
Highlights Training of an AI agent able to drive on herbaceous crop fields efficiently with no redundancy (65.2% of area covered with a clipped number of actions). Creation of a framework to train AI agents to drive autonomously on herbaceous crop fields. Development of a DRL policy to mask forbidden actions, easing the training phase of AI agents. Abstract. Considering the landscape of today’s global agricultural sector, it is essential to align resource optimization and productivity with the development of long-term sustainable practices. Additionally, challenges such as labor shortages and the high costs of developing and maintaining traditional machinery arise. In a context defined by increasing competitiveness and the advancement of new automation technologies, developing deep reinforcement learning (DRL) algorithms emerges as an ideal solution to meet the sector’s demands. A review of the existing literature reveals that previous research on the subject encompasses various applications of DRL models for specific agricultural tasks and regions. This work proposes a trained AI agent capable of driving agricultural tractors autonomously on any kind of 2D field, regardless of its shape or size. A novel framework based on deep reinforcement learning has been developed to train the model. This framework incorporates a fully customizable reward-penalty layer, reinforcement learning policies, field shapes and sizes, tractor configurations, and neural network architectures. A novel DRL policy incorporating action masking to exclude forbidden actions is also proposed, accelerating convergence and enhancing the agent's learning efficiency. A comprehensive statistical test compares distinct agents trained on different policies and approaches. Selecting the best performing agent renders a mean covered area of 65.2% with a clipped number of actions (250). Keywords: Agriculture, Autonomous driving, Coverage path planning, Deep reinforcement learning, Navigation, PPO.