Curriculum Learning Influences the Emergence of Different Learning Trends

Romina Mir, Pegah Ojaghi, Andrew Erwin, Ali Marjaninejad, Michael F. Wehner, Francisco J. Valero‐Cuevas · 2024

Reinforcement learning (RL) algorithms are traditionally evaluated and compared by their learning trends (i.e., average performance) over trials and time. However, the presence of a single learning trend in a curriculum is, in fact, an assumption. To test this assumption, we used the performance of Proximal Policy Optimization (PPO) under five different curricula aimed at learning dynamic in-hand manipulation tasks. The curricula consisted of different combinations of rewards for lifting and rotating a 5g ball with a three-finger hand with the palm facing down. Mining the performance of all 60 individual trials as time series, we find there are learning trends distinct from the average. We conclude researchers should look beyond the average learning trends when evaluating curriculum learning to fully identify, appreciate, and evaluate the progression of autonomous learning of multi-objective tasks.

Read the paper · More papers on PaperTik