Learning collusive strategies with function approximation algorithms
Miguel Á. González Ballester, Penélope Hernández · Journal of Dynamics and Games · 2025
This paper presents an algorithm called Q-function based on the gradient-decent function approximation that learns a dynamic environment that depends on the market dynamics. Q-function is able to learn a pricing strategy in a non-stationary environment estimating a function that is linear in parameters. Throughout the game, both firms face a trade-off between setting the prices that they have learned to be optimal at each state and exploring new pricing strategies that can allow them to learn better estimations of our function approximation. After running 1000 simulations for the baseline parameters in a duopolistic market where each firm is modeled by a Q-function, it is observed that the Q-function firms converged to supra-competitive prices beyond the Nash equilibrium price for the one shot game. Hence, Q-function with a somewhat restrictive structure can converge to collusive outcomes. However, the collusive behaviour of these algorithms can be altered by applying external shocks that force the algorithms to follow a competitive path.