ARMC-RL: Adaptive Caching With Reinforcement Learning for Efficient 360° Video Streaming in Edge Networks
Minji Choi, Somin Park, Jin-Hyun Ahn, Dong Ho Kim, Cheolwoo You · IEEE Access · 2025
The demand for high-capacity content such as 3D and 360° videos continues to grow. To address this, multi-access edge computing (MEC) is gaining attention as a next-generation computing technology in the B5G/6G network. However, since MEC systems have relatively limited memory capacity compared to the amount of data they must process, cache replacement strategies play a critical role. We propose an adaptive recency masking caching (ARMC) algorithm to optimize MEC caching for 360° video streaming. The proposed caching mechanism efficiently replaces cached data by combining two techniques: recency masking and Frequency-Filtered Least Recently Used. The use of the techniques is determined by the importance of the cache data and the available cache capacity. In addition, we introduce the concept of an observation window to improve cache performance by reflecting the recency of data request patterns. We conducted experiments using a field of view (FoV) dataset recorded from real users watching 360° videos via head-mounted displays. Given the characteristics of 360° videos, we assumed that the MEC cache would require high-quality tiles matching the user’s FoV at each moment. Through experiments, we confirmed that the proposed method achieved a higher cache hit rate compared to existing cache replacement techniques. In particular, ARMC improved the hit rate by up to 29% compared to Least Frequency Used algorithm under constrained cache conditions of 6% or less of the total data. Higher cache hit rates contribute to reducing transmission latency and lowering bandwidth consumption. Meanwhile, the size of the observation window, a key variable in the proposed technique, varied in terms of its optimal size depending on the cache size and user viewing patterns. To address this issue, we proposed ARMC-RL, a variant of ARMC with reinforcement learning (RL) assistance, which is designed to dynamically estimate the optimal observation window size according to the given environment. Based on the experiments, ARMC-RL, depending on the learning model, achieves cache hit rates that are similar to or higher than those of ARMC and converges to the optimal observation window size by the second episode. This result confirms that ARMC-RL enables the stable automation of ARMC. Ultimately, ARMC-RL can be effectively applied to MEC-based caching systems, enabling adaptive cache optimization in response to dynamic user demand and network conditions.