Popularity-Aware Data Placement in Erasure Coding-Based Edge Storage Systems
Ruikun Luo, Jiadong Zhao, Qiang He, Feifei Chen, Song Wu, Hai Jin, Yun Jin Yang · IEEE Transactions on Parallel and Distributed Systems · 2025
Edge computing allows app vendors to store popular data on edge servers, enabling users to retrieve data with low latency. However, edge servers may become unavailable at runtime due to expected exceptions. Data requests are routed to cloud servers, resulting in increased data retrieval latency. To address this issue,erasure coding(EC) has been employed to improve data availability, aiming to ensure full data access for all the users in anedge storage system(ESS). However, in real-world scenarios, data popularity differs and varies. Existing approaches for edge data placement place coded blocks across the entire system without considering data popularity. As a result, they often suffer from high data retrieval latency. In addition, they are designed to process data items individually. Data placed earlier will limit the placement options for subsequent files because edge servers with the most neighbors in the system can be easily exhausted. Some files cannot be placed properly to accommodate user demands. This increases users' data retrieval latency further. This paper tries to study the placement of multiple files in an edge storage system, considering their popularity. We first model theedge data placement(EDP) problem as a mixed-integer programming problem and prove its$\mathcal {NP}$-hardness. Then, we present an optimal algorithm named EDP-O, decoupling the EDP problem into three convex optimization subproblems for solving with an iterative algorithm. In addition, we propose an approximation algorithm named EDP-A that quickly solves the EDP problem in large-scale scenarios with a guaranteed approximation ratio of$\ln N$. The results of experiments conducted on a real-world dataset show that EDP-O and EDP-A reduce the average data retrieval latency against four representative approaches by an average of 18.4% and 15.6% in small-scale scenarios. EDP-A reduces the average data retrieval latency against four representative approaches by an average of 54.7% and reduces the data discard rate by an average of 34.9% in large-scale scenarios.