Meta-Policy Gradient Reinforcement Learning for Energy Efficiency in Hierarchical 6G Task Offloading

C. Rajinikanth, Gunasekaran Thangavel, Dharmalingam Murugesan, A. Rajaram · Journal of Circuits Systems and Computers · 2025

The era of 6G will see a shift in how connected systems design, allocate and repurpose computationally-driven workloads to take advantage of distributed systems. The increasingly data-driven nature and the ever-growing demands for ultra-low-latency functionality make the cloud-centric approach insufficient. In this paper, we propose a novel-Edge–Fog–Cloud–Device (EFCD) architecture, which uses the federated digital twin intelligence and autonomous orchestration to optimize the task offloading in extremely dynamic and STRE source-limited 6G networks. The architecture describes a multi-layered cognitive orchestration model based on a Meta-Policy Gradient Reinforcement Learning (MPG-RL) algorithm that dynamically adjusts to changes in user mobility, network latency, energy limitations and computational urgency. We introduce a federated learning-based digital twin system to ensure data privacy and minimize communication costs across layers, where learning is synchronized without a centralized sharing of raw data. The approach is demonstrated and validated on our own generated synthetic smart city dataset of a dense urban scenario. By extensive simulation, we reduce task latency by 47%, improve energy efficiency by 33% and increase throughput by 38%, over traditional deep reinforcement and centralized orchestration approaches. This work is a significant step toward self-adaptive, sustainable and privacy-preserving computation for the forthcoming 6G environments. This work demonstrates that the proposed MPG-RL achieves a success rate of 91.3% and adapts rapidly within 7.1[Formula: see text]s. The core philosophy behind this mechanism is to enable efficient and scalable task offloading by leveraging digital twins and federated meta-learning, thereby allowing the system to quickly generalize and adapt to diverse network conditions.

Read the paper · More papers on PaperTik