An Adaptive and Scalable Framework for Resource-Efficient Deployment of Mixture of Experts in LLM-Based Intelligent IoT Networks

C.M. Liu, Yangfan Li, Cen Chen, Hailan Kuang, Xiaolin Ma, Xiaofeng Zou, Jing Liu, Zeyu Lu, Zhaoyuan Zhang, Xinhua Liu · IEEE Internet of Things Journal · 2025

The exponential growth of the Internet of Things (IoT) necessitates the deployment of large-scale models capable of processing the complex and diverse data generated by IoT devices. However, the substantial memory requirements of these models pose significant challenges, especially in scenarios where rapid decision-making and low-latency responses are critical. To address these challenges, we propose three innovative strategies for optimizing large model usage in IoT environments. The first strategy is an adaptive loading scheme, which enables dynamic loading of individual model experts. The second strategy involves an expert-by-expert loading approach, further enhancing the ability to load experts as needed, which optimizes memory usage and accelerates computations. The third strategy employs an interlayer expert reuse mechanism, facilitating the efficient reuse of experts across different layers, thus enhancing response rates without compromising model accuracy. Importantly, these strategies can be directly applied to Mixture of Experts (MoE) large language models without requiring additional training, thereby providing a seamless and efficient solution for leveraging these models in memory-constrained, high-performance IoT environments.

Read the paper · More papers on PaperTik