Research on efficient inference of large language model based on routing policy

Penghui Shan, Xi Wei, Yizhong Zhang, Dong Liu · 2025

To solve the problem of balancing high cost and high performance in large language model (LLMs) inference scenarios, an adaptive routing strategy (MA-Router) with multi-modal attention mechanism is proposed in this paper. By constructing a hierarchical data enhancement framework and a two-channel attention routing network, fine-grained perception of problem complexity and optimization of dynamic model selection are realized. Firstly, a data enhancement method based on syntax tree reconstruction and knowledge graph extension is designed to effectively alleviate the sparsity and domain bias of training data. Secondly, a semantic-grammatical dual-modal feature extraction module is proposed, which uses DeBERTa-v3 to encode semantic information and capture dependency syntactic structure with graph neural network (GNN). Furthermore, the cross-attention mechanism is introduced to integrate multi-modal features, and the model capability-problem complexity bipartite graph is constructed for relational reasoning. Finally, hierarchical reinforcement learning Agent dynamically adjusts the routing threshold to realize the environment adaptive decision optimization. In the three benchmark tests of MMLU, MT-Bench and GSM8K, compared with the optimal baseline model RouteLLM, MA-Router reduced the call rate of strong model (GPT-4) by 32.7% (CPT(80%) to 26.15%), and the inference cost was reduced by 2.8 times. At the same time, the response quality of 95.2% was maintained (APGR=0.682). Ablation experiments show that syntactic feature channels and reinforcement learning modules contribute 14.2% and 29.3% performance gains, respectively. This research provides an efficient and scalable solution for multi-model collaborative reasoning system, which shows significant application value in real-time scenarios such as customer service dialogue and code generation.

Read the paper · More papers on PaperTik