Towards Dual Optimization of Efficiency and Accuracy:Hyper Billion Scale Model Computing Platform for Address Services

Dongxiao Jiang, Hailin Sun, Jiahua Lin, Jie Jiang, Xing Sun, Yanchi Li, Bocheng Xia, Siyuan Xiao, Yulong Liu · 2023

Whether intelligent algorithm services can be successfully implemented into production for enterprises mainly depends on: How to simultaneously achieve dual optimization of efficiency and accuracy at a rational low cost. In this paper we demonstrate an applied computing platform for global user addresses of J&T Express Ltd., to introduce some well-thought-out designs that enable address services to be not only efficient but also precise. These original designs include: (1) a complementary model composed of BERT for deep phased learning and GNN for continuous online learning; (2) biaffine layer to capture the relations between time and space that can improve address representation ability; (3) the last hidden layer is reformed and shared to GNN's construction for embedding consistency; (4) Qdrant is used as an high-efficient vector retrieval engine for inference performance; and (5) intensive computing tools (e.g., TensorRT, Triton) from NVIDIA is deeply utilized to accelerate training and inference. More than one year operation experience has verified the effectiveness, online evaluation indicators such as accuracy rate, inference response time and computing expense in section 5 illustrate the excellent performance even for hyper-billions scale and multilingual queries.

Read the paper · More papers on PaperTik