Gatekeeper to save COGS and improve efficiency of Text Prediction
Nidhi Tiwari, Sneha Kola, Milos Milunovic, Siqing Chen, Marjan Slavkovski · 2023
The text prediction (TP) workflow in editor calls a Large Language Model (LLM), after a character is typed by the user to get subsequent sequence of characters.The confidence score of the prediction is used for filtering the results to ensure that only correct predictions are shown to user.As LLMs require massive amount of computation and storage, such an approach incurs high execution cost.So, we propose a Model gatekeeper (GK) to stop the LLM calls that will result in incorrect predictions at client application level itself.This way a GK can save cost of model inference and improve user experience by not showing the incorrect predictions.We demonstrate that use of a model gatekeeper saved ≈ 46.6% of COGS (Cost Of Goods Sold) for TP, at the cost of ≈ 4.5% loss in character saving.Use of GK also improved the efficiency (suggestion rate) of TP model by 73%.