NetOps-VL: Leveraging Multimodal LLM for Enhanced Network Operations

Shun Lu, Jing Shen, Xu Guan, Jie-ling Yao, Ying Si, Longgang Zhao, Bing Qian · 2024

Multimodal large language models (LLMs) have demonstrated powerful abilities in interpreting various common image scenarios, attributed to their extensive parameters and varied datasets. However, in specific domains like Network Operations (NetOps), their effectiveness is limited by insufficient knowledge about network equipment features, stemming from a lack of domain-specific expertise. Meanwhile, small visual models with convolutional layers, despite having fewer parameters and faster inference speeds, often require separate training for each NetOps visual task, and their cost may not always be lower than that of general LLMs. To tackle this challenge, we constructed a Visual Question Answering (VQA) dataset focused on image recognition and OCR tasks for network infrastructure within the NetOps domain. Subsequently, we employed GPT-4 as a teacher model to generate Chain of Thought (CoT) prompts, and utilized the created dataset to fine-tune the original multimodal LLM. This approach ensures that the enhanced multimodal LLM attains proficiency in NetOps domain knowledge. We propose the method for generating CoT prompts as “multimodal fine-tune-CoT”. The fine-tuned LLMs is referred to as NetOps-VL. In our experiments, we conducted multiple tests to validate multimodal fine-tune-CoT's effectiveness. The results indicate that NetOps-VL exhibits some advantages over other methods on the test dataset. Moreover, NetOps-VL improves NetOps image and text comprehension without sacrificing its general visual capabilities.

Read the paper · More papers on PaperTik