Research on Deep Data Annotation Methods Based on Expertise in the Power Industry
Nan Jiang, Chenyu Liu, Shaofei Yu · 2025
This study presents the development of an automated annotation tool for energy industries such as power, leveraging large language models to address challenges in processing unstructured text data efficiently and accurately. The tool comprises two main subsystems: data cleaning and preprocessing, and data generation. The data cleaning subsystem normalizes multi-format documents, converting them into structured formats suitable for annotation. The data generation subsystem, based on a large language model, includes modules for text, table, terminology, and image annotation, enabling the automatic generation of diverse question-answer pairs. The tool's performance was evaluated on 574 documents of eight common types of documents. Results show the tool's strong annotation capability, achieving high accuracy and coverage rates. However, limitations were observed in areas such as term recognition, context understanding, and multimodal content handling, highlighting avenues for future improvement. This automated annotation tool offers a scalable, cost-effective solution for data annotation in the power industry and lays the groundwork for enhancing data-driven decision-making and knowledge extraction across similar industries.