Feature Extraction and Compliance Classification of Text Files Using Large Language Models

Xiang Liu, Yanghao Liao, Zusheng Zhang, Jingcheng Hu, Lifeng Huang · IEEE Transactions on Computational Social Systems · 2025

In industries such as finance, healthcare, and new energy vehicles, data classification and grading standards ensure regulatory compliance and protect sensitive information. However, automating text file classification under these standards presents several challenges. Traditional machine learning and deep learning approaches require large labeled datasets, which are often scarce. Existing classification methods are typically domain-specific, limiting cross-domain adaptability. Moreover, many approaches simply categorize documents as regulatory or nonregulatory and assign security levels, but fail to map them accurately to specific rules. To address these challenges, this article proposes prompt-driven grading and classification algorithm (PGCA), a prompt learning-based method for text classification and grading. PGCA integrates a structured feature repository and SQL-inspired prompt templates to efficiently extract and match features from text documents, establishing mappings between text, and classification rules and grading standards. Furthermore, the integration of a preclassification strategy enables the filtration of irrelevant rules, thereby substantially reducing computational overhead. Experiments show that PGCA achieves classification accuracy between 95.0% and 99.0%, outperforming baselines such as TsF-KNN, Gen-DT, bt-SVM, AGCRCNN, and AC-BiLSTM by 4%–25%. Additionally, the preclassification stage cuts computational costs by 63.7% while keeping accuracy loss to within 1%.

Read the paper · More papers on PaperTik