Prompt-Based Generation Strategy for Imbalanced Information Security Rating Dataset Augmentation
Yuna Han, Simon S.Y. Shim, Deva Kumar Gajulamandyam, Hangbae Chang · 2024
As information leakage due to internal and external threats continues to escalate in the era of technological dominance, safeguarding confidential information managed by companies and organizations has become a critical issue across industries. Information security rating involves evaluating the sensitivity of information by identifying key details within it, typically following decision-making processes to classify data as either confidential or public. While existing research on security rating has primarily focused on advancing training models, there has been limited attention to addressing the issue of imbalanced secret data, which can lead to overfitting. To mitigate this issue, this paper proposes a prompt-based technique to augment the minority “Secret” class, specifically for the WikiLeaks dataset. The proposed methodology consists of three phases: first, enhancing state-of-the-art augmentation prompts with additional constraints; second, tuning few-shot generation prompts using a sliding-window approach centered on the frequency-weighted median to generate augmented data that aligns with the distribution of the training dataset; and third, training a fine-tuned large language model using Llama 3.1 8B with LoRA to perform security rating on the augmented data. Experimental results demonstrate that the proposed methodology yields incremental improvements, achieving 98% accuracy with one-shot generation and 99% accuracy with few-shot generation. To the best of our knowledge, this paper contributes by being the first study to generate highquality “Secret” data for the domain-specific task of information security rating and enhances security rating performance by introducing our proposed augmentation method. In future work, we aim to further optimize the fine-tuning of large language models to apply security rating data for real-world scenarios.