Security Assessment and Generation Improvement Strategies for Large Language Models
Yu Zhang, Yongbing Gao, Lidong Yang · International Journal of Asian Language Processing · 2025
This study evaluates the performance of mainstream large language models (LLMs) in Chinese security generation tasks, examines the potential security risks associated with these models, and proposes strategies for mitigating these risks. To this end, we developed the multidimensional security question answering (MSQA) dataset and the multidimensional security scoring criteria (MSSC). This study compares the performance of three models across six distinct security tasks. Pearson correlation analysis was conducted using GPT-4 and questionnaires, while automatic scoring was implemented using GPT-3.5-Turbo and Llama-3. Experimental results reveal that ERNIE Bot excels in ideology and ethics evaluation, ChatGPT demonstrates strong performance in assessing rumors, false information and privacy security, and Claude performs well in evaluating factual fallacies and social biases. Additionally, the fine-tuned model showed effectiveness in security scoring tasks, and the proposed Security Tips Expert (ST-GPT) successfully mitigates security risks. Despite the promising results, all models exhibit inherent security risks. Based on these findings, we recommend that both domestic and international models adhere to the legal frameworks of their respective jurisdictions, minimize AI hallucinations, continuously expand training corpora, and undergo regular updates and iterations to enhance their reliability and safety.