Silent LLMs: Using LoRA to Enable LLMs to Identify Hate Speech

Yuzhang Wang, Peng Yu, Guoxiu He · 2024

The detection of hate speech on social networks presents significant challenges due to its increasingly subtle nature. The advent of large language models (LLMs) has revolutionized text understanding and generation, presenting novel avenues for hate speech detection. This study evaluates the performance of the ChatGLM model in detecting hate speech through a small sample prompt method. Our analysis uncovers three key limitations when employing the LLMs to directly generate answers: inconsistency in output format, the illusion of comprehensiveness inherent in LLMs, and the inability to respond due to security concerns. To mitigate these limitations, we investigate four LLMs: LLaMA, Llama-2, Llama-3, and ChatGLM-3. These models are equipped with Multi-Layer Perception (MLP) and Low-Rank Adaptation (LoRA), which specifically tailored for hate speech detection. Extensive comparisons with baseline models are conducted across three hate speech datasets with six classification tasks. Our findings demonstrate that our improved LLMs can surpass traditional methods in detecting hate speech, highlighting their potential for further improvement and refinement in addressing this critical societal issue.

Read the paper · More papers on PaperTik