Mamba-Hybrid Language Models: Advancing Detection of AI-Generated Text in Academic Contexts

Manish Prajapati, Santos Kumar Baliarsingh, Shuvam Das, Amiya Kumar Dash, Arup Sarkar, Roshni Pradhan · 2025

The emergence of ChatGPT, an advanced generative artificial intelligence (AI) tool, has posed significant challenges to academic integrity in educational environments. This paper explores effective approaches and essential strategies to address these challenges, ensuring that educators can uphold standards of honesty and authenticity in student work. This study introduces an efficient language model using Mamba-Hybrid Language Models (MHLMs) algorithms to detecting AI-generated text. In the preprocessing stage, text was standardized through various steps such as lowercase conversion, tokenization, stop-word removal, feature extraction, digit removal, and space elimination to ensure high data quality. Our investigation into distinguishing Essays generated by Large Language Models (LLMs) from those written by students yielded significant findings. Mamba achieves state-of-the-art results on a diverse set of domains, where it matches or exceeds the performance of strong Transformer models. Our results show that while pure Selective state-space models (SSMs) based models match or exceed Transformers on many NLP tasks, MHLMs models lag behind Transformer models on tasks which require strong detection of AI generated and human generated Essays. In contrast, we found that the 10B-parameter MHLMs model outperformed the 8B-parameter Transformer across all standard tasks we evaluated, with an average improvement of +1.25 points. Additionally, MHLMs are projected to be up to 5 times faster in token generation during inference.

Read the paper · More papers on PaperTik