Identifying AI-Generated Text Sources via Linguistic Style Fingerprints
Wanyi Feng · 2025
With the widespread adoption of large language models (LLMs), AI-generated texts have rapidly proliferated across various domains, posing the identification of their source models as a pivotal challenge in AI governance. Traditional binary classification approaches, such as distinguishing “AI versus human,” are ill-equipped to handle the complexities of a multi-model ecosystem, underscoring the need for more nuanced and interpretable detection methods. This study proposes a multimodel text source tracing approach grounded in linguistic style fingerprints, introducing the TabPFN classifier for the first time to identify the origins of five prominent English generative models: ChatGPT, Claude, Gemini, ERNIE Bot, and Grok. We constructed a dataset encompassing ten distinct semantic tasks and developed an enhanced, multi-dimensional style feature system that incorporates TF-IDF lexical features, syntactic structures, lexical diversity, punctuation habits, and linguistic complexity. Without relying on deep semantic modeling, we demonstrated the efficacy of linguistic style in differentiating models and the efficiency of TabPFN in small-sample, high-dimensional tasks. The findings offer a highly scalable and practical solution for AI content tracing and platform governance.