Designing Politics and IR Assessments in the Era of AI: An Empirical Investigation into ChatGPT’s Output Across Bloom’s Revised Taxonomy

Matthias Dilling, Leah Owen · Journal of Political Science Education · 2024

AI breakthroughs have sent shockwaves through the HE sector. Reactions have ranged from concern over academic integrity to excitement that large language models (LLM) might allow automating lower-stakes intellectual tasks. This paper tests some of these arguments using a mixed-methods research design to analyze an original dataset of 120 ChatGPT-generated essays on far-right politics based on prompts on different levels of Bloom’s revised taxonomy. Our findings echo results that LLM-generated essays generally achieve solid to good marks. However, we find no linear negative trend between the level on Bloom and the quality of essays produced. Essays’ quality was weaker on the lowest and highest levels of the taxonomy and strongest on intermediate levels, even when controlling for the version of ChatGPT and the round of data collection. The in-depth reading revealed distinctive limitations in argument construction and empirical detail—particularly regarding “hallucinations” and source misattributions—which could only be marginally improved through prompt engineering. In turn, ChatGPT was better at suggesting “places to look” for helpful information, meaning it might be more accurately compared to a signpost rather than a calculator. The implications for politics assessments and teaching are discussed.

Read the paper · More papers on PaperTik