Lessons from a Jihadi Corpus
David B. Skillicorn · 2012
We analyze the posts in the Islamic Awareness forum, using models for frequent words (content), for Salafist-Jihadist language, and for deception. These last two models each produce a single-factor ranking enabling, in each case, the most useful subset of posts to be selected for further analysis. Posts that rank highly for Salafist-Jihadist language rank low for deception, suggesting that faking extremist websites is probably an ineffective strategy. The process described here is a template for analysis of many kinds of open-source corpora where language models of what makes posts interesting are known.