Applying Text Analytics to Derive Value from Blog Posts

Ivan Tang · DigitalCommons@Pace (Pace University) · 2019

Text analytics consists of examining the unstructured data contained in natural language using various methods, techniques, and tools.This form of analytics has been growing in popularity due in part to social media and its ability to record public opinion on topics of interest.Blogs are a powerful form of social media in that they are easy to create and maintain and enable any Internet user to become a content creator.Deriving meaning from such open source content can be applied to many domains including marketing, product development, and intelligence gathering.We apply text analytic tools to derive value from a cybersecurity blog authored by a prominent figure in the cybersecurity field.The blog is popular among cybersecurity aficionados who have an interest in the author's expert opinion on cybersecurity topics such as vulnerabilities, patches, and privacy.The author has been actively blogging on this platform since 2004.We implement website scraping tools and generate a text data set by harvesting text from the cybersecurity blog.We process and analyze this text using the Natural Language Toolkit in Python.We process the removing stop-words and punctuation, converting all letters to lowercase, and removing inflectional word endings.We apply word frequency measures to analyze the processed text and reveal the most frequent topics discussed by the blog's author.In addition, we map these frequent topics to a timeline to uncover meaningful patterns in topics of interest across the blog and how discussions of cybersecurity topics have evolved over 15 years in this author's blog.

Read the paper · More papers on PaperTik