Content-based File Type Identification

Kireet Bhat, Jason T. Lam, Farhana H. Zulkernine · 2018

Data is encoded and stored in a wide variety of file types having different schema. As a result, File Type Identification (FTI) has become essential in data pre-processing and security. FTI traditionally involves interpretation of the key signature bytes found in the beginning of each file. However, when these signature bytes are corrupted, the contents of the file must be analysed to correctly identify the file type. Recently, studies leveraging machine learning algorithms have seen great success in content-based classification. In addition, when the signature bytes are not damaged, Commercial Off The Shelf (COTS) systems like Libmagic are preferred. We present an end-to-end framework combining signature and content-based file type classification. We utilize Libmagic in combination with machine learning algorithms, implemented in Spark, to accurately identify file types. The results show that our system can successfully identify the data schema and file types using a pool of workflow templates.

Read the paper · More papers on PaperTik