Fast file-type identification

Irfan Ahmed, Kyung-suk Lhee, Hyunjung Shin, Manpyo Hong · 2010

This paper proposes two techniques to reduce the classification time of content-based file type identification. The first is a feature selection technique, which uses a subset of highly-occurring byte patterns in building the representative model of a file type and classifying files. The second is a content sampling technique, which uses a subset of file content in obtaining its byte-frequency distribution. Our initial experiments show that the proposed approaches are promising even the simple 1-gram features are used for the classification.

Read the paper · More papers on PaperTik