Improved of content-based file type identification algorithm
Junyong Luo · Jisuanji gongcheng yu sheji · 2011
On the basis of the content-based file type identification algorithm,both fixed and variable size window are adopted to extract statistic characteristic of files' binary content,which improves feature extraction in current algorithm.Feature selection is introduced into file type identification and a novel evaluation function combing feature width and stability is used for feature selection,which are used to establish models for different file types as standard to determine a tested file type.Our aim is not to use the structure and Key words of any specific file types as this allows the approach to be applied to general file types.Experiments show that the proposed approach improves the precision and recall of file type identification.