Adversarial face recognition and phishing detection using multi-layer data fusion
Harry Wechsler, Venkatesh Ramanathan · 2012
This thesis addresses digital identity for biometric / face recognition screening and cyberspace security subject to denial and deception characteristic of adversarial behavior. The adversarial aspect concerns defense and offense operations that involve impostors and identity theft. Denial and deception correspond to occlusion and disguise for biometrics, while for cyberspace security they correspond to spoofing and obfuscation. To prevent or mitigate the impacts of adversarial behavior from offensive attacks this thesis proposes the use of multi-layer data fusion. Multi-layer aspect of fusion refers to features, representations, algorithms, decision-making, adversarial aspects and their purposeful combinations. This novelty, feasibility, and utility of our research is illustrated in the physical and cyber worlds: (i) robust face recognition in the presence of occlusion and disguise, and (ii) phishing detection to prevent identity theft through spoofing and obfuscation. The novel face recognition methodologies include: (i) Adaptive and Robust Correlation Filters (ARCF) built around match filters and recognition-by-parts, and (ii) hybrid anthropometric and appearance based biometric authentication using boosting for feature level fusion and backpropagation learning for decision level fusion. The cluster and strength of the ARCF correlation peaks indicate the confidence in the face authentications. Experimental evidence using the AR benchmark database shows that our methods are highly reliable in the presence of occlusion, disguise, and illumination, expression and temporal variability. The novel phishing detection methodologies address: (i) phishing email detection using semantic topics and Probabilistic Latent Semantic Analysis (PLSA), boosting, and Co-Training for both labeled and unlabeled examples, (ii) phishing website detection using Latent Dirichlet Allocation (LDA) and boosting, and (iii) impersonated entity discovery using LDA, boosting, and Condition Random Field (CRF). The phishing detection methodology handles the adversarial use of synonyms, polysemy (words with multiple meanings) and other linguistic variations. In addition, the same methodology requires only a small percentage of data to be annotated thus saving time, labor, and avoiding errors incurred during human annotation. The phishing website detection methodology is device and language neutral. The impersonated entity discovery methodology automatically extracts the entity the attacker is trying to spoof. This helps service providers to collaborate with each other to exchange attack information and protect their customers. Experimental results on SPAM Archive, which is one of the largest public corpus, show that our phishing detection methodology outperforms state of the art phishing detection methods.