Automated Gender Identification for Arabic and English Handwriting
Amira E. Youssef, Ahmed S. Ibrahim, Amos Lynn Abbott · 2013
This paper is concerned with off-line handwriting analysis for the purpose of identification of the writer's gender. Such identification is a useful tool in forensic handwriting analysis. Previous studies have indicated that handwriting by males and females tend to exhibit distinctive characteristics, even across different languages and cultures. We have investigated the hypothesis of language independence by applying machine-learning techniques to handwriting samples in two languages. The two languages, Arabic and English, are representative of other languages that use the same character sets. In particular, Arabic uses the same characters as Urdu and Persian, while English is based on the same character set as French, Spanish, Italian, and other Latin languages. Using a database in which 282 individuals provided handwriting samples in both Arabic and English along with gender information, several classifiers were implemented and compared. Each classifier utilized support vector machines (SVM) to identify the gender of the writer. For classifiers that were trained separately for the two languages, an accuracy of 68.6% was observed for Arabic, while an accuracy of 85.7% was observed for English. When trained using handwriting samples for both languages, an accuracy of 74.3% was observed. These results indicate that language-independent analysis can eventually be employed in forensic analysis.