Computer-based solutions to certain linguistic problems arising from the Romanization of Arabic names (Vols. I and II)
Paul Roochnik · 1993
If the concept of personal names represents a host of linguistic issues, the Romanization of Arabic names is particularly complex. Most Arabic Romanization schemata are hybrid systems of (a) transliteration, i.e., the substitution of one Roman letter for each Arabic letter, and (b) transcription, i.e., the conversion of the source phonemes into Roman letters and extra-alphabetical symbols such as the apostrophe. The very multiplicity of systems leads to ambiguity and confusion. The corpus under investigation in this research, called 'GDATA', contains over 100,000 Arabic names randomly selected from eighteen Arab countries, that have been Romanized without regard to any particular schema. For this reason, one Arabic name may be represented in as many as fifty different ways. This thesis is a synchronic study which examines a wide-ranging specimen of GDATA names and their variants, and tries to detect the patterns or categories of orthographic variation. It is then necessary to determine whether or not the variations discovered from the analysis of this sample apply to GDATA as a whole: can global variation patterns be extrapolated? A computer program called STANDARD.exe has been written which attempts to standardize the spelling of a Romanized Arabic name by encoding the variations discerned in GDATA analysis. STANDARD.exe produces reasonable results, which suggests that while some aspects of the program need refinement, the program as a whole accomplishes its objective. It is concluded that further analysis be done to uncover more patterns in the variants of Romanized Arabic names, and that wherever possible, computational tools and 'intelligent' approaches such as neural networks, be deployed to that end.