A STATISTICAL APPROACH FOR PATTERN SEARCH IN INDUS WRITING
Nisha Yadav, M. N. Vahia, Iravatham Mahadevan, Hrishikesh Joglekar · IJDL. International journal of Dravidian linguistics · 2008
We search for potentially grammatical patterns in the Indus writing based on the concordance of Mahadevan (1977). We make no assumptions about its structure or meaning. We only attempt to check if the Indus writing is meaningfully structured with specific rules to code useful information. In order to avoid possible errors in interpretation due to incompletely read or multiple copies of a piece of writing and other possible sources of error, we create an Extended Basic Unique Data Set (EBUDS) based on the original electronic concordance of Mahadevan (1977). EBUDS consists of completely read single line texts on any side of the writing material. We exclude multi-lined text or partially read text. This gives us a set of 1548 lines of data consisting of 7000 signs. We show that the ordering of the signs in the writing is much more significant than random association. The unit length of information is 2, 3 or 4 signs at a time. We then study the most frequently occurring two, three and four sign combinations in EBUDS. We find that in many cases, most common two-sign combinations are also parts of most common three-sign combinations, which in turn also appear in four-sign combinations. However, in texts with just 2, 3 or 4 signs we do not often find these frequent two, three or four sign combinations. We therefore conclude that while the information is given in units of two, three or four signs, these are more like phrases where an additional sign is required to complete the grammatical structure.