Quantifying Corruption: Computational Analysis of Textual Transmission in the Diwan of Mah Laqa Bai Chanda
Akash Arsh · 2026
For over a century, the textus receptus of the Diwan of Mah Laqa Bai Chanda (c. 1768–1824)—one of the first women to compile a complete diwan in Urdu—has been based on the 1906 lithograph Gulzar-e-Mah Laqa, edited by Ghulam Samdani Khan Gauhar and published at Nizam ul Matabe, Hyderabad. Systematic collation with the three surviving lifetime manuscripts demonstrates that this edition does not constitute an authentic witness to Mah Laqa Bai's text. A systematic collation of the three surviving lifetime manuscripts—the British Library copy (BL, 1799), the Telangana Oriental Manuscript Library copy (TO, 1811), and the Salar Jung Museum copy (SJ, 1818)—against this print edition reveals that the print text is an ideological editorial reconstruction rather than an authentic witness. Using a NeighborNet phylogenetic network in SplitsTree4, we show that the 1906 print edition is isolated at a distance of twelve times the total spread of the manuscript tradition. While the poet altered 8.2% of couplets substantially across nineteen years of documented authorial revision, the print editor modified 50.2% of couplets. This paper introduces the Dakhni Retention Rate (DRR), a formalised token-level lexical tracking metric designed to quantify the systemic suppression of regional linguistic variants during the transition from manuscript to print. We demonstrate that the 1906 text represents a systematic linguistic normalisation of SJ 1818. The complete collation dataset underlying all quantitative results reported here is openly archived at https://doi.org/10.5281/zenodo.20587113 and may be used to reproduce all findings independently. The DRR framework offers a generalizable methodology for digital humanities scholars to track vernacular suppression and canon formation across other South Asian print traditions.