A Diachronic Corpus of Modern Indonesian Literature Between 1920 and 2000

Mohammad Rokib, Noorhidawati Abdullah, Moh. Mudzakkir, Lutfiyah Alindah · Journal of Open Humanities Data · 2026

This dataset presents a diachronic corpus of modern Indonesian literature comprising 58 literary works published between 1920 and 2000. The corpus contains approximately 1.97 million tokens and captures three orthographic periods in Indonesian literary writing. Texts were digitized from printed sources using Optical Character Recognition (OCR), then semi-manually cleaned, tokenized, orthographically standardized, lemmatized, and annotated for historical orthography. Each work is documented using a metadata schema adapted from Dublin Core, supplemented with fields for literary genre, orthography system, and annotation level. The repository provides open corpus extracts and accompanying metadata for works eligible for public distribution. The dataset supports research in corpus linguistics, language change, Indonesian literary studies, and digital humanities.

Read the paper · More papers on PaperTik