A General-Purpose Multilingual Document Encoder
Onur Galoğlu Robert Litschko, Robert Litschko, Goran Glavaš · 2023
Massively multilingual pretrained transformers (MMTs) have tremendously pushed the state of the art on multilingual NLP and cross-lingual transfer of NLP models in particular.While a large body of work leveraged MMTs to mine parallel data and induce bilingual document embeddings, much less effort has been devoted to training general-purpose (massively) multilingual document encoder that can be used for both supervised and unsupervised documentlevel tasks.In this work, we pretrain a massively multilingual document encoder as a hierarchical transformer model (HMDE) in which a shallow document transformer contextualizes sentence representations produced by a stateof-the-art pretrained multilingual sentence encoder.We leverage Wikipedia as a readily available source of comparable documents for creating training data, and train HMDE by means of a cross-lingual contrastive objective, further exploiting the category hierarchy of Wikipedia for creation of difficult negatives.We evaluate the effectiveness of HMDE in two arguably most common and prominent crosslingual document-level tasks: (1) cross-lingual transfer for topical document classification and (2) cross-lingual document retrieval.HMDE is significantly more effective than (i) aggregations of segment-based representations and (ii) multilingual Longformer.Crucially, owing to its massively multilingual lower transformer, HMDE successfully generalizes to languages unseen in document-level pretraining.We publicly release our code and models. 1 .