Software demonstration: meaning-based querying of historical corpora with MacBERTh
Lauren Fonteyn, Enrique Manjavacas · Zenodo (CERN European Organization for Nuclear Research) · 2022
This is an abstract for a software demonstration at DHBenelux 2022. This software demonstration will focus on MacBERTh, a BERT-based model pre-trained on Early Modern and Late Modern English (3.9B (tokenized) words, time span: 1450-1950; Manjavacas & Fonteyn 2021, 2022). We will demonstrate how MacBERTh may help researchers (i) access and (ii) analyse the semantic information encoded in linguistic corpus data in a (semi-)automatic way.