Extracting Information from Medieval Notarial Deeds.

Charlene Ellul, Charlie Abela, Joel Azzopardi · OAR@UM (University of Malta) · 2018

The Notarial Archives in Valletta houses a collection of Latin Notarial deeds that has not been exploited yet. In this paper, Machine Learning techniques are proposed and implemented to extract entities such as people, place names, dates, deed types and keywords from these historical texts. Both supervised and unsupervised techniques are considered and compared with baseline models. Experimental results on a subset of these documents are already showing results that outperform the baselines for Latin text such as those from CLTK. Evaluation was carried out using indexes of four published Notarial Registers.

Read the paper · More papers on PaperTik