Large language vision models for zero-shot handwriting recognition of historical herbarium labels

Matthias Körschens, Solveig Franziska Bucher, Christiane M. Ritz, Sebastian Gebauer, Jens Wesenberg, Christine Römermann · Ecological Informatics · 2026

Herbaria contain large numbers of conserved specimens with lots of information for biodiversity research, since they offer a track record of the morphology as well as temporal and spatial distribution of plant species worldwide. Besides the dried plant itself, a lot of additional information is usually provided with the herbarium specimens, typically captured in printed or handwritten labels, such as the date of collection, the location and the collector’s name. While, due to historical reasons, the specimens have been collected and labeled manually, considerable efforts are underway to digitize entire herbaria and therewith make the specimens available for analysis with automated methods. However, the extraction of information from handwritten labels is a considerable challenge, since the handwritings do not only differ from one collector to another, but they are also often in old types of writing (e.g., Sütterlin, an old German script). Therefore, they are often hard to decipher both manually and automatically, and barely any substantial consistent data of this kind exists to train state-of-the-art vision models. Since the location of the labels differs depending on the record, they need to be detected before the automated analysis of the writing, which also proved challenging in the past. In this work we show that state-of-the-art Large Language and Vision Models (LLVM) possess capabilities to extract such handwriting zero-shot, i.e., completely without training or fine-tuning, to a high degree of accuracy. Additionally, we show that the results can be refined and improved considerably by performing zero-shot detection of the labels beforehand. We evaluate our approach on two novel datasets, one containing handwritten and one printed labels, respectively, based on herbarium scans from the virtual herbarium of the flora of Germany. In our evaluations, the approaches achieve a mean similarity of 84.5% for handwritten, and one of 93.1% for printed labels. Thus, we conclude that still some evaluation is needed before the LLVMs can be fully applied to transcribe herbarium specimen labels, as sometimes the species taxonomies as well as the collection sites are not correctly identified. Still these models can support the transcription process in large collections. Our code and a graphical web application is publicly available under https://github.com/Atlas8008/herbarium_label_reader .

Read the paper · More papers on PaperTik