Improving Toponym Recognition Accuracy of Historical Topographic Maps

Kenzo Milleville, Steven Verstockt, Nico Van de Weghe · 2020

Scanned historical topographic maps contain valuable geographical information.Often, these maps are the only reliable source of data for a given period.Many scientific institutions have large collections of digitized historical maps, typically only annotated with a title (general location), a date, and a small description.This allows researchers to search for maps of some locations, but it gives almost no information about what is depicted on the map itself.To extract useful information from these maps, they still need to be analyzed manually, which can be very tedious.Current commercial and open-source text recognition tools underperform when applied to maps, especially on densely annotated regions.They require additional processing to provide accurate results.Therefore, this work presents an automatic map processing approach focusing mainly on detecting the mentioned toponyms and georeferencing the map.Commercial and open-source tools were used as building blocks, to provide scalability and accessibility.As lower-quality scans generally decrease the performance of text recognition tools, the impact of both the scan and compression quality was studied.Moreover, because most maps were too large to effectively process as a whole with state-of-the-art commercial recognition tools, a tiling approach was used.The tile size affects recognition performance, therefore a study was conducted to determine the optimal parameters.First, the map boundaries were detected with computer vision techniques.Afterward, the coordinates surrounding the map were extracted using a commercial OCR system.After projecting the coordinates to the WGS84 coordinate system, the maps were georeferenced.Next, the map was split into overlapping tiles, and text recognition was performed.A small region of interest was determined for each detected text label, based on its relative position.This region limited the potential toponym matches given by publicly available gazetteers.Multiple gazetteers were combined to find additional candidates for each text label.Optimal toponym matches were selected with string similarity metrics.Furthermore, the relative positions of the detected text and the actual locations of the matched toponyms were used to filter out additional false positives.Finally, the approach was validated on a selection of 1 : 25 000 topographic maps of Belgium from 1975-1992.By automatically georeferencing the map and recognizing the mentioned place names, the content and location of each map can now be queried.

Read the paper · More papers on PaperTik