NLMaps Web: A natural language interface to OpenStreetMap

Minghini, M. [Hrsg.], Ludwig, C. [Hrsg.], Anderson, J. [Hrsg.], Mooney, P. [Hrsg.], Grinberger, A.Y. [Hrsg.] · Florida International University Digital Commons (Florida International University) · 2021

Nominatim [1] and Overpass [2] are powerful ways of querying OSM, but the Overpass Query Language is somewhat impractical for quick queries for unfamiliar users.In order to query OSM using natural language (NL) queries such as "Show me where I can find drinking water within 500m of the Louvre in Paris", Lawrence and Riezler [3] created the first NLMaps dataset mapping NL queries to a custom machine-readable language (MRL), which can then be used to retrieve the answer from OSM via a combination of queries to Nominatim and Overpass.They extended their dataset in a subsequent work by auto-generating synthetic queries from a table mapping NL terms to OSM tags -calling the combined dataset NLMaps v2.[4] The proposed purpose of these datasets is training a parser that can parse NL queries into their MRL representation, as done in [4][5][6][7].The main aim of this work was to build a web-based NLMaps interface that can be used to issue queries and to view the result.In addition, the web interface should enable the user to give feedback on the returned, either by simply marking the parser-produced MRL query as correct or incorrect, or by explicitly correcting it with the help of a web form.This feedback should be directly used to improve the parser by training it in an asynchronous online learning procedure.After observing that parsers trained on NLMaps v2 perform poorly on new queries, an investigation into the causes for this revealed several shortcomings in NLMaps v2, mainly: (1) Train and test split are extremely similar limiting the informativeness of evaluating on the test split.(2) Various inconsistencies exist mapping from NL terms to OSM tags (e.g."forest" sometimes mapping to natural=wood, sometimes to landuse=forest).(3) The NL queries' linguistic diversity is limited since most of them were generated with a very simple templating procedure, which leads to parsers trained on the data not being very robust to new wordings of a query.(4) In a similar vein, there is only a small amount of different area names in NLMaps v2 with the names "Paris", "Heidelberg" and "Edinburgh" being so dominant that parsers are biased towards producing them.(5) Some generated NL queries are not a good representation of natural language, which makes them counter-productive learning examples.(6) Usage of OSM tags is sometimes incorrect, which affects the usefulness of produced parses.The detailed analysis is used to eliminate some of the shortcomings -such as incorrect tag usage -from NLMaps v2.Additionally, a new approach of auto-generating Will, S. (2021).

Read the paper · More papers on PaperTik