Lexicography Comes into Contact with Natural Language Processing

Pedro A. Fuertes-Olivera · 2025

In today&s;s world, lexicography is totally dependent on Natural Language Processing (NLP) and humans&s; (mostly lexicographers&s;) own knowledge. NLP tools and methods are being used for automating the creation of lexicographic data that are later analysed and interpreted by humans (lexicographers and non-lexicographers), who give them their blessing. This process is called "post-editing lexicography" and is in the vanguard of lexicographic theory and practice. Until the coming of age of Large Language Models (LLMs), the process of automation rested on a range of tools, methods and procedures that started to be developed around 20 years ago and were constantly being updated and improved. The launch of ChatGPT in November 2022 stirred the lexicographic waters and opened a lexicographic debate on its merits as a "lexicographer". Although research published so far offers mixed results regarding its lexicographic possibilities, I think that chatbots will be very influential and will shape the future of lexicography under the guidance of human lexicographers. This accounts for my conviction that in the near future lexicography will be mostly shaped by the concept of post-editing lexicography, because this facilitates work, enhances productivity, and reduces costs. As in the rest of the book, I am not focusing on dictionaries or other reference works, but on the lexicographic square, especially on lexicographic data. After reviewing some of the well-known tools and procedures used for automating lexicography, as represented by state-of-the-art NLP tools, methods and procedures, the chapter offers a methodology for working with LLMs in order to extract appropriate data for storing in the database of the lexicographic editing software of a lexicographic project, e.g. that of DIDES . The methodology is based on a combination of three ideas - developed in the field of Computer Science - that are jointly applied, with the aim of reducing possible hallucinations and increasing LLM output: (a) semantic entropy, (b) lexical entropy, and (c) multi-agent systems (MAS).

Read the paper · More papers on PaperTik