An Unsupervised Ontology Construction Method Based on Pre-trained Language Model
Hanqi Zheng, Guige Ouyang, Yongzhong Huang · 2025
With the fast growth of text data, the importance of automatic ontology construction has grown significantly. The proposed article provides a novel approach by applying pre-trained language models to automatically construct the ontology. The framework consists of two sequential phases: automatic concept discovery and automatic relation discovery. In the context of automatic concept discovery, instances and their embedding vectors are extracted first through Named Entity Recognition (NER). Then, the unsupervised affinity propagation (AP) clustering algorithm is applied to classify these embedding vectors, resulting in the discovery of the concepts. A denoising method is discussed to obtain higher accuracy with respect to the concepts obtained to reduce noise caused by complete clustering. Related to automatic relation discovery between concepts (mentioned in the next section), the process generates inexplicable concepts that are contextually similar based on the embedded vectors of instances mapping over the two entities. This can enable unsupervised automatic discovery of relations between the contextually related concepts. This method shows a certain feasibility and achieves early effectiveness in unsupervised automatic ontology construction with the experimental results.