Expertise Identification Using Transformers
T M Sreekanth, Ajith K. John, Rohitashva Sharma, Prathamesh Berde, C. S. R. C. Murthy · 2024
Expertise Identification involves extracting expertise/skills of a person from a set of documents related to his work. This has many important applications in large multi-disciplinary organizations such as ours. Most of the existing approaches for Expertise Identification apply unsupervised learning techniques such as those based on TF-IDF to extract keyphrases from the documents, which are then used as expertise. However, keyphrases represent the main ideas covered in a document, whereas expertise should be more domain-specific and detailed to be practically usable. Moreover, these unsupervised learning techniques fail to extract expertise which are not explicitly present within the documents. We cast Expertise Identification problem as an abstractive text generation problem, and use supervised learning with transformer based language models to solve this problem. We also show that existing metrics that are based on exact syntactic match between ground truth expertise and the predicted expertise are not suitable for performance evaluation of Expertise Identification techniques. Instead, we propose to use an evaluation metric based on semantic similarity. Experiments reveal that our approach based on transformers clearly outperforms the unsupervised learning techniques.