Scalable Analysis of English Dictionary Files on HPCC Systems Big Data Platform
C Jayanth, U Adarsh, David de Hilster, Hugo Martinelli Watanuki, G. Shobha, Jyoti Shetty · 2024
This paper presents a scalable approach to address the challenge of enhancing and scaling English dictionaries, employing a scalable integration of NLP++ VisualText analyzers with the HPCC Systems Big Data platform. The methodology involves parsing English Wiktionary files to generate knowledge bases and dictionary files, deploying large XML files on the HPCC platform via the ECL-NLP++ plugin. This study focuses on utilizing ECL and HPCC System architecture to analyze files. The work demonstrates the effectiveness of HPCC Systems in handling large linguistic datasets and compares it with Hadoop MapReduce, highlighting the effectiveness of HPCC in processing real-time data and its user-friendly functionality The results show extraction of English words and their parts of speech from Wiktionary files. This research contributes to improving Natural Language Processing (NLP) applications and lays the groundwork for future enhancements, such as incorporating semantic information and exploring other linguistic variations.