Word Net-enriched text classification with compressed distance based word networks
Aarish Shah Mohsin, Mohammed Tayyab Ilyas Khan, Nadeem Akhtar · 2025
The use of deep neural networks (DNNs) in text categorization is restricted in out-of-distribution (OOD) and few shot scenarios because they typically require extensive labeled datasets and significant computing resources. A recent study presented a compressor-based, non-parametric method that uses gzip and k-nearest neighbors to achieve com petitive performance. Using word windows, WordNet-based word embeddings, and word network centrality features obtained from the normalized compression distance metric, we provide a unique text representation method that builds upon this. Our approach, which is parameter-free, achieves competitive results on in-distribution datasets and performs competitively to non-pretrained DNNs and pre-trained models like BERT in OOD and few-shot settings, as demonstrated by experiments conducted on six in-distribution and five OOD datasets, including low-resource languages. Notably, we use compressed distance in our word network building to capture semantic word similarities, with centrality measures enhancing the text representation. Our work offers a lightweight alternative to DNNs, excel ling in OOD and few-shot scenarios while maintaining robust in-distribution performance.