Parallel hypernym relation extraction based on partition index dividing

Juxin Yin, Lei Pan · 2023

Aiming at the shortcoming of existing pattern methods in handling massive data, a parallel hypernym relation extraction method on Spark is proposed. To improve the extraction accuracy, Combined with Spark's RDD programming model, an improved credibility algorithm(ppmit) is designed to identify the inverse hypernym relation; to address the data skew problem when calculating the credibility value of the hypernym relation, a data partitioning strategy - PID(Partition Index Dividing) algorithm is proposed to calculate the partition balance, add partition index to the overflow data, redivide the data, and ensure the partition data balance, thereby reducing the calculation time. Experiments conducted on the Chinese Wikipedia dataset show that the proposed method can guarantee the extraction accuracy and effectively improve the operational efficiency of the pattern extraction method.

Read the paper · More papers on PaperTik