BayesShrink threshold estimation-based multi-metric space clustering algorithm for power grid project cost data
Handong Lu, Ming Fang, Yuanliang Zhang, Xin Li · Results in Engineering · 2026
Accurate analysis of power grid project cost data is crucial for project investment control and resource allocation optimization. However, such data often contain substantial noise due to measurement errors, design changes, and market fluctuations. Their high-dimensional, multi-source, and heterogeneous characteristics can trigger the “curse of dimensionality,” which reduces the accuracy and robustness of traditional clustering methods. Existing approaches predominantly rely on single-distance metrics and the assumption of spherical-like distributions, rendering them ineffective for processing real-world cost data that are multimodal, noisy, and non-uniformly distributed. In highly noisy environments, these methods often suffer from over-smoothing or insufficient denoising, leading to clustering results that are susceptible to noise interference and exhibit limited cluster structure recognition. To address these challenges, this paper proposes a multi-metric space clustering algorithm for power grid project cost data based on BayesShrink threshold estimation. First, a cost data system covering five dimensions - including standards, specifications, and pricing information - is constructed. Subsequently, an innovative method is introduced: sliding-window dynamic noise variance estimation combined with generalized Gaussian distribution prior optimization. This enhances the BayesShrink threshold selection mechanism and enables adaptive suppression of high-frequency noise. Finally, by integrating L 1 /L 2 distances, edit distances, and word-vector cosine distances within a multi-metric space, a multi-metric clustering framework is established. This approach overcomes the limitations of traditional methods regarding data distribution patterns and allows simultaneous processing of multimodal cost data containing numerical, textual, and geographic information. Experimental results demonstrate that across four datasets - equipment, labor, geographic, and synthetic - the proposed algorithm maintains an effective data retention rate exceeding 97.38% even under strong noise interference. It effectively preserves genuine signals while avoiding the “over-cleaning” that would cause the loss of authentic cost features. The clustering results achieve a BSS/WSS ratio as high as 9.74 and a maximum silhouette coefficient of 0.83. These metrics significantly outperform those of advanced benchmark algorithms, indicating superior inter-cluster separation and intra-cluster compactness. The algorithm prioritizes flagging anomaly clusters or outliers for manual review, enabling rapid identification of “suspected anomaly records.” This supports more accurate and interpretable cost estimation for power grid projects, supplier evaluations, and regional cost optimization.