Distributed Speaker Tracking in IoAuT Networks via Tempered Stein Variational Gradient Descent Sampling
Yuan Jing, Hao Song · IEEE Internet of Things Journal · 2025
This paper addresses the problem of distributed speaker tracking in wireless acoustic sensor networks, also known as the Internet of Audio Things (IoAuT). We propose a novel Tempered Stein Variational Gradient Descent (T-SVGD) algorithm for distributed fusion of local likelihood probability density functions (PDFs) and approximation of the posterior distribution of the speaker state vector. In the proposed approach, the Gaussian Mixture Model (GMM) is employed to approximate the posterior predictive distribution, and a tempered initial sampling distribution is designed to enhance SVGD sampling. The average consensus fusion method is then utilized to achieve distributed fusion of the local likelihood PDFs across the network nodes. Subsequently, the T-SVGD algorithm is applied to approximate and sample the posterior distribution, thereby enabling distributed state estimation of the speaker position. Unlike conventional particle-based methods, the proposed algorithm deterministically transports and optimizes the sampling particles using the gradient information of the posterior distribution, resulting in superior tracking performance with fewer particles. This design strikes a favorable balance between tracking accuracy and computational efficiency, making it particularly suitable for IoAuT devices operating in complex acoustic environments, even under non-Gaussian noise and reverberation conditions. Extensive simulations and real-world experiments demonstrate that the proposed method achieves an average root mean square error (ARMSE) below 0.5 m under non-Gaussian noise at SNRs as low as 0 ~ 5 dB.