AgeNet-AT: An End-to-End Model for Robust Joint Speaker Age Estimation and Gender Recognition Based on Attention Mechanism and Titanet
Mahsa Zamani Tarashandeh, Amirhossein Torkanloo, Mohammad Hossein Moattar · 2023
The potential uses of speaker age estimation in a variety of domains, such as forensics and human-computer interaction, have made it popular in recent years. However, the performance of the implemented techniques is significantly influenced by the noise and utterance length. In this study, the AgeNet-AT, which is based on Titanet model and attention mechanism, is proposed as a reliable age estimation and gender recognition model. To develop an end-to-end architecture for age estimation, the suggested method uses Titanet as the embedding extractor and an attention mechanism to determine the importance of each feature in the embedding vector. It is hypothesized that certain of Titanet’s extracted features may include characteristics that can distinguish between speakers of different ages and genders because Titanet is a model made to discriminate between different speaker identities. Titanet was therefore selected for this study’s embedding extractor. To further concentrate on the most important features for age estimate, an attention layer is employed. In addition, gender classification is included in the model as an auxiliary task to enhance estimation performance. On the TIMIT dataset, trials are run under multiple assessment conditions, such as variable utterance lengths and noise levels. The AgeNet-AT model’s robustness is demonstrated by the experimental findings. With Root Mean Square Error (RMSE) of 5.92 and 6.85 and Mean Absolute Error (MAE) of 4.30 and 4.73 for male and female speakers, respectively, the model exceeded the most recent age estimation results on the TIMIT dataset.