ViT-Mixer: A Robust Automatic Modulation Classification Model Based on Vision Transformer
Boyu Xu, Xinfeng Gao, Li Guo · 2024
Recently, an increasing number of automatic modulation classification techniques are utilizing deep learning to classify the modulation types of input shortwave signals. Existing research has made significant progress in feature extraction from signals using deep learning. However, many deep learning networks require high computational and parameter costs, making them expensive for practical deployment. Therefore, we chose the Vision Transformer (ViT) as the foundation for our signal recognition model. We improved it with a new structure based on basic matrix multiplication, data dimension transformation, and scalar nonlinear operations, resulting in a new model: Vision Transformer Mixer (ViT-Mixer). Moreover, existing models are often validated only on public datasets that lack several common signal types, limiting their effectiveness in real-world applications. To address this, we constructed a new dataset through actual collection and simulation methods to adapt the model to our task. We tested several models on this new dataset. Experimental results show that our proposed model, ViT-Mixer, achieves satisfactory accuracy in classifying specific signals in our dataset, while significantly reducing the computational and parameter costs compared to the original ViT.