Simple Transformer with Single Leaky Neuron for Event Vision
Himanshu Kumar, Aniket Konkar · 2025
There is an increasing interest in integrating the self-attention mechanism and transformer architecture in event-based vision. Several event-based and spiking neural network transformers have been proposed with complex architectures. We propose a simple transformer with a multi-head attention mechanism, a feature extractor, and a single leaky neuron. We have used ResNet, an effective feature extractor for RGB images, on frame-based event data. Experimental results show that our architecture surpasses many event and spiking-based methods, including transformers, and achieves competitive performance in DVS Gesture, N-MNIST, and CIFAR10-DVS datasets. Notably, we achieve an accuracy of 98.3% on the DVS Gesture and 99.3% on the N-MNIST dataset. Source code available at GitHub.