A Multibranch CNN Framework Based on Time–Frequency Feature Fusion for SSVEP Detection
Yiming Shao, Hongliang Zheng, Xiaogang Chen · IEEE Sensors Journal · 2025
Steady-state visual evoked potential (SSVEP) has emerged as a prominent paradigm in brain-computer interface (BCI) research due to its non-invasiveness, high signal-to-noise ratio, and robust time-locking/phase-locking properties. Current limitations in both complex frequency-encoded stimulation methods and neural networks constrained to single-domain feature extraction necessitate the development of advanced feature fusion frameworks. Although existing deep learning approaches attempt to integrate multi-domain features, they largely overlook the subject-specific variability in feature-domain contributions to classification outcomes. To comprehensively extract discriminative information while enhancing user adaptability, we proposed a novel filter bank multi-branch convolutional neural network (FB-MultiCNN) architecture for SSVEP signal detection. Our contributions are summarized as follows: First, a multi-branch network structure consisting of two dedicated subnetworks is employed to extract features from both the time domain and frequency domain. Secondly, we introduce a tunable feature fusion coefficient α, which allows the model to be custom-optimized for individual users, thereby significantly improving its generalizability and user-specific performance. Offline evaluation on public benchmarks achieved an accuracy of 85.51 ± 20.77% and information transfer rate (ITR) of 166.27 ± 53.34 bits/min within a 1.0-s window. Real-time online validation further demonstrated practical efficacy, yielding 96.70 ± 5.36% accuracy and 199.49 ± 19.82 bits/min ITR. By extracting time-domain and frequency-domain features, FB-MultiCNN significantly improves the performance of SSVEP-based BCIs.