Parallel Feature Fusion for Multimodal Scene Recognition on Dual-Core MCUs
Kieran Woodward, Eiman Kanjo · IEEE Pervasive Computing · 2025
Multimodal sensing promises more robust environmental understanding for pervasive computing applications, but implementing sophisticated sensor fusion on resource-constrained devices remains challenging. We present a novel approach that leverages dual-core microcontrollers to enable the parallel processing of visual and audio data for scene recognition. Our system concurrently executes specialized neural networks across separate cores while efficiently fusing their intermediate features, achieving 93% classification accuracy across nine different environments. This represents a 12.27% improvement over single-modality approaches while reducing latency by 48 ms compared to sequential processing. The entire system operates within just 258 KB of memory, demonstrating that complex multimodal AI is achievable even on highly constrained devices. By enabling sophisticated environmental understanding without cloud connectivity, this work advances privacy-preserving edge intelligence for a wide range of future applications.