BiFuseNet: A Multimodal Network for Estimating Blood Alcohol Concentration via Bidirectional Hierarchical Fusion

Abdullah Tariq, Arooba Maqsood, Martin Mašek, Syed Zulqarnain Gilani · 2025

Drunk driving remains a significant public safety challenge, demanding innovative alternatives to conventional methods such as field sobriety tests and breathalysers. Estimating a driver's level of intoxication through facial cues is particularly challenging due to the subtle and person-specific nature of alcohol-induced behaviours. In this paper, we present BiFuseNet, a 3D spatio-temporal multi-modal network designed to classify alcohol impairment levels into three categories: sober, moderate, and severe. Unlike prior approaches that rely on either uni-modal RGB video or hand-crafted facial features, our method exploits complementary physiological cues from RGB and infrared (IR) facial videos. We introduce a Bi-directional Hierarchical Fusion (BiHF) module that applies cross-attention mechanisms at multiple semantic levels of our BiFuseNet, including early, middle, and late feature stages. This enables deep integration of modality-specific signals across varying temporal and spatial contexts. To capture both short-term facial movements and sustained facial dynamics, we implement a sliding window strategy that samples over 30 frames across ten-minute recordings. Extensive experiments on a public dataset demonstrate that BiFuseNet outperforms uni-modal and traditional fusion baselines, achieving a classification accuracy of 88.41% and an AUC-ROC of 0.91, establishing a new state of the art in estimating blood alcohol concentration.

Read the paper · More papers on PaperTik