Smart object counter: a lightweight RGB-T crowd counting framework with adaptive fusion
Kenneth Lemuel · DR-NTU (Nanyang Technological University) · 2026
Crowd counting systems in urban surveillance cannot afford to fail after dark. Especially that in Singapore, where camera networks spread across transport hubs, shopping malls and public walkways, they need to function reliably day and night. But the problem is that RGB cameras lose useful detail under low light, which is exactly when continuous monitoring matters most. This is where thermal imaging comes in by helping pick up infrared radiation from people regardless of lighting conditions, yet the models that use both modalities most effectively tend to be too large and slow for practical usage on standard hardware. This paper looks at whether that tradeoff can be resolved. The model proposed in this paper is Adaptive FPN Lite + Calibration, running both RGB and thermal channels through on shared backbone rather than two separate heavy streams. Then a small feature pyramid with per-level gating lets the model adjust how much it relies on different scales depending on the scene presented, and a calibration branch at the end nudges the final count closer to its true value. All of this fits into roughly 16.4 million parameters compared to the 40.6 million for the state-of-the-art Broker Modality model on the same benchmark. Getting to this design took testing nine model variants along the way, including two earlier adaptive designs that were dropped after they had failed to beat a much simpler baseline. On the official RGBT-CC test split, the Adaptive FPN achieves an MAE of 22.127 and RMSE of 37.938, which is a 43 percent improvement over the RGB-only baseline at MAE 38.711. On speed, it runs at 66.91 frames per second against the Broker Modality model’s 9.75, which is roughly 7 times faster at only 60 percent of the parameter count and noticeably less memory usage. Per-scene measurements across five crowd-density ranges in Section 5.6 back this up, showing the speed gap holds whether the scene is sparse or packed. What this project ultimately shows is that one does not need a heavy dual-stream architecture to get useful RGB-T counting, and that a single shared backbone with per-level gating and a small calibration step gets you competitive accuracy at a fraction of the cost.