Prototype-aware Feature Selection for Multi-view Action Prediction in Human-Robot Collaboration

Bohua Peng, Bin Chen, He Wei, Visakan Kadirkamanathan · 2024

Semantic representation of actions is essential for the development of mutual cognition towards efficient human-robot collaboration. As such, it has become an exciting emerging venue for robots to understand human intentions with Transformers, as a family of highly expressive models. However, due to high computational complexity, local redundancy in video frames can add up to inference latency of Transformers. To address this issue, we propose a prototype-aware token pruning method, namely ProtoPrune, to select important features for efficient action recognition. Specifically, for a video sequence, we first encode knowledge encapsulated in keyframes as prototype representations with a pretrained Transformer. Next, these prototypes teach the pretrained Transformer to preserve important visual features, pruning task-irrelevant tokens for improved throughput. In our experiments, the probabilistic token pruning method reduces over 37% of GFLOPs, with a performance retention rate of 92.9% without retraining ViT backbones. Our visualization showcases the improved robustness of learned action representations.

Read the paper · More papers on PaperTik