Retrieval-augmented-generation-enhanced Dense Video Caption for Human Indoor Activities: Disambiguating Caption Using Spatial Information Beyond Field of View Constraints
Bin Chen, Yugo Nakamura, Shogo Fukushima, Yutaka Arakawa · Sensors and Materials · 2024
BLEU-3, METEOR, CIDEr, and Precision metrics, but shows a small decline in ROUGE-L and Recall metrics compared with GVL.These results demonstrate that our method effectively incorporates spatial information and reduces ambiguity for indoor human activity caption applications.