Diffusion-grounded VideoLLM for Entity-aware temoporal grounding

Pengcheng Fang, Yuxia Chen, Wang, Yong, Cai, Xiaohao, Ruan, Xiangning · arXiv (Cornell University) · 2025

Precise temporal grounding requires distinguishing when a queried event occurs from when its participating entities are merely visible. We propose Diffusion-Grounded VideoLLM, which conditions temporal feature extraction on query-relevant entities before language reasoning. The framework tracks entities named in the query and uses their masks to condition a frozen video diffusion backbone. Intermediate spatiotemporal features are extracted through truncated denoising and combined with entity tokens and timestamp embeddings. The language model uses this evidence together with the full query to generate temporal intervals and answers to grounded questions. On Charades-STA and NExT-GQA, the model obtains 43.5 mIoU and 28.4 Acc@GQA, improving the reported Grounded-VideoLLM reference by 6.7 and 1.7 points, respectively. Component and entity-pathway ablations support the usefulness of conditioning diffusion features on query-relevant entities for temporal grounding.

Read the paper · More papers on PaperTik