Language-guided visual retrieval

He Su · 2021

List of Tables3.1 Comparison with recent state-of-the-art approaches on three popular datasets RefCOCO, RefCOCO+, and RefCOCOg with ground-truth information (Acc%).We achieve state-of-the-art performance on all datasets, demonstrating the effectiveness of our model. . . . . . . . . . . . . . . . . . . .27 3.2 Comparison with state-of-the-art approaches on automatically detected proposals of Faster R-CNN (Acc%).SVMN also achieves state-of-the-art performance, which proves the robustness of our model. . . . . . . . . . .27 3.3 Ablation studies of the main modules of SVMN (Acc%).'w/o' represents replacing the module by a straightforward replacement or removing it directly.The results of erased models are compared with that of the full model and prove the effectiveness of the Residual Attention Parser, the symbolic module, and the relation network in the visual module. . . . . .28 4.1 Performance comparisons of our approach with existing state-of-the-art methods. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .46 4.2 Ablation results of the layer composition analysis. . . . . . . . . . . . . .47 4.3 Comparison of results of sub-graphs ablation in Hybrid Graph Network. .48 4.4 Results of ablation experiments of the feature analysis. . . . . . . . . . .

Read the paper · More papers on PaperTik