PCATNet: Position-Class Awareness Transformer for Image Captioning
Ziwei Tang, Yaohua Yi, Changhui Yu, Aiguo Yin · Computers, materials & continua/Computers, materials & continua (Print) · 2023
Existing image captioning models usually build the relation between visual information and words to generate captions, which lack spatial information and object classes. To address the issue, we propose a novel Position-Class... | Find, read and cite all the research you need on Tech Science Press