PCATNet: Position-Class Awareness Transformer for Image Captioning

Ziwei Tang, Yaohua Yi, Changhui Yu, Aiguo Yin · Computers, materials & continua/Computers, materials & continua (Print) · 2023

Existing image captioning models usually build the relation between visual information and words to generate captions, which lack spatial information and object classes. To address the issue, we propose a novel Position-Class... | Find, read and cite all the research you need on Tech Science Press

Read the paper · More papers on PaperTik