Image and Video Captioning with Augmented Neural Architectures

Rakshith Shetty, Hamed R. Tavakoli, Jorma T. Laaksonen · IEEE Multimedia · 2018

Neural-network-based image and video captioning can be substantially improved by utilizing architectures that make use of special features from the scene context, objects, and locations. A novel discriminatively trained evaluator network for choosing the best caption among those generated by an ensemble of caption generator networks further improves accuracy.

Read the paper · More papers on PaperTik