Transformer Network for video to text translation

N Mubashira, Ajay James · 2020 International Conference on Power, Instrumentation, Control and Computing (PICC) · 2020

Recently generation of natural language descriptions for videos has created a lot of focus in computer vision and natural language processing research. Video understanding involves detecting scene's visual and temporal elements and reasoning it for description generation.Several real world implementations such as video indexing and retrieval,video to sign language translation etc, are based on this.Because of the complicated nature and diversified content, the captioning problem becomes more challenging.This is a machine translation problem which uses encoder decoder architecture of GRU or LSTM to deal with this kind of problem.But here the decoding starts with the final hidden state of encoder as input.This can't be a good summary of the input sequence,because all the intermediate states of encoder are ignored.This paper proposes a transformer network with deep attention based encoder and decoder to generate the natural language description for video sequence data. This network processes the sequences as a whole and learns relationship between each elements in the sequence by providing attention.

Read the paper · More papers on PaperTik