Sequence to Sequence Pre-Trained Model for Natural Language Processing

Gulshan Dhasmana, Prasanna Kumar H R, Guru Prasad M S · 2023

Pre-trained model is a transfer learning model trained on huge dataset. These model can be reused to solve new problems. Pre -trained model uses the encoder structure. Sequence - to Sequence model uses the encoder - decoder structure. Deep learning model can do the better task whenever large dataset is available but it may not be Sequence - to - Sequence pre - trained model (SSPT). SSPT model takes the chunks of random text and it will try to decode the clean text. For the natural language processing both BERT and BART can be used. BERT is more kind of bidirectional encoder structure with an objective of mask filling language, whereas BART is both encoder and decoder structure with left to right language modelling task. T5 is the text to text transformer which is aims to achieve state of the art result. mT5 is the massively multilingual text to text transformers which is the variant of T5. This paper focus more on SSPT model such as BART and T5 Model, The result indicates that BART performance is better over the other trained model. Prompting and fine tuning shall be used to improve the performance of the existing model.

Read the paper · More papers on PaperTik