A Novel Approach to Text Summarization of Document using BERT Embedding
Kumkum S. Adwani, P. P. Shelke · 2023
This study examines and compares existing studies on various text summaries of documents and the processes associated with them. Despite the fact that the literature contains a large number of research contributions, we have critically and fully analysed current research and review papers that are relevant to text summary of document systems. The various approaches are classified based on the main ideas used in their procedures. The emphasis is on the concept utilised by the concerned authors, the approach used for experiments, and the performance evaluation measures. The researchers' claims are also emphasised. Our findings from the exhaustive literature review are presented together with the detected flaws. This study is very important for the comparative examination of various text summarizer approaches, which is necessary for addressing associated difficulties. On a review of the literature, we developed our own way in which we constructed a programme in Python and wrote it in the Spyder console using the Anaconda culture. The following standard libraries are used: NLTK for text processing, TensorFlow hub for getting the BERT pretrained model, correlation from the Sklearn library, and certain other standard libraries such as Panda for importing the dataset, Numpy, and so on. The data set must be imported after all of the libraries have been included. To summarise, the Kaggle datasets news summary training dataset was utilised. There are headlines, the full text, a synopsis, and links to related news stories for each story in the collection. In this experiment, the first 100 articles from the news summary collection were used. We analysed the outcome using the Rouge scoring system on both the created summary and the original summary. Rouge's average score for the first 100, 50, and 30 documents. The results obtained are encouraging.