Automatic FAQ Generation Using Text-to-Text Transformer Model

Santosh Vasisht, Varun Tirthani, Akhil Eppa, Punit Koujalgi, Ramamoorthy Srinath · 2022 3rd International Conference for Emerging Technology (INCET) · 2022

Given a passage of text, the manual generation of FAQ’s (Frequently Asked Questions) is a very tedious process. The ability to automate the generation of relevant and useful FAQ’s along with appropriate answers within an acceptable time frame is an interesting problem to work upon. Due to the ever-increasing data and content on the Internet, netizens can leverage FAQ’s and their corresponding answers for better understanding, especially for technical and corporate websites. Previous research in this domain has been promising and we intend to add novel features to enhance the quality of QA (question-answer) pairs generated. In this paper, we propose an architecture that performs end-to-end processing of the input text and displays QA pairs with the help of an automated QA generation pipeline. The pipeline consists of a pre-trained Text-To-Text Transformer (T5) Model for span extraction and subsequent question-answer generation. A fine-tuned SpanBERT Model is utilized for answer generation and ranking of the QA pairs. The architecture also incorporates the simultaneous computation of the coherence factor between input sentences along with the usage of Knowledge Graph for additional domain knowledge. A web application was also developed to demonstrate the generation of a comprehensive list of FAQ’s for any user-provided passage. The proposed approach improves on existing pipelines by having BLEU scores in excess of 0.4 for answers and 0.6 for questions.

Read the paper · More papers on PaperTik