A Framework to Query in Natural Language on Big Data Platform

B. Sowmiya, S. Saravanan, Srinivasa Rao Payyavula, Santanu Bhattacharjee, Munagala Revanth Kumar, Sonakshi Mugrai, Pratyusha Routh, Aadit Bhargava · 2024

Many dashboards have been developed in contemporary Big Data systems to help business subsidiaries make data-driven decisions. The fact that there is currently no interface that enables business stakeholders to conduct natural language queries directly on the Big Data platform in order to acquire timely insights and information highlights another crucial gap in the quest for efficient natural language querying. The abstract highlights the critical need for a natural language querying system inside the Big Data ecosystem, which would enable non- technical users to engage with and derive insightful information from complicated datasets using ordinary language. Most existing frameworks focus on finding the models best working on structured or unstructured data, this study additionally also focuses on data type. By bridging this gap and developing a user-friendly querying interface, stakeholders can improve data accessibility, speed decision-making processes, and foster a data-driven culture across their organizations. There are several types of files of data sets, types of data, and various algorithms that are used for making a query interface. The data sets can be numeric, textual, or string. The string data can contain a mixture of numeric and textual data. The research essentially shows the working of 2 models with different data sets of different data types: the BERT Model and the Lang chain model. Each model works with different accuracy on different types of data sets and a detailed explanation of the same is provided. The proposed framework successfully utilizes LLMs, namely BERT and Lang Chain to efficiently increase the throughput of queries written by users for databases of multiple data types at a best performance percentage of 91.35 %

Read the paper · More papers on PaperTik