Development of an NLP-Based Automatic Data Retrieval Model

P Thejas, S. Anupama Kumar, Y S Kiran Kumar · 2023

Understanding and analysing data are essential in today's data-rich environment. Users can engage with modern programmes by speaking or typing commands; they are designed to be user-friendly. This work's primary objective is to record audio, convert it to text, preprocess the text, and then pull out relevant information. With the use of natural language processing, machines can currently comprehend and generate human language (NLP). Using deep learning libraries, algorithms, and language rules, NLP-based query retrieval aims to comprehend the semantics of the query and match it with the most pertinent facts in a dataset. The complexity of the data, including its sparsity, diversity, and dimensionality, as well as the dynamic nature of datasets, present the key hurdles in this evolution. The technologies and tools used include Speech Recognition API for audio-to-text conversion, SpaCy and NLTK (Natural Language Toolkit) is a popular Python library used for NLP (Natural Language Processing). NoSQL databases for scalable data storage, and information retrieval techniques like BM25F and vector space model. The application's goal is to deliver pertinent solutions to user inquiries in a variety of formats, and its performance is assessed by contrasting several algorithms to identify the most precise and efficient strategy for the best outcomes.

Read the paper · More papers on PaperTik