LLM Based News Research Tool Using LangChain with Enhancing Similarity Search and Token Limit
Raushan Kumar, Subbulakshmi P, Khushi Habbu · International Journal of Research Publication and Reviews · 2024
In the digital age, with an overwhelming abundance of news sources and information outlets, people increasingly struggle to effectively navigate and absorb news content.This has led to a growing need for advanced technologies that assist in the categorization, distillation, and extraction of insights from vast amounts of news data.This abstract introduces a novel approach known as LangChain, which leverages blockchain infrastructure and Language Model (LM) technology to develop a sophisticated news research tool.LangChain is a framework designed specifically for Large Language Models (LLMs) and offers several essential features for document processing.It includes multiple text loaders capable of handling various formats such as text files, CSVs, and URLs.Once documents are loaded, LangChain employs character and recursive text splitters to divide the text into manageable chunks.Text encoding is handled through the Hugging Face and OpenAI modules, which convert text into numeric vectors that can be stored in vector databases.These libraries provide the necessary transformations and embeddings to represent words numerically.LangChain incorporates fundamental principles of classical information retrieval (IR) for tasks such as retrieval, summarization, search, and keyword extraction.It integrates with FAISS, a library used for similarity search and clustering of dense vectors, facilitating efficient storage and retrieval from FAISS indexes.Techniques such as TF-IDF are utilized for generating similar search results.Additionally, LangChain integrates RetrievalQA with sources chain, a critical component where information chunks are processed through LLMs.Filtered chunks are used to trigger further processing and are continually refined.These refined sections are then used in summarization techniques, enhancing the effectiveness of subsequent queries.Ultimately, the final response is derived from the processed and condensed portions of the input query.