Advanced Retrieval Augmented Generation for Local LLMs
Leonardo Marques Rocha, Rian Manoel Pessoa · 2024
This paper presents an extension of Retrieval Augmented Generation workflow for Large Language Models executing in a local processor with limited resources. The novelty presented is an improvement on such workflows to consider limitations in the total budget of token usage and privacy of data using local storage of data. This has the potential to leverage applications without the necessity of online services that can have a high cost and latency due the Chain of Thought used in most data retrieval cases. The presented workflow has a lightweight usage of computing and can be fully implemented in a low-resource compute environment.