GenAI based Data Extraction with Query-Based Insights
Swati Shilaskar, Sharayu Chakole, Amit Samarth, Prem Shejole · 2025
An innovative method has been proposed to enhance data extraction and analysis from PDF files by combining OCR technology with advanced AI models. The process begins with a PDF Reader module that employs OCR to extract text from PDFs, addressing challenges posed by the unstructured nature of document content. To tackle these difficulties, the Gemini LLM model transforms semi-structured or unstructured text into structured JSON formats. Once the data is structured, the Pandas AI component leverages the GROQ and Mistral LLM models to generate visualizations such as charts, graphs, and tables. These visual representations provide a more intuitive and insightful way to interpret the extracted data. This approach not only ensures more accurate data extraction but also significantly enhances the analysis and comprehension of complex documents.