AI-Driven Multi-Modal Information Synthesis: Integrating PDF Querying, Speech Summarization, and Cross-Language Text Summarization

K. Suresh Manic, Ahmed-Al Balushi, Al-Bemani A.S., Saleh Al Araimi, N. Balaji, Uma Suresh, Asiya Najeeb · Procedia Computer Science · 2025

This research presents an innovative AI-powered system designed to revolutionize information retrieval and summarization by integrating multiple data modalities. The system is composed of three key components: a PDF Question Answering System, Speech-to-Text Summarization, and Multi-Source Text Summarization with Cross-Language Translation capabilities. The PDF Question Answering System processes documents by segmenting them into 5000-word chunks and generating embeddings using the all-MiniLM-L6-v2 model. These embeddings are then stored in ChromaDB, a specialized vector database, enabling precise querying through similarity searches. The Speech-to-Text Summarization feature converts audio files or live streams into text using the OpenAI Whisper model. It then creates a summary of the text, which can be translated into different languages, making it easier for users to access the information. The Multi-Source Text Summarization and Translation module further broadens the system’s capabilities by processing content from various sources, including PDFs, YouTube videos, and websites. This content is summarized using the Gemma-7b-It model, with additional translation options available to the user. Test results showed that the system accurately converts speech to text, even in difficult audio conditions, with very few mistakes. It creates clear and concise summaries that keep the important details and processes tasks quickly. For example, it converts 5 minutes of audio into text in about 2 seconds and answers questions from PDF documents in less than 12 seconds. These results demonstrate the system’s ability to handle large amounts of data and different types of content efficiently, making it a flexible tool for users who need to collect and summarize information from multiple sources. It also supports multiple languages through its built-in translation feature.

Read the paper · More papers on PaperTik