Leveraging Large Language Models for Generating Training Datasets for Text Extraction from Thumbnails
Eric Xu, Chimezie Onwuegbuchulem, Sarfraz Shaikh, Lin Deng · 2025
Obtaining datasets for training AI models can often be an expensive endeavor in the domain of digital forensics. This research leverages large language models (LLMs) to automate the creation of a training dataset aimed at the extraction of text in Word thumbnails. Our method unfolds in three stages: initially, an LLM generates the text for a Word document based on a randomly chosen title. The generated text acts as the training label for a training instance. Subsequently, we apply a predefined Word format template, which organizes the document into sections with specified fonts and sizes. This content is integrated with the template to produce a formatted Word document. From this document, we extract the thumbnail and also capture a high-resolution screenshot. Therefore, each training instance comprises two types of labels (i.e., a high-resolution image and text label of a Word document) and one blurry thumbnail of the Word document. Through this method, we successfully generated 10 datasets. Each dataset contains the same font with 30K training instances. The 30k instances are further divided into three groups in terms of three different thumbnail resolution sizes: medium, large, and extra large. Each group has 10k instances.