The Di2Win Document Intelligence Platform

Afonso Ferreira, Cleber Zanchettin, Rômulo César Dias de Andrade, Byron Leite Dantas Bezerra · 2025

We present the Di2Win Document Intelligence Platform (DIP). This modular AI-driven pipeline transforms raw document images --- captured by scanners or mobile phones --- into structured data and business actions in a single pass. The system comprises five loosely-coupled micro-services: (1) image-quality verification using a contrast-invariant model that flags blur, skew, and illumination issues above 100 ms per page; (2) document classification via a Transformer-base model with layout embeddings, delivering top-k types with calibrated confidence; (3) information extraction through i) Dilbert, a multimodal Token-Layout-Language model fine-tuned on weakly-labeled forms or ii) Delfos, a Large Language Model Mixture of Experts fine-tuned with well-defined prompts; (4) DataDrift, a powerful rules engine to avoid inconsistent outputs concerning the business process; and (5) process automation orchestrated by a Business Process Model Notation (BPMN) plus a Robot Process Automation (RPA) engine that routes results to databases, APIs, or human-review queues. All AI components are orchestrated through a messaging service to control the information flow, and the application exposes REST/gRPC endpoints to communicate with outside consumers. This enables the hot-swapping of models without downstream code changes by plugging a new message consumer into the messaging system. This also provides horizontal scalability since to increase the application throughput, we only need to add new AI engine consumers to the messaging system. Deployed in banking, insurance, and healthcare, the Di2Win DIP has processed more than 30 million pages, reducing average handling time by 79% and re-keying errors by 86 %, speeding up the workflows up to ten times. Our DocEng demonstration allows attendees to upload documents, observe live quality and confidence dashboards, and edit extracted fields with immediate feedback to the active-learning loop.

Read the paper · More papers on PaperTik