Scalable and Cost-effective Serverless Architecture for Information Extraction Workflows
Dheeraj Chahal, Surya Chaitanya Palepu, Rekha Singhal · 2022
Information extraction from an image or scanned document is a complex and challenging process since it involves recognizing various visual structures such as tables, boxes, logos, text, charts, etc. Hence, the content extraction applications contain a pipeline of multiple computer vision algorithms, APIs, and models. Deploying such applications for document processing requires a resilient system to deliver high performance. Such applications can be deployed on cloud to leverage the flexible infrastructure and multiple supporting services available there.