A framework for high performance image analysis pipelines

Raúl Ramos-Pollán, Ángel Cruz-Roa, Fabio A. González · 2012

This paper describes the software framework being developed to enable the execution of large-scale image analysis pipelines. Images are analyzed through algorithms (feature extraction, annotation, classification, etc.) assembled into processing pipelines and managed by the framework to be run onto the available computing resources, whether cloud, opportunistic or dedicated clusters. Underneath, we use Google's big table storage model for image sources and metadata, showing both flexibility and performance. Additionally, our architecture provides a clear separation between framework developers, providers of algorithms and experimenters, enabling the organization of teams and software repositories to ensure its organizational sustainability in the long term. Altogether, we integrate best practices in pattern recognition, software engineering and high performance computing to enable large scale experiments in image analysis. We herewith describe the framework and present preliminary results demonstrating its scalability and ease of use.

Read the paper · More papers on PaperTik