Machine Learning Model: Perspectives for quality, observability, risk and continuous monitoring

Diego Nogare, Ismar Frango Silveira, Pedro Pinheiro Cabral, Rafael Jorge Hauy, Veronica Leão Neves · 2024

The transition of machine learning (ML) and artificial intelligence (AI) projects from experimental stages to fully operational solutions presents substantial challenges. This is especially true for applications where these technologies play a critical role, demanding high-quality, reliable, and observable ML models. This paper explores the crucial aspects of continuous monitoring in ML models and emphasizes the need for a comprehensive approach that goes beyond technical development. It highlights that ensuring the reliability and robustness of deployed ML models requires a multifaceted framework encompassing data governance, model lifecycle management, and thorough team training. The paper addresses key aspects such as model quality, risk management, and the crucial role of observability in maintaining model stability and reliability in production environments. Using Itaú Unibanco as a case study, the paper showcases a robust model risk management approach and a dual monitoring system: an independent validation team oversees riskier models, while smaller models are monitored by their development team. The paper concludes by emphasizing the significance of a robust Model Risk Management (MRM) framework in the evolving landscape of AI and ML, particularly as these technologies become deeply integrated into various business operations. Highlighting that Itaú Unibanco’s rigorous approach to model quality, observability, low risk, and continuous integration aligns with the regulatory requirements set by the Brazilian central bank.

Read the paper · More papers on PaperTik