D3.1 XAI methods and Benchmark Suite for Human-AI systems v1

George Makridis, Philip Mavrepis, Maria Margarita Separdani, Christos Diou, George Fragiadakis · Zenodo (CERN European Organization for Nuclear Research) · 2025

This document presents the progress made in the HumAIne project under Work Package 3 (WP3), which focuses on Explainable and Interpretable AI for Human-AI collaboration (HAIC). WP3 plays a key role in designing tools and methodologies that align with HumAIne's vision of a human-centered and ethical development platform for digital and industrial technologies in healthcare, finance, smart cities, manufacturing, and energy. The deliverable provides an overview of the work conducted to develop glass-box models, explainable frameworks, and benchmarking tools. These efforts collectively aim to enhance transparency, trustworthiness, and usability in Human-AI interactions. The work reported in this deliverable is a collaborative effort among all HumAIne consortium partners. It represents the initial phase of an incremental development process, capturing the state of WP3 components as of the submission date. The deliverable emphasizes WP3's integration into the HumAIne Reference Architecture (RA) and its alignment with the platform's goals. Future updates will incorporate feedback, ongoing technical advancements, and pilot-specific validations, ensuring a comprehensive and robust framework by the project's conclusion. The development of WP3 tools leverages the project's RA, as established in D2.3, which employs the architectural view model. This includes Logical, Process, Development, Physical, and Scenario views, ensuring consistency and alignment across all HumAIne components. The deliverable highlights the role of WP3 in supporting advanced learning paradigms (Active Learning (AL), Swarm Learning (SL), and Neurosymbolic AI (NSAI)) and integrating explainable AI (XAI) techniques to meet diverse user needs. These contributions aim to bridge the gap between technical sophistication and human accessibility, fostering trust and understanding in AI-driven systems. Key contributions detailed in this deliverable include: Development of Glass-Box Models (T3.1): Creation of inherently interpretable models, such as Concept Bottleneck Models (CBMs), alongside custom explainability methods like WeakSpot Analysis and Conformalized Quantile Regression, designed to improve transparency and reliability. Explainability for Black-Box Models (T3.2): Introduction of an XAI framework tailored to user needs, offering explanations for complex AI paradigms and pilot-specific applications. Benchmarking Suite (T3.5): Establishment of a comprehensive evaluation framework to measure the performance, transparency, and adaptability of HAIC processes. This suite spans horizontally across the platform, ensuring consistent validation and iterative improvement. WP3 tools are designed as Dockerized microservices, ensuring modularity, scalability, and seamless integration into the HumAIne platform. These tools interact with the platform's Workflow Definition, Model Registry, and Deployment Layer, supporting both end-to-end pipelines and real-time evaluations. Additionally, the Benchmarking Suite evaluates all Human-AI interactions, encompassing the use of advanced AI paradigms and user-centered explainability tools.

Read the paper · More papers on PaperTik