Supporting transparency in AI systems: designing training dataset explanations and assessing AI literacy
Ariful Islam Anik · Mspace (University of Manitoba) · 2026
Training datasets play a central role in shaping the behavior, limitations, and impacts of AI systems, yet transparency about training datasets remains limited in practice. As a result, users can lack the context needed to meaningfully interpret system outputs, form informed judgments, and engage with AI systems effectively. One approach to addressing this lack of context is through training dataset explanations, which communicate key characteristics of the data used to train AI systems. This dissertation investigates how such explanations can support data-centric transparency in AI systems. As part of this exploration, the dissertation presents three complementary studies. The first two studies examine how explanation design factors influence how users use training dataset explanations in an onboarding setting and how they assess the systems that provide them. The first study examines the influence of presentation style by comparing a data storytelling–based presentation with a Q&A-based presentation, while the second study investigates the effects of information depth by comparing summary and detailed training dataset explanations. These studies show that, in an onboarding context, differences in presentation style and information depth lead to different patterns of use, perceived understanding, and trust. These findings highlight how design choices in training dataset explanations can shape users’ assessment of AI systems. The third study introduces and validates a dual-format AI literacy scale grounded in an established AI literacy framework. The study focuses on scale development and psychometric validation of its measurement properties. The scale combines self-assessed and demonstrated knowledge to support systematic assessment of users’ conceptual understanding of AI. Follow-up analyses using the scale reveal mismatches between perceived and demonstrated understanding, highlighting the limitations of relying solely on self-reported measures in explanation research. Overall, this work contributes to the field of human-centered explainable AI by foregrounding training datasets as a critical aspect of system transparency and by contributing both design insights and measurement tools to support explanation research. In doing so, it supports explanation design and evaluation practices that can better account for user differences and interaction contexts, advancing the long-term objective of effective human–AI collaboration.