Repairing Raw Data Files with TASHEEH
Mazhar Hameed, Gerardo Vitagliano, Fabian Panse, Felix Naumann · ACM SIGMOD Record · 2025
Data files serve as a vital resource for all data-driven applications. Among these, comma-separated value (CSV) files are particularly popular with users and businesses due to their flexible standard. However, also due to this loose standard, the data in these files are often indeed ''raw'', fraught with many types of structural inconsistencies that hinder seamless ingestion into a data system. We say that rows in CSV files with such structural inconsistencies are ill-formed. Traditionally, data practitioners write custom code to repair the structure and format of ill-formed rows, even before they can leverage data cleaning tools and libraries, which typically assume that data are already properly loaded. Writing such code and configuring loading scripts is tedious, time-consuming, and requires both expertise and frequent human intervention. To address these challenges, we present Tasheeh - a system that automatically detects ill-formed rows containing data and then standardizes their structure into a uniform format based on the structure of well-formed rows. By automating these essential steps, our system frees up valuable time and resources, enabling practitioners to focus on downstream stages of the data processing pipeline.