Data transformation processes in Extract, Transform, Load (ETL) pipelines are crucial in creating the inputs to AI and analytics systems, despite the fact that they often operate as "black boxes" with little transparency. This paper presents Explainable ETL, a platform for transparent, traceable, and comprehensible data transformations. We explore the integration of lineage tracking, semantic annotations, and interpretability tools like as SHAP, LIME, and metadata graphs into ETL orchestration to enhance auditability, bias detection, and regulatory compliance. The suggested architecture's direct integration of explainability modules into ETL tools allows data engineers and business users to understand why data appears as it does at each stage of the pipeline. According to experimental results, explainable ETL reduces bias propagation by 65% and improves error traceability by 92%. This tactic encourages more accountability and trust in data-driven systems by bridging the gap between responsible AI and data engineering.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex