As AI accelerators become integral in various domains from healthcare to autonomous systems, the reliability and verifiability of their data pipelines have become critical. Current data pipelines for AI accelerators are typically opaque, lacking systematic mechanisms to ensure data provenance, integrity, and explainability at every stage. This paper presents a data-centric framework named Trust-Aware Data Pipelines (TDP) for AI accelerators, structured around five core stages: Data Ingestion, Lineage Tracking, Explainable Preprocessing, Conformance Verification, and Deployment Validation. By embedding robust validation methods, systematic lineage tracking, and explainable data transformations, TDP ensures comprehensive end-to-end visibility. The framework is demonstrated through two practical case studies—image classification and sensor fusion tasks—highlighting significant improvements in traceability, reproducibility, and accountability, with minimal performance overhead. The proposed approach provides a foundation for building verifiable, trustworthy AI systems capable of meeting stringent safety and accountability requirements in critical applications.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex