With the boom of machine learning (ML) techniques, software practitioners are building ML systems to process massive volumes of streaming data for diverse software engineering tasks, such as failure prediction in AIOps. Trained on historical data, these models can suffer performance degradation due to concept drift, where the inter-relationship between features and labels (concepts) changes between training and production. Therefore, concept drift detection is essential to monitor deployed ML models and retrain them as needed. In this work, we explore applying state-of-the-art (SOTA) semi-supervised concept drift detection techniques on synthetic and real-world datasets in an industrial setting, which requires minimal manual labeling and maximal model compatibility. We find that current SOTA methods not only require significant labeling effort but also are model-specific. To address these limitations, we propose CDSeer, a novel model-agnostic technique to detect concept drift. Our evaluation shows that CDSeer outperforms the SOTA in precision and recall while requiring significantly less manual labeling. We demonstrate CDSeer's effectiveness at concept drift detection by evaluating it on eight datasets from different domains and use cases. Internal deployment of CDSeer on a proprietary industrial dataset shows a 57.1% improvement in precision while using 99% fewer labels compared to the SOTA method. Its performance is also comparable to the supervised (fully labeled) concept drift detection method. The improved performance and ease of adoption make CDSeer valuable in enhancing the reliability of ML systems.
Paper
References (90)
Scroll for more · 38 remaining