In many applications of machine learning (ML), updates are performed with the\ngoal of enhancing model performance. However, current practices for updating\nmodels rely solely on isolated, aggregate performance analyses, overlooking\nimportant dependencies, expectations, and needs in real-world deployments. We\nconsider how updates, intended to improve ML models, can introduce new errors\nthat can significantly affect downstream systems and users. For example,\nupdates in models used in cloud-based classification services, such as image\nrecognition, can cause unexpected erroneous behavior in systems that make calls\nto the services. Prior work has shown the importance of "backward\ncompatibility" for maintaining human trust. We study challenges with backward\ncompatibility across different ML architectures and datasets, focusing on\ncommon settings including data shifts with structured noise and ML employed in\ninferential pipelines. Our results show that (i) compatibility issues arise\neven without data shift due to optimization stochasticity, (ii) training on\nlarge-scale noisy datasets often results in significant decreases in backward\ncompatibility even when model accuracy increases, and (iii) distributions of\nincompatible points align with noise bias, motivating the need for\ncompatibility aware de-noising and robustness methods.\n
Paper
References (51)
Scroll for more · 38 remaining